ConceptioArchivearXiv CS
arXiv CSopen access

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective Sijie Mai* Shiqin Han School of Computer Science, South China Normal University {sijiemai,20222121019}@m.scnu.edu.cn

arXiv:2604.18460v1 [cs.LG] 20 Apr 2026

Abstract Multimodal affective computing aims to predict humans’ sentiment, emotion, intention, and opinion using language, acoustic, and visual modalities. However, current models often learn spurious correlations that harm generalization under distribution shifts or noisy modalities. To address this, we propose a causal modality-invariant representation (CmIR) learning framework for robust multimodal learning. At its core, we introduce a theoretically grounded disentanglement method that separates each modality into ‘causal invariant representation’ and ‘environment-specific spurious representation’ from a causal inference perspective. CmIR ensures that the learned invariant representations retain stable predictive relationships with labels across different environments while preserving sufficient information from the raw inputs via invariance constraint, mutual information constraint, and reconstruction constraint. Experiments across multiple multimodal benchmarks demonstrate that CmIR achieves state-of-theart performance. CmIR particularly excels on out-of-distribution data and noisy data, confirming its robustness and generalizability.

1

Figure 1: A case study on CMU-MOSI. Vanilla model without causal inference makes incorrect prediction for the test sample where the speaker delivers negative comment while smiling, while CmIR accurately predicts the label based on correct causal relationships.

test data distributions differs from training distributions under distribution shifts or noisy modality conditions (Zhuang et al., 2025; Peters et al., 2016). This limitation restricts their practical deployment in real-world scenarios where environmental conditions, speaking styles, lighting conditions, and background contexts constantly vary (Arjovsky et al., 2019). For instance, models might over-rely on the speaker’s consistent smile (a spurious visual feature) in training data rather than focusing on semantic content and genuine emotional expressions, leading to performance degradation when tested on different speakers (see Figure 1). Similarly, noisy modalities (e.g., background noise or low-resolution visual frames) further disrupt spurious correlations, exacerbating the generalization gap. To address this limitation, some approaches focus on domain adaptation to learn shared representations across domains (Dai et al., 2020; Zhu et al., 2025; Zhang et al., 2022), while others employ disentanglement strategy to separate different factors in multimodal data (Hazarika et al., 2020; Zhuang et al., 2025). However, these methods lack causal interpretation and cannot guarantee the disentangled/learned features align with causal and spurious components. Recent works incorporate causal inference principles to identify and eliminate spurious correlations (Xu et al., 2025b; Jiang

Introduction

Multimodal affective computing (MAC) aims to integrate information from language, acoustic, and visual modalities to predict high-level semantics such as sentiment, emotion, intention, and opinion (Mai et al., 2025; Poria et al., 2017). Recently, MAC has achieved remarkable progress and become increasingly important for applications such as human-computer interaction, customer service automation, and affective computing systems. Despite remarkable advances, existing approaches often learn spurious cross-modal correlations from training data rather than genuine causal relationships, harming model generalization when * Corresponding Author

1

• Empirical validation: Extensive experiments on multiple multimodal tasks show that CmIR achieves state-of-the-art results in both standard and out-of-distribution (OOD) benchmarks. Moreover, it exhibits superior robustness under noisy modality testing.

et al., 2025; Sun et al., 2022). However, they often lack rigorous theoretical analysis or focus on specific biases (such as speaker bias and modality bias) rather than providing a general framework, which rely on prior knowledge or assumptions about the biases and may not generalize well to unseen environments or other types of distribution shifts. To resolve these issues, we propose a Causal modality-Invariant Representation (CmIR) learning framework for robust multimodal learning. Drawing on causal inference principles (Arjovsky et al., 2019; Peters et al., 2016), CmIR is theoretically grounded in causal inference and information theory, with a novel feature disentanglement method that separates each modality into two complementary components: (1) invariant causal represeninv ), which carries stable causal relationtation (Zm ships with labels across different environments; and (2) environment-specific spurious representation spu (Zm ), which captures non-causal, environmentdependent noise and has no causal link to the label. The core insight is that while the distribution of raw features may vary across environments, the causal mechanisms between invariant features and prediction targets remain stable. CmIR is achieved through an elegant objective function that incorporates: (1) an invariance constraint to ensure invariant representations maintain stable relationships with labels across environments; (2) a mutual information constraint to minimize correlations between invariant and spurious components; and (3) a reconstruction constraint to preserve sufficient information from raw inputs. These constraints guarantee that invariant representations satisfy critical properties: environmental independence and causal sufficiency for prediction. We then use invariant modality representations for prediction, ensuring robustness to distribution shifts and noisy modalities. Compared to previous methods (Xu et al., 2025b; Yang et al., 2024), CmIR does not rely on specific bias or assumption. It directly learns invariant representations that are stable across all environments, which is more general and applicable to various distribution shifts. Our contributions are summarized as follows:

• Theoretical guarantees: We prove the existence and extractability of invariant representations given multi-environment training data. We also show that predictors based on invariant representations achieve lower worst-case OOD risk than those using raw features.

2

Related Work

2.1

Multimodal Affective Computing

Most works for MAC center on devising fusion techniques to learn discriminative multimodal representations (Zadeh et al., 2018a; Wang et al., 2025; Zadeh et al., 2017) or employing techniques such as information bottleneck to regularize unimodal distributions (Shankar, 2022; Mai et al., 2023c; Luo et al., 2025b). Recently, multimodal large language models have enabled the direct processing of multimodal signals using large pre-trained models, enhancing the interpretation of human affective states (Zhao et al., 2025; Xu et al., 2025a). However, these methods often neglect to improve the generalizability of models for OOD data. To enhance generalization and robustness, some methods focus on domain adaptation to learn shared representations across domains (Dai et al., 2020; Zhu et al., 2025; Zhang et al., 2022), while others employ disentangled learning to separate different factors in multimodal data (Yang et al., 2023; Hazarika et al., 2020; Zhuang et al., 2025; Tsai et al., 2019b). However, these methods lack causal interpretation and cannot guarantee the disentangled/learned features align with causal and spurious components. 2.2

Causal Inference

Causal inference can detect and remove non-causal associations in complex datasets to improve model robustness and generalization (Wang et al., 2022; Niu et al., 2021). Many causality-based methods have been introduced to reduce cross-modal bias in multimodal learning. Researchers employ counterfactual reasoning to refine attention distributions (Huang et al., 2025), apply front-door and back-door adjustments to decouple spurious links between text and vision (Liu et al., 2023), develop counterfactual and debiasing frameworks (Sun

• Methodological innovation: We propose a novel framework name CmIR, which, for the first time, systematically disentangles each modality into causal and spurious components in MAC to comprehensively learn causalitysufficient invariant representations. 2

causally influenced only by invariant components inv } and is independent of E and {Z spu }. {Zm m Theorem 1 (Definition of Causal Invariant Modality Representations). Assume there exists a funcinv } and a distrition class Φm = {ϕm : Xm → Zm bution distance measure D such that the following optimization problem has a solution: ϕ∗m = arg min max D(P (Y |ϕm (Xm ), E = e1 ),

Figure 2: The SCM of CmIR for prediction process. It includes language (l), visual (v), acoustic (a) modalities.

ϕm ∈Φm e1 ,e2 ∈E

P (Y |ϕm (Xm ), E = e2 )) Then constitutes a causal invariant modality representation satisfying: P (Y |ϕ∗m (Xm ), E = e1 ) =P (Y |ϕ∗m (Xm ), E = e2 ),

et al., 2022; Huan et al., 2024; Sun et al., 2023), and design causal intervention modules to separate misleading connections between expressive style and semantic content (Xu et al., 2025b). However, most methods are restricted to single or specific modality pairs, or they rely on explicitly annotated bias types that require domain knowledge (Nam et al., 2020). This limitation hinders their broader application to complex multimodal data where biases are often implicit and not predefined. In contrast, we propose a general invariant representation learning framework without requiring predefined bias types. Moreover, Invariant Risk Minimization (Arjovsky et al., 2019) and Invariant Causal Mechanism of CLIP (Song et al., 2025) that aim to learn invariant features are related to our work, but they focus on single modalities or specific modality combinations in particular application scenarios. In contrast, we provide a more general multimodal causal framework with feature disentanglement to understand and learn the properties of invariant features more comprehensively and accurately.

3

ϕ∗m (Xm )

∀e1 , e2 ∈ E Proof. See Section A.1 for the proof. Theorem 1 suggests that ϕ∗m (Xm ) is a valid invariant modality representation capturing only causal features for robust prediction. If we can learn a representation ϕ(Xm ) such that the conditional distribution of the label Y given this representation is the same across all environments (i.e., P (Y |ϕ(Xm ), E = e) does not depend on e), then this representation must capture only the causal features (i.e., causal features exist). Intuitively, causal relationships between features and labels are invariant under changes of the environment. If the prediction rule based on ϕ(Xm ) remains unchanged when the environment varies, it means that ϕ(Xm ) does not contain any environment-specific spurious information. In other words, it blocks all backdoor paths from environment E to label Y . Thus, the invariance condition is a signature of causality.

Theoretical Analysis

Here we establish a theoretical foundation for our causal approach to multimodal learning. We first establish the existence and extractability of causal invariant modality representations, then prove their advantages in terms of generalization performance. 3.1

3.2

Extraction of Invariant Representations

While Theorem 1 establishes the definition of invariant representations, practical implementation requires extracting these representations from raw modalities while preserving all relevant information. This motivates our disentanglement approach.

Definition of Invariant Representations

We begin by formalizing the causal structure of multimodal learning. Consider M modalities {Xm }M m=1 and prediction target Y . Following the structural causal model (SCM) framework (Peters et al., 2016), we assume the data generation process involves an environment variable E that induces distribution shifts, and each modality Xm can be decomposed into an invariant component inv (containing causal features) and a spurious Zm spu component Zm (containing environment-specific features). As shown in Figure 2, we assume Y is

Theorem 2 (Theoretical Guarantee for Extracting Disentangled Representations). Consider encoder inv , Z spu ) and decoder functions gm : Xm → (Zm m spu inv functions rm : (Zm , Zm ) → Xm that optimize the following objective: min

{gm ,rm ,h}M m=1

+ λ1

M X m=1

3

inv M Ee∈E [Lpred (Y, h({Zm }m=1 ))]

(m)

Rinv + λ2

M X m=1

(m)

Rdec + λ3

M X m=1

(m)

Rrec

· · ·×XM is the M -modal feature space and Y is the label space. Let Eall denote the set of all possible environments, each corresponding to a distribution P e (x, y). Let hinv ∈ H be a predictor using ininv }M , and variant representations Z inv = {Zm m=1 hraw ∈ H be a predictor using raw multimodal representations X = {Xm }M m=1 . Assume: 1. Invariance Condition: The invariant repre′ sentations satisfy P e (Y |Z inv ) = P e (Y |Z inv ) for all e, e′ ∈ Eall . 2. Information Sufficiency: The mutual information between invariant representations and raw features satisfies I(Z inv ; X) > c for some constant c > 0 (Z inv contains enough information from X). 3. Loss Function Regularity: The loss function ℓ is L-Lipschitz continuous and bounded. Then the worst-case OOD risk satisfies:

where Lpred is the task prediction loss, h is the prediction head, P P (m) inv Rinv = = e1 ∈E e2 ∈E D(P (Y |Zm , E inv e1 ), P (Y |Zm , E = e2 )) enforces invariance, (m) inv ; Z spu ) minimizes mutual informaRdec = I(Zm m tion between invariant and spurious components, (m) inv , Z spu )∥2 ensures Rrec = ∥Xm − rm (Zm m reconstruction capability, and λ1 , λ2 , λ3 > 0 are hyperparameters. Assuming the label Y is independent of environment E, the function classes {gm , rm , h} have sufficient capacity and the data follows the SCM described in Section 3.1, then as λ1 , λ2 , λ3 → ∞, the optimal solution satisfies:   inv spu 2 1. lim E ∥Xm −rm (Zm , Zm )∥ = 0 (perλ3 →∞

fect reconstruction is achieved) inv 2. Zm ⊥ ⊥ E (invariant component is environment-independent)

ROOD (hinv ) < ROOD (hraw )

spu inv |Zm , E) = 0 (spurious component 3. I(Y ; Zm

contains no additional causal information)

where ROOD (h) = maxe∈Etest Re (h) and Re (h) = E(x,y)∼P e [ℓ(h(x), y)].

Proof. See Section A.2 for the proof. Theorem 2 suggests that causal invariant representations can be learned using our CmIR. The invariance constraint forces Z inv to have the same predictive relationship with Y across environments, making it environment-independent (capturing only causal features). The reconstruction loss ensures that the pair (Z inv , Z spu ) retains all information from the original input, preventing information loss and avoiding degenerate decomposition solutions. The mutual information minimization pushes Z inv and Z spu to be statistically independent, so that Z spu cannot carry any causal information about Y that is already in Z inv . Reconstruction loss and mutual information minimization together force Z inv to contain all causal information in the original input. These three constraints lead to a clean decomposition: Z inv contains only causal factors and encompasses all causal information from the original input, and Z spu contains only environment-specific noise. 3.3

Proof. See Section A.3 for the proof. Theorem 3 proves that under realistic conditions, predictors based on invariant representations achieve strictly lower worst-case out-ofdistribution risk than those using raw features. The intuition is straightforward: raw features contain both causal and spurious parts. The spurious part may change arbitrarily in new environments, causing large errors in the worst-case scenario. In contrast, the invariant representation relies only on the stable causal mechanism, which remains unchanged across environments, thereby guaranteeing more reliable performance even under the most adverse distribution shifts.

4

Algorithm Implementation

Here we elaborate on the implementation of CmIR proposed in Theorem 2. Our objective is to learn the disentangled modality representations, and perform prediction by fusing invariant representations inv . The overall framework (see Figure 3) consists Zm of unimodal networks that produces raw unimodal features Xm ∈ R1×d (see Appendix B for unimodal networks), encoders gm , decoders rm , and a prediction head (predictor) h. Next, we detail the implementation of each loss component.

Distributionally Robust Risk Advantage of Invariant Representations

Having established how to learn invariant representations, we now prove their theoretical advantages for worst-case OOD risk under distribution shift. Theorem 3 (Distributionally Robust Risk Advantage of Invariant Representations). Let H be a hypothesis class over X × Y, where X = X1 × X2 × 4

Figure 3: The overall framework of CmIR and the visualization of the proposed constraints.

4.1

where α(e) is an environment-dependent coefficient controlling the noise intensity, and Σm is the modality-specific covariance matrix (which can be set as the identity matrix or estimated from the data). The noise coefficient α(e) is different for different environments, which is defined as α(e) = α(1) ∗ e, e ∈ {1, 2, ..., K}. α(1) is the noise coefficient for Environment 1, which is a hyperparameter whose values are shown in Table 7. The (e) perturbed feature X̃m is fed into the encoder gm inv,(e) to obtain the invariant representation Zm for that environment. Finally, the invariance constraint is implemented by minimizing the discrepancy beinv , E) tween the conditional distributions P (Y |Zm across different environments. To realize this, we can adopt a common strategy that enforces consistency in the output distributions of predictor across environments for classification tasks. Specifically, Kullback-Leibler (KL) divergence can be used as the distribution distance measure D:   X (m) inv,(e1 ) inv,(e2 ) Rinv = KL P (Y |Zm )∥P (Y |Zm )

Prediction Loss Lpred

The prediction loss ensures that invariant representations effectively predict the target label Y . Given a batch of data, the prediction loss is computed as: inv spu = gm (Xm ) Zm , Zm

(1)

N

  1 X inv M Lpred = }m=1 , Yi ℓ h {Zm,i N

(2)

i=1

where gm is the encoder for modality m, N is the batch size, and ℓ is the corresponding loss function. For classification tasks (e.g., humor detection and sarcasm detection), cross-entropy loss is utilized. For regression tasks (e.g., sentiment analysis), mean squared error (MSE) or mean absolute error (MAE) is used. The predictor is implemented using unimodal feature concatenation and a few multi-layer perception layers (see Figure 7). 4.2

(m)

The Invariance Constraint Rinv

The invariance constraint requires that inv , E) remains invariant across difP (Y |Zm ferent environments. However, real-world data often lack explicit environment labels E. To address this, we draw inspiration from data augmentation and simulate different virtual environments by injecting varying degrees of noise into the raw features Xm . For each sample, we assign a random virtual environment label e ∈ {1, 2, ..., K}, where K is the number of environments (a hyperparameter, see Table 7). Then we perform noise perturbation via applying (e) an additive noise ϵm = α(e) · ϵm to Xm based on the environment label e: (e) X̃m = Xm + α(e) · ϵm ,

e1 ̸=e2 inv,(e)

where P (Y |Zm ) is given by the output of predictor (after Softmax) on the corresponding environment’s representation. However, this strategy requires to implement a unimodal predictor for each invariant modality representation which increases the model complexity, and training noise might be introduced if unimodal predictors are not well trained. Moreover, it is hard for regression tasks to calculate the KL-divergence between output distributions. To this end, we adopt a simpler implementation that encourages the learned invariant representations to be identical across different environments, which is a stronger constraint that satisfies the invariance constraint because the out-

ϵm ∼ N (0, Σm ) (3) 5

where diag(C m ) denotes the diagonal matrix of C m , and α is a hyperparameter that is between zero and one. In Eq. 6, we use the Frobenius norm of matrix ∥ · ∥F (standard for matrix regularization). The term α balances the constraint strength between diagonal (same-sample) and offdiagonal (cross-sample) terms in the correlation matrix, which is a hyperparameter that depends on datasets (see Table 7). When α is less than 1, it can down-weight off-diagonal terms to focus on sample-wise orthogonality. In this way, we can enforce a stricter constraint on the invariant and spurious representations from the same sample, and also encourage invariant and spurious representations from different samples to be orthogonal, promoting the statistical independence between two representations.

put distributions must be the same if input features were the same. We have provided the comparison results of these two variants in Appendix H. Specifically, the invariance constraint is implemented as: X (m) inv,(e1 ) inv,(e2 ) Rinv = ∥Zm − Zm ∥1 (4) e1 ̸=e2

Minimizing this term encourages the model to extract features from Xm that are insensitive to noise perturbations (simulating environmental changes).In practice, for each sample in a batch, we can assign an environment label e and generate (e) X̃m using Eq. 3. For K environments, we can generate K + 1 variants of unimodal representations (including the original unimodal representation it(m) self). The invariance loss Rinv is computed over all K(K + 1)/2 pairs for each sample in the batch, ensuring strong invariance constraints. 4.3

4.4

(m)

Mutual Information Constraint Rdec

The reconstruction loss ensures that the disentaninv , Z spu ) retain all informagled representations (Zm m tion from the original input Xm , preventing the loss of crucial content during representation learning. Firstly, the encoder gm maps the input Xm to inv , Z spu ). the disentangled representation pair (Zm m The decoder rm then attempts to reconstruct the original input from this pair:

Theorem 2 requires minimizing the mutual inforinv ; Z spu ) between the invariant repremation I(Zm m inv and the spurious representation Z spu sentation Zm m to promote their disentanglement and capture independent information. Directly computing mutual information is intractable. We employ a widely used and effective alternative: approximating mutual information minimization by enforcing orthogonality (zero linear correlation) between the two representations in the feature space, which is a practical and computationally efficient proxy for miniinv and Z spu . mizing mutual information between Zm m Orthogonality is a necessary condition for statistical independence, and we augment this constraint with invariance and reconstruction constraints to ensure semantic separation of causal and spurious factors. This proxy is widely used in disentanglement learning for its scalability to large multimodal inv and datasets. Minimizing this term encourages Zm spu Zm to learn in orthogonal directions, thereby reducing information redundancy between them. Specifically, for each batch of training data, the correlation matrix C m can be calculated as: inv spu ⊤ C m = N or(Zm )N or(Zm )

(m)

Reconstruction Constraint Rrec

inv spu X̂m = rm (Zm , Zm )

(7)

The reconstruction loss is computed using MSE: (m)

Rrec = ∥Xm − X̂m ∥22

(8)

The encoder and decoder are implemented as multilayer perceptron networks (see Figure 7). 4.5

Overall Optimization Objective

The complete optimizable objective function is: L = Lpred +

M X

(m)

(m)

(m)

λ1 Rinv +λ2 Rdec + λ3 Rrec (9)

m=1

where λ1 , λ2 , λ3 are hyperparameters that balance the importance of each constraint.

(5)

5

where N or(x) = x−mean(x) denotes feature norstd(x)

Experiments

CmIR is evaluated on multiple tasks of MAC, including multimodal sentiment analysis (MSA), multimodal humor detection (MHD) and multimodal sarcasm detection (MSD). The used datasets include CMU-MOSI (Zadeh et al., 2016), CMUMOSI (OOD) (Sun et al., 2022), CMU-MOSEI

inv ∈ RN ×d denotes a batch of invarimalization, Zm ant modality representations, and C m ∈ RN ×N is the correlation matrix for modality m. Then, we enforce orthogonality via the following operation: (m)

Rdec = ∥diag(C m )+α · (C m−diag(C m ))∥2 (6) 6

(Zadeh et al., 2018b), CH-SIMS-v2 (Liu et al., 2022), UR-FUNNY (Hasan et al., 2019) and MUStARD (Castro et al., 2019). Due to space limitation, we introduce experimental settings, baselines, datasets and additional results in Appendix. Codes are available at: https://github. com/TmacMai/CmIR. 5.1

(a) Results on MHD task

(b) Results on MSD task

Figure 4: The results on (a) UR-FUNNY (Hasan et al., 2019) and (b) MUStARD (Castro et al., 2019) datasets.

Performance on the MSA Task

The performance of CmIR on MSA is summarized in Table 1 and Table 2. On CMU-MOSI, CmIR surpasses strong baseline ITHP (Xiao et al., 2024) by more than 2 points in Acc7 and 1 points in Acc2. For CMU-MOSEI, it outperforms GSCon (Shi et al., 2025) and achieves the best scores in Acc2, F1, MAE, and Acc7.Compared with feature disentanglement method FDMER (Yang et al., 2022) and previous backdoor-adjustment work that focuses on specific confounders (Xu et al., 2025b), CmIR demonstrates considerable improvement. Similar superiority is observed on CH-SIMS-v2 (Table 2), where CmIR outperforms all baselines across every metric, including an improvement of 2.5 points in Acc5. Overall, CmIR establishes state-of-the-art results on MSA across three standard benchmarks. This strong performance is primarily attributed to CmIR’s causal learning strategy, which effectively learns invariant representations across all environments that eliminate general bias and enable a more robust multimodal learning.

The results under OOD scenarios is depicted in Table 3. It is observed that: I) All models degrade when moving from in-distribution to OOD settings, verifying that spurious correlations impede generalization; II) CmIR delivers significantly stronger OOD performance than standard multimodal baselines. Its advantage over ITHP (Xiao et al., 2024) grows notably, with Acc2 improvement rising from 1.5 points to 3.5 points, and Acc7 improvement increasing from 2.1 points to 7.2 points, underscoring the efficacy of our causal strategy; III) Compared to recent causality-based methods (CLUE, GEAR, MulDeF), CmIR consistently outperforms them in all metrics, highlighting its robustness in mitigating broad spurious correlations. This is because instead of focusing on specific bias or assumption, CmIR directly learns invariant representations that are stable across all environments, which is more general and applicable to various distribution shifts.

5.2

5.4

5.3

Performance on the MHD and MSD Tasks

Results under OOD scenarios.

Discussion on Noisy Modalities

(1) To assess the robustness of CmIR to modality noise, we corrupt all modalities of all training and testing samples with Gaussian noise (the noise rate NR is set at 10% -70%). The compared baselines include TMDC (Zhuang et al., 2025), C-MIB (Mai et al., 2023c) and Multimodal Boosting (Mai et al., 2024), which adopt the same training and testing settings as CmIR. Following prior work (Mai et al., 2024), we report Acc2 and MAE. Table 4 shows that CmIR outperforms competitive baselines across most metrics (particularly in MAE), and its performance advantage becomes even more pronounced as the noise level increases. This is mainly because CmIR can more accurately identify and extract causal features from noisy inputs and maintain stable predictive ability for labels via the proposed constraints. These results indicate the robustness of CmIR in handling noise data. (2) To assess the resilience of CmIR to out-

To assess the task generalizability of CmIR, we evaluate it on MHD and MSD (classification tasks) using UR-FUNNY and MUStARD datasets. The baselines include MulT (Tsai et al., 2019a), SelfMM (Yu et al., 2021), MMIM (Han et al., 2021), HKT (Hasan et al., 2021), DMD (Li et al., 2023), DMD+SuCI (Xu et al., 2025b), AtCAF (Huang et al., 2025), MAG-XLNet (Rahman et al., 2020), MCL (Mai et al., 2023a), MGCL (Mai et al., 2023b), ITHP (Xiao et al., 2024), MISA (Hazarika et al., 2020), and MIL (Zhang et al., 2024), where DMD+SuCI and AtCAF are causality-based methods. As shown in Figure 4, CmIR surpasses the strongest baselines (AtCAF and MGCL) by margins exceeding 4 and 2 points on UR-FUNNY and MUStARD, respectively. Overall, CmIR achieves competitive performance on both MHD and MSD, confirming its effectiveness and strong generalizability to diverse multimodal tasks. 7

Table 1: Comparisons on the CMU-MOSI and CMU-MOSEI datasets. The results labeled with † are obtained from original papers, and other results are obtained from our experiments. The best results are highlighted. Model

Venue

Self-MM (Yu et al., 2021) ConFEDE (Yang et al., 2023) FDMER† (Yang et al., 2022) SuCI† (Xu et al., 2025b) C-MIB (Mai et al., 2023c) EMOE (Fang et al., 2025) Multimodal Boosting (Mai et al., 2024) ITHP (Xiao et al., 2024) Diffusion Bridge (Lee et al., 2025) GSCon (Shi et al., 2025) CmIR

AAAI 21 ACL 23 ACM MM 22 AAAI 25 TMM 23 CVPR 25 TMM 24 ICLR 24 CVPR 25 TIP 25 -

CMU-MOSI Acc2↑ F1↑ MAE↓ 84.9 84.8 0.731 85.1 85.2 0.728 84.6 84.7 0.724 84.6 84.5 87.8 87.8 0.662 84.8 84.8 0.723 88.5 88.4 0.634 88.5 88.5 0.663 86.9 86.8 0.649 88.1 88.0 0.696 89.6 89.5 0.616

Acc7↑ 45.8 43.3 44.1 42.2 47.7 45.2 49.1 47.7 47.3 45.7 49.8

Model MISA (Hazarika et al., 2020) MAG-BERT (Rahman et al., 2020) Self-MM (Yu et al., 2021) MMIM (Han et al., 2021) AV-MC (Liu et al., 2022) KuDA (Feng et al., 2024) DLF (Wang et al., 2025) Diffusion Bridge (Lee et al., 2025) CmIR

Acc5↑ 47.5 49.2 53.5 50.5 52.1 53.1 47.5 52.5 56.0

Acc3↑ 68.9 70.6 72.7 70.4 73.2 74.3 70.0 70.7 75.0

MAE↓ 0.342 0.346 0.315 0.339 0.301 0.289 0.346 0.323 0.286

Corr↑ 0.671 0.641 0.691 0.641 0.721 0.741 0.683 0.677 0.742

Table 3: Results on CMU-MOSI (OOD). Following causal-based methods (Sun et al., 2022, 2023), Acc2† and F1† denote results considering neutral samples. Methods CLUE (Sun et al., 2022) GEAR (Sun et al., 2023) MulDeF (Huan et al., 2024) Self-MM (Yu et al., 2021) ITHP (Xiao et al., 2024) KAN-MCP (Luo et al., 2025a) CmIR

Acc7 Acc2† Acc2

F1†

F1

41.8 42.9 40.3 41.3 41.8 48.5

78.8 80.4 79.7 76.7 79.5 79.1 83.0

79.9 82.1 81.6 78.1 81.3 81.4 84.4

78.8 80.5 79.6 76.7 79.5 79.3 83.0

79.9 82.1 81.5 78.1 81.3 81.5 84.4

of-distribution (OOD) noises not encountered in training, we adopt a mixed-noise evaluation strategy: training samples are contaminated with Gaussian noise, while testing samples are perturbed with distinct noise types (Laplace and random erasing). As presented in Table 4, ‘CmIR (OOD)’ obtains competitive results and significantly surpasses strong baselines. These findings confirm that CmIR generalizes effectively to unseen noises, highlighting its promising application potential. 5.5

Acc7↑ 53.0 52.7 54.1 54.6 52.7 52.5 54.0 52.2 53.1 50.8 55.1

CMU-MOSEI Acc2↑ F1↑ MAE↓ 85.2 85.2 0.540 85.7 85.6 0.538 86.1 85.8 0.536 85.8 85.7 86.9 86.8 0.542 85.0 85.0 0.542 86.5 86.5 0.523 87.1 87.1 0.550 87.1 87.0 0.531 87.4 87.4 0.561 87.8 87.7 0.513

Corr↑ 0.763 0.772 0.773 0.784 0.760 0.779 0.792 0.800 0.750 0.793

because it is the core constraint to learn causal invariant representations. Without it, we cannot eninv is environment-invariant sure that the learned Zm and the core idea of CmIR cannot be realized; (3) (m) Mutual Information Constraint: When Rdec ’ is removed, the extent of performance decline is (m) similar to that observed when Rinv is removed, indicating the necessity of minimizing the mutual inv and Z spu , and demoninformation between Zm m strating the effectiveness of our disentanglement framework. Compared with learning only invariant representations (Song et al., 2025), CmIR simultaneously learns both invariant and spurious representations while minimizing their mutual information, enabling the model to understand and learn the properties of invariant representations more easily and comprehensively; (4) Reconstruction Con(m) straint: When Rrec is removed, the performance also shows a noticeable decline, although the drop is the smallest among all ablations. This occurs because the reconstruction loss ensures the causal and spurious representations fully retain all information from the raw features, thereby preventing information loss and suboptimal disentanglement.

Table 2: The results on CH-SIMS-v2. CH-SIMS v2 Acc2↑ F1↑ 78.2 78.3 77.1 77.1 78.7 78.6 77.8 77.8 80.6 80.7 80.2 80.1 78.1 77.9 78.6 79.0 81.9 82.0

Corr↑ 0.785 0.784 0.788 0.835 0.790 0.855 0.856 0.839 0.832 0.853

5.6

Hyperparameter Robustness Analysis

We evaluate the effectiveness of hyperparameters on CMU-MOSI (OOD), including the weights for invariance constraint λ1 , mutual information constraint λ2 , reconstruction constraint λ3 , and the number of environments K. As shown in Figure 5 (a), (b), and (c), when the values of weights are small, the performance of CmIR experiences a certain degree of degradation, as the effect of the constraints is not fully utilized. Among them, the performance drop is most pronounced when the weight of invariance constraint decreases, highlighting its importance. Conversely, when the values of λ2 and λ3 are too large, the performance also declines, likely because mutual information

Ablation Experiments

(1) Causal Inference: As presented in Table 5, in the case of ‘Vanilla Framework’, we directly use the raw modality features for prediction. The model exhibits its sharpest performance decline (over 3 points in Acc2 and Acc7), suggesting the importance of learning causal invariant representations for more robust multimodal prediction and verifying our claim; (2) Invariance Constraint: (m) As shown in ‘W/O Rinv ’, removing invariance constraint leads to a noticeable performance drop, 8

Table 4: Discussion on noisy modalities on CMU-MOSI and CMU-MOSEI. MuBo denotes Multimodal Boosting. NR

MOSI

MOSEI

0.1 0.2 0.3 0.4 0.5 0.6 0.7 Avg 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Avg

TMDC Acc2/MAE 87.4 / 0.748 86.6 / 0.741 86.9 / 0.733 85.5 / 0.792 85.0 / 0.912 84.8 / 1.181 84.0 / 1.277 85.7 / 0.912 86.6 / 0.618 86.2 / 0.593 85.3 / 0.603 84.4 / 0.596 84.6 / 0.592 83.8 / 0.602 83.4 / 0.623 84.9 / 0.604

Gaussian Noise C-MIB MuBo Acc2/MAE Acc2/MAE 87.8 / 0.670 86.7 / 0.678 87.5 / 0.726 86.1 / 0.738 86.4 / 0.912 86.4 / 0.785 83.2 / 1.366 85.5 / 0.841 84.9 / 1.660 86.1 / 1.172 80.8 / 2.595 82.0 / 1.355 82.1 / 3.146 84.4 / 1.750 84.7 / 1.582 85.3 / 1.046 86.1 / 0.545 86.4 / 0.544 84.5 / 0.582 86.6 / 0.557 85.6 / 0.622 85.5 / 0.623 84.4 / 0.703 85.3 / 0.682 83.7 / 0.875 84.1 / 0.724 82.4 / 1.054 85.4 / 0.924 80.5 / 1.404 80.3 / 1.125 83.9 / 0.826 84.8 / 0.740

(a) Weight λ1

CmIR Acc2/MAE 88.1 / 0.615 87.4 / 0.621 87.3 / 0.650 86.4 / 0.669 86.9 / 0.685 87.1 / 0.680 86.1 / 0.740 87.0 / 0.666 87.5 / 0.528 87.2 / 0.522 86.8 / 0.531 86.6 / 0.535 85.9 / 0.558 85.8 / 0.548 86.3 / 0.671 86.6 / 0.556

(b) Weight λ2

OOD Noises (Laplace Noise and Random Erasing Noise) TMDC C-MIB MuBo CmIR Acc2/MAE Acc2/MAE Acc2/MAE Acc2/MAE 87.2 / 0.769 87.8 / 0.666 87.4 / 0.639 87.8 / 0.638 86.7 / 0.861 87.5 / 0.689 87.3 / 0.681 88.2 / 0.644 86.8 / 0.741 85.1 / 1.019 87.1 / 0.710 86.9 / 0.678 85.2 / 0.772 85.6 / 1.303 85.3 / 0.900 85.8 / 0.663 84.6 / 0.867 83.9 / 1.900 86.9 / 1.114 85.6 / 0.715 83.1 / 0.884 86.5 / 2.272 85.0 / 1.060 87.0 / 0.689 82.0 / 0.966 84.0 / 3.516 84.7 / 1.731 85.6 / 0.771 85.1 / 0.837 85.8 / 1.624 86.2 / 0.976 86.7 / 0.685 86.4 / 0.613 86.9 / 0.564 86.1 / 0.561 87.1 / 0.523 85.5 / 0.631 85.1 / 0.594 85.2 / 0.585 86.6 / 0.523 85.0 / 0.667 85.9 / 0.665 85.4 / 0.610 86.5 / 0.528 84.9 / 0.682 84.8 / 0.767 84.4 / 0.714 86.7 / 0.535 84.3 / 0.751 73.1 / 0.768 84.4 / 0.763 86.3 / 0.563 82.7 / 0.829 82.0 / 0.856 84.1 / 0.795 86.3 / 0.644 81.7 / 0.888 86.5 / 2.366 85.0 / 1.158 85.4 / 0.663 84.4 / 0.723 83.5 / 0.940 84.9 / 0.741 86.4 / 0.568

(c) Weight λ3

(d) Environment Number K

Figure 5: Acc2 of CmIR w.r.t the change of constraint weights and the number of environments. Table 5: Ablation experiments on CMU-MOSI. Model Vanilla Framework

46.3

85.8

0.681

(m) W/O Rinv (m) W/O Rdec (m) W/O Rrec

45.7

86.7

0.652

45.7

86.9

0.671

47.7

88.1

0.623

CmIR

49.8

89.6

0.616

(a) Regular Features

5.7 Visualization of Invariant Representations

Acc7↑ Acc2↑ MAE↓

To demonstrate that CmIR indeed learns invariant representations that maintain stable predictive power for labels, we intervene with OOD noise on test samples and visualize the extracted causal language representations, alongside visualizing the language representations obtained without causal inference training. As shown in Figure 6, the causal representations learned by CmIR for different classes are well-separated in the feature space, with neutral samples concentrated between the positive and negative ones. In contrast, regular representations learned without CmIR exhibit substantial overlap, and neutral samples appear more scattered. This indicates that even under noisy conditions, the causal representations learned by CmIR more accurately reflect label information.

(b) Invariant Features

Figure 6: T-SNE visualization of language features without/with causal inference learning.

6

and reconstruction constraints dominate the learning process, preventing sufficient attention to the invariance constraint and the prediction loss. Moreover, as shown in Figure 5 (d), the performance remains relatively stable as the value of K varies, indicating the robustness of CmIR. Overall, when hyperparameters are varied across a wide range, CmIR’s performance consistently maintains a good level (Acc2 ≥ 81%), which to some extent demonstrates the stability of CmIR.

Conclusion

We propose CmIR for robust multimodal learning. By disentangling each modality into invariant causal and spurious components, CmIR learns stable and causally-aware representations. We provide theoretical guarantees for CmIR, and demonstrate state-of-the-art results in multiple tasks. CmIR excels under distribution shifts and noisy-modality conditions, highlighting its practical robustness. 9

Limitations

J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT.

While CmIR demonstrates strong performance and robustness, our work has certain limitations. First, the environmental simulation via feature perturbation, while effective, may not fully capture the complexity of real-world distribution shifts. Future work could explore more sophisticated environment generation strategies or incorporate realworld multi-environment datasets. Second, the mutual information minimization constraint implemented via feature orthogonality is an approximation. More precise mutual information estimation techniques could be integrated for potentially better disentanglement, albeit at increased computational cost. These limitations, however, do not undermine the core theoretical contributions or the empirical effectiveness of the proposed framework.

7

Florian Eyben. 2010. Opensmile: the munich versatile and fast open-source audio feature extractor. In ACM International Conference on Multimedia, pages 1459– 1462. Yiyang Fang, Wenke Huang, Guancheng Wan, Kehua Su, and Mang Ye. 2025. Emoe: Modality-specific enhanced dynamic emotion experts. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 14314–14324. Xinyu Feng, Yuming Lin, Lihua He, You Li, Liang Chang, and Ya Zhou. 2024. Knowledge-guided dynamic modality attention fusion framework for multimodal sentiment analysis. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 14755–14766. Wei Han, Hui Chen, and Soujanya Poria. 2021. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9180–9192, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.

Acknowledgment

This work is supported by the Guangdong Philosophy and Social Sciences Planning Project (No. GD26YJY34).

References

Md Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, and Ehsan Hoque. 2021. Humor knowledge enriched transformer for understanding multimodal humor. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 12972–12980.

Emile HL Aarts and 1 others. 1987. Simulated annealing: Theory and applications. Reidel. Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019. Invariant risk minimization. In International Conference on Learning Representations.

Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, LouisPhilippe Morency, and Mohammed Ehsan Hoque. 2019. Ur-funny: A multimodal language dataset for understanding humor. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2046–2056.

Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency. 2018. Openface 2.0: Facial behavior analysis toolkit. In 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018), pages 59–66. IEEE. Santiago Castro, Devamanyu Hazarika, Verónica PérezRosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019. Towards multimodal sarcasm detection (an _obviously_ perfect paper). In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4619–4629.

Devamanyu Hazarika, R. Zimmermann, and Soujanya Poria. 2020. Misa: Modality-invariant and -specific representations for multimodal sentiment analysis. ACM MM, pages 1122–1131.

Yong Dai, Jian Liu, Xiancong Ren, and Zenglin Xu. 2020. Adversarial training based multi-source unsupervised domain adaptation for sentiment analysis. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7618–7625.

Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021.

Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer. 2014. Covarep: A collaborative voice analysis repository for speech technologies. In ICASSP, pages 960–964.

Ruohong Huan, Guowei Zhong, Peng Chen, and Ronghua Liang. 2024. Muldef: A model-agnostic debiasing framework for robust multimodal sentiment analysis. IEEE Transactions on Multimedia.

10

Changqin Huang, Jili Chen, Qionghao Huang, Shijin Wang, Yaxin Tu, and Xiaodi Huang. 2025. Atcaf: Attention-based causality-aware fusion network for multimodal sentiment analysis. Information Fusion, 114:102725.

Sijie Mai, Ya Sun, Ying Zeng, and Haifeng Hu. 2023a. Excavating multimodal correlation for representation learning. Information Fusion, 91:542–555. Sijie Mai, Ying Zeng, and Haifeng Hu. 2023b. Learning from the global view: Supervised contrastive learning of multimodal representation. Information Fusion, 100:101920.

Menghua Jiang, Yuxia Lin, Baoliang Chen, Haifeng Hu, Yuncheng Jiang, and Sijie Mai. 2025. Disentangling bias by modeling intra-and inter-modal causal attention for multimodal sentiment analysis. arXiv preprint arXiv:2508.04999.

Sijie Mai, Ying Zeng, and Haifeng Hu. 2023c. Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations. IEEE Transactions on Multimedia, 25:4121–4134.

Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations.

Sijie Mai, Ying Zeng, and Haifeng Hu. 2025. Learning by comparing: Boosting multimodal affective computing through ordinal learning. In Proceedings of the ACM on Web Conference 2025 (WWW ’25), pages 2120–2134, New York, NY, USA. Association for Computing Machinery.

Jeong Ryong Lee, Yejee Shin, Geonhui Son, and Dosik Hwang. 2025. Diffusion bridge: Leveraging diffusion model to reduce the modality gap between text and vision for zero-shot image captioning. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 4050–4059.

Sijie Mai, Ying Zeng, Shuangjia Zheng, and Haifeng Hu. 2023d. Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis. IEEE Transactions on Affective Computing, 14(3):2276– 2289.

Yong Li, Yuanzhi Wang, and Zhen Cui. 2023. Decoupled multimodal distilling for emotion recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6631– 6640.

Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. 2020. Learning from failure: Debiasing classifier from biased classifier. Advances in Neural Information Processing Systems, 33:20673– 20684.

Yang Liu, Guanbin Li, and Liang Lin. 2023. Crossmodal causal relational reasoning for event-level visual question answering. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):11624–11641.

Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021. Counterfactual vqa: A cause-effect look at language bias. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12700– 12710.

Yihe Liu, Ziqi Yuan, Huisheng Mao, Zhiyun Liang, Wanqiuyue Yang, Yuanzhe Qiu, Tie Cheng, Xiaoteng Li, Hua Xu, and Kai Gao. 2022. Make acoustic and visual cues matter: Ch-sims v2. 0 dataset and avmixup consistent module. In Proceedings of the 2022 international conference on multimodal interaction, pages 247–258.

Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. 2016. Causal inference by using invariant prediction: identification and confidence intervals. Journal of the Royal Statistical Society Series B: Statistical Methodology, 78(5):947–1012. Soujanya Poria, Erik Cambria, Rajiv Bajpai, and Amir Hussain. 2017. A review of affective computing: From unimodal analysis to multimodal fusion. Information Fusion, 37:98–125.

Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations. Miaosen Luo, Yuncheng Jiang, and Sijie Mai. 2025a. Towards explainable fusion and balanced learning in multimodal sentiment analysis. In ACM MM, pages 1997–2006.

Wasifur Rahman, M. Hasan, Sangwu Lee, Amir Zadeh, Chengfeng Mao, Louis-Philippe Morency, and E. Hoque. 2020. Integrating multimodal information in large pretrained transformers. ACL, 2020:2359–2369.

Yuanyi Luo, Wei Liu, Qiang Sun, Sirui Li, Jichunyang Li, Rui Wu, and Xianglong Tang. 2025b. Triagedmsa: Triaging sentimental disagreement in multimodal sentiment analysis. IEEE Transactions on Affective Computing.

Shiv Shankar. 2022. Multimodal fusion via cortical network inspired losses. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1167– 1178.

Sijie Mai, Ya Sun, Aolin Xiong, Ying Zeng, and Haifeng Hu. 2024. Multimodal boosting: Addressing noisy modalities and identifying modality contribution. IEEE Transactions on Multimedia, 26:3018–3033.

QingHongYa Shi, Mang Ye, Wenke Huang, Bo Du, and Xiaofen Zong. 2025. Gradient and structure consistency in multimodal emotion recognition. IEEE Transactions on Image Processing.

11

Zeen Song, Siyu Zhao, Xingyu Zhang, Jiangmeng Li, Changwen Zheng, and Wenwen Qiang. 2025. Learning invariant causal mechanism from vision-language models. In Forty-second International Conference on Machine Learning.

Dingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du, and Lihua Zhang. 2022. Disentangled representation learning for multimodal emotion recognition. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1642–1651.

Teng Sun, Juntong Ni, Wenjie Wang, Liqiang Jing, Yinwei Wei, and Liqiang Nie. 2023. General debiasing for multimodal sentiment analysis. In Proceedings of the 31st ACM International Conference on Multimedia, pages 5861–5869.

Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu, Kun Yang, Zhaoyu Chen, Yuzheng Wang, Peng Zhai, Ke Li, and Lihua Zhang. 2024. Towards multimodal sentiment analysis debiasing via bias purification. In European Conference on Computer Vision, pages 464–481. Springer.

Teng Sun, Wenjie Wang, Liqaing Jing, Yiran Cui, Xuemeng Song, and Liqiang Nie. 2022. Counterfactual reasoning for out-of-distribution multimodal sentiment analysis. In Proceedings of the 30th ACM International Conference on Multimedia, pages 15–23.

Jiuding Yang, Yakun Yu, Di Niu, Weidong Guo, and Yu Xu. 2023. Confede: Contrastive feature decomposition for multimodal sentiment analysis. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7617–7630.

Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019a. Multimodal transformer for unaligned multimodal language sequences. In ACL, pages 6558–6569.

Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu. 2021. Learning modality-specific representations with selfsupervised multi-task learning for multimodal sentiment analysis. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pages 10790–10797.

Yao Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis Philippe Morency, and Ruslan Salakhutdinov. 2019b. Learning factorized multimodal representations. In ICLR.

Ziqi Yuan, Yihe Liu, Hua Xu, and Kai Gao. 2024. Noise imitation based adversarial training for robust multimodal sentiment analysis. IEEE Transactions on Multimedia, 26:529–539.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.

Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In EMNLP, pages 1114–1125.

Pan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen, and Jingtong Hu. 2025. Dlf: Disentangled-languagefocused multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 21180–21188.

Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis Philippe Morency. 2018a. Memory fusion network for multiview sequential learning. In AAAI, pages 5634–5641.

Wenjie Wang, Xinyu Lin, Fuli Feng, Xiangnan He, Min Lin, and Tat-Seng Chua. 2022. Causal representation learning for out-of-distribution recommendation. In Proceedings of the ACM Web Conference 2022, pages 3562–3571.

Amir Zadeh, Paul Pu Liang, Jonathan Vanbriesen, Soujanya Poria, Edmund Tong, Erik Cambria, Minghai Chen, and Louis Philippe Morency. 2018b. Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. In ACL, pages 2236–2246.

Xiongye Xiao, Gengshuo Liu, Gaurav Gupta, Defu Cao, Shixuan Li, Yaxing Li, Tianqing Fang, Mingxi Cheng, and Paul Bogdan. 2024. Neuro-inspired information-theoretic hierarchical perception for multimodal learning. In The Twelfth International Conference on Learning Representations.

Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis Philippe Morency. 2016. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intelligent Systems, 31(6):82– 88.

Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, and 1 others. 2025a. Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215.

Yazhou Zhang, Yang Yu, Dongming Zhao, Zuhe Li, Bo Wang, Yuexian Hou, Prayag Tiwari, and Jing Qin. 2024. Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation. IEEE Transactions on Artificial Intelligence, 5(3):1349–1361.

Zhi Xu, Dingkang Yang, Mingcheng Li, Yuzheng Wang, Zhaoyu Chen, Jiawei Chen, Jinjie Wei, and Lihua Zhang. 2025b. Debiased multimodal understanding for human language sequences. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 14450–14458.

Yuhao Zhang, Ying Zhang, Wenya Guo, Xiangrui Cai, and Xiaojie Yuan. 2022. Learning disentangled representation for multimodal cross-domain sentiment analysis. IEEE transactions on neural networks and learning systems, 34(10):7956–7966.

12

Step 2 (Sufficiency of the Optimization Objective): The optimization objective directly minimizes the maximum discrepancy of P (Y |ϕm (Xm ), E) across environments. By the properties of the distribution distance measure D, D(P1 , P2 ) = 0 if and only if P1 = P2 . Therefore, the optimal solution ϕ∗m must satisfy:

Jiaxing Zhao, Qize Yang, Yixing Peng, Detao Bai, Shimin Yao, Boyuan Sun, Xiang Chen, Shenghao Fu, Xihan Wei, Liefeng Bo, and 1 others. 2025. Humanomni: A large vision-speech language model for human-centric video understanding. arXiv preprint arXiv:2501.15111. Aoqiang Zhu, Min Hu, Xiaohua Wang, Jiaoyun Yang, Yiming Tang, and Ning An. 2025. Multimodal invariant sentiment representation learning. In Findings of the Association for Computational Linguistics: ACL 2025, pages 14743–14755.

D(P (Y |ϕ∗m (Xm ), E = e1 ), P (Y |ϕ∗m (Xm ), E = e2 )) = 0

Yan Zhuang, Minhao Liu, Yanru Zhang, Jiawen Deng, and Fuji Ren. 2025. Tmdc: A two-stage modality denoising and complementation framework for multimodal sentiment analysis with missing and noisy modalities. arXiv preprint arXiv:2511.10325.

A

for all environment pairs e1 , e2 ∈ E, which implies: P (Y |ϕ∗m (Xm ), E = e1 ) = P (Y |ϕ∗m (Xm ), E = e2 ) Step 3 (Causal Interpretation): The condition P (Y |ϕ∗m (Xm ), E = e1 ) = P (Y |ϕ∗m (Xm ), E = e2 ) for all e1 , e2 implies that ϕ∗m (Xm ) blocks all backdoor paths from E to Y that pass through modality m. By the backdoor criterion (Peters et al., 2016), this means ϕ∗m (Xm ) contains only inv and excludes spurious causal features from Zm spu features from Zm , as the latter would create environment-dependent associations with Y . This completes the proof that ϕ∗m (Xm ) is a valid causal invariant modality representation capturing only causal features.

Detailed Theoretical Analysis

Here we establish a detailed theoretical foundation for our causal approach to multimodal learning. We first establish the existence/definition and extractability of causal invariant modality representations, then prove their advantages in terms of generalization performance. A.1

Definition of Invariant Representations

Theorem 1 (Definition of Causal Invariant Modality Representations) Assume there exists a function inv } and a distribution class Φm = {ϕm : Xm → Zm distance measure D (e.g., KL divergence or Wasserstein distance) such that the following optimization problem has a solution: ϕ∗m = arg min

A.2

max D(P (Y |ϕm (Xm ), E = e1 ),

ϕm ∈Φm e1 ,e2 ∈E

P (Y |ϕm (Xm ), E = e2 )) Then ϕ∗m (Xm ) constitutes a causal invariant modality representation satisfying: P (Y |ϕ∗m (Xm ), E = e1 ) =P (Y |ϕ∗m (Xm ), E = e2 ), ∀e1 , e2 ∈ E

Extractability of Invariant Representations

While Theorem 1 establishes the definition of invariant representations, practical implementation requires extracting these representations from raw modalities while preserving all relevant information. This motivates our disentanglement approach. Theorem 2 (Theoretical Guarantee for Disentangled Representations) Consider encoder functions inv , Z spu ) and decoder functions gm : Xm → (Zm m inv , Z spu ) → X that optimize the followrm : (Zm m m ing objective: min

{gm ,rm ,h}M m=1

Proof. The proof follows from the invariance principle in causal inference (Peters et al., 2016). We proceed in three steps: Step 1 (Necessity of Invariance): If a feature Z contains causal information about Y , then the conditional distribution P (Y |Z) should remain invariant across different environments (Arjovsky et al., 2019). This is because causal mechanisms are stable under interventions on non-descendant variables in the causal graph.

+ λ1

M X m=1

inv M Ee∈E [Lpred (Y, h({Zm }m=1 ))]

(m) Rinv + λ2

M X m=1

(m) Rdec + λ3

M X

(m)

Rrec

m=1

where: • Lpred is the prediction loss and h is the prediction head P P (m) inv • Rinv = e1 ∈E e2 ∈E D(P (Y |Zm , E = inv , E = e )) enforces invariance e1 ), P (Y |Zm 2 13

(m)

spu

inv ; Z • Rdec = I(Zm m ) minimizes mutual information between invariant and spurious components (m)

Part 2 (Environment Independence of Invari(m) ant Component): The term λ1 Rinv with λ1 → ∞ inv , E = e ), P (Y |Z inv , E = forces D(P (Y |Zm 1 m e2 )) = 0 for all e1 , e2 . By Theorem 1, this iminv , E = e ) = P (Y |Z inv , E = e ) plies P (Y |Zm 1 2 m for all e1 , e2 . inv is not Now, assume for contradiction that Zm independent of E. Then there exist values z inv , e1 , inv = z inv |E = e ) ̸= P (Z inv = e2 such that P (Zm 1 m z inv |E = e2 ). By the law of total probability: Z inv P (Y |E = ei ) = P (Y |Zm = z inv )

spu

inv , Z 2 • Rrec = ∥Xm − rm (Zm m )∥ ensures reconstruction capability

• λ1 , λ2 , λ3 > 0 are hyperparameters Assuming the label Y is independent of environment E, the function classes {gm , rm , h} have sufficient capacity and the data follows the SCM described in Section 3.1, then as λ1 , λ2 , λ3 → ∞, the optimal solution satisfies:   inv spu 2 1. lim E ∥Xm −rm (Zm , Zm )∥ = 0 (per-

inv × P (Zm = z inv |E = ei )dz inv

λ3 →∞

inv ) = P (Y |Z inv ) but P (Z inv |E = Since P (Y |Zm m m inv |E = e ), we must have P (Y |E = e1 ) ̸= P (Zm 2 e1 ) ̸= P (Y |E = e2 ). However, in our SCM, Y is causally independent of E given the invariant inv }, and consequently given {Z inv }. features {Xm m inv ⊥ This contradiction implies Zm ⊥ E. Part 3 (No additional causal information in spu spu inv , E) = 0 by Zm ): We prove I(Y ; Zm | Zm contradiction, using the results from Part 1 and Part 2, the mutual information constraint, and the construction of virtual environments. Assume, for contradiction, that

fect reconstruction is achieved) inv 2. Zm ⊥ ⊥ E (invariant component is environment-independent) spu

inv , E) = 0 (spurious component 3. I(Y ; Zm |Zm contains no additional causal information)

Proof. We prove the three claims in order. The limit λi → ∞ is understood in the sense of tightening constraints: as λi grows, the corresponding regularization term must vanish to keep the loss finite, provided the optimal loss remains bounded. We assume the feasible set (where all terms are finite) is non-empty, which is reasonable given sufficient model capacity. Part 1 (Perfect Reconstruction): The term (m) λ3 Rrec with λ3 → ∞ forces the reconstruction error to zero. By of the squared L2  the properties  inv spu 2 norm, lim E ∥Xm − rm (Zm , Zm )∥ = 0 if λ3 →∞

spu inv I := I(Y ; Zm | Zm , E) > 0

From Part 2, the invariance constraint gives inv , E) = P (Y | Z inv ); hence P (Y | Zm m inv inv spu I = H(Y | Zm ) − H(Y | Zm , Zm , E) (1)

The inequality I > 0 implies

spu

inv , Z and only if Xm = rm (Zm m ) almost surely. Thus, in the limit, the decoder can reconstruct the input with arbitrarily small error. Note that exact zero error may be unattainable for finitedimensional representations, but the limiting statement suffices for theoretical analysis; in practice, taking λ3 sufficiently large yields negligible reconstruction error. This ensures the disentangled repinv , Z spu ) preserve all information resentations (Zm m in the original modality Xm . (m) The reconstruction constraint Rrec is crucial for preventing degenerate solutions. It ensures that the disentangled representations form a sufficient statistic for Xm , preserving all information while separating causal from non-causal components. This is essential for maintaining performance in the source environments while improving generalization to new environments

inv spu inv H(Y | Zm , Zm , E) < H(Y | Zm ).

(2)

spu Thus, conditioning on (Zm , E) strictly reduces inv the entropy of Y compared to conditioning on Zm alone. In particular, there exists a measurable set of positive measure on which the conditional disinv , Z spu ) depends on Z spu in a tribution P (Y | Zm m m non-degenerate way. By Part 1 (perfect reconstruction in the limit), inv , Z spu ) determines X almost surely. the pair (Zm m m Since Y is a function of the multimodal input, we can write:  inv spu Y = Ψ Zm , Zm ,η

where η is an independent noise term capturing spu irreducible uncertainty. The dependence on Zm is essential because of (2). 14

Now consider two different environments e1 and e2 , because the encoder gm is deterministic and the same for all environments, the conditional distrispu inv and E changes with e. bution of Zm given Zm spu inv , E = e) is Formally, the map e 7→ P (Zm | Zm injective; i.e., for e1 ̸= e2 we have

Theorem 3 (Distributionally Robust Risk Advantage of Invariant Representations) Let H be a hypothesis class over X × Y, where X = X1 × X2 × · · · × XM is the M -modal feature space and Y is the label space. Let Eall denote the set of all possible environments, each corresponding to a distribution P e (x, y). Let hinv ∈ H be a predictor spu inv spu inv inv }M , P (Zm | Zm , E = e1 ) ̸= P (Zm | Zm , E = e2 ) using invariant representations Z inv = {Zm m=1 and hraw ∈ H be a predictor using raw multimodal in the sense that the two conditional distributions representations X = {Xm }M m=1 . Assume: are not equal almost everywhere. 1. Invariance Condition: The invariant repreUsing the law of total probability, for any envi- sentations satisfy P e (Y |Z inv ) = P e′ (Y |Z inv ) for ronment e, all e, e′ ∈ Eall . 2. Information Sufficiency: The mutual inforinv , E = e) = P (Y | Zm Z mation between invariant representations and raw (3) inv spu spu inv features satisfies I(Z inv ; X) > c for some constant P (Y | Zm , Zm ) dP (Zm | Zm , E = e) c > 0 (Z inv contains enough information related to Because the environments can vary significantly, X). spu 3. Loss Function Regularity: The loss function the family of mixing distributions {P (Zm | inv ℓ is L-Lipschitz continuous and bounded. Zm , E = e)}e∈E is rich enough to distinguish Then the worst-case out-of-distribution risk satdifferent integrands. In particular, if the integrand spu spu inv isfies: P (Y | Zm , Zm ) is not constant in Zm (which ROOD (hinv ) < ROOD (hraw ) follows from I > 0), then the value of the integral in (3) changes continuously with e when α(e) where ROOD (h) = maxe∈E Re (h) and Re (h) = test varies. Hence for e1 ̸= e2 , the two integrals cannot E(x,y)∼P e [ℓ(h(x), y)]. be equal. Consequently, Proof. We prove the theorem through four key inv inv steps, with additional explanations regarding the P (Y | Zm , E = e1 ) ̸= P (Y | Zm , E = e2 ) validity of assumptions and practical implications. which contradicts Part 2 where we established that Step 1 (Information Retention Property): inv , E = e) is constant across all environP (Y | Zm The information sufficiency assumption ments. Therefore, our assumption I > 0 is false, I(Z inv ; X) > c is typically reasonable in practice and we must have because invariant features often capture fundamental aspects of data that remain stable across spu inv I(Y ; Zm | Zm , E) = 0 environments. For instance, in vision-language tasks, semantic content tends to remain consistent Together, these three parts prove that the optimal even when visual appearance changes. Formally, solution satisfies all three claimed properties. by the mutual information chain rule: Remark on finite λ. The limit λi → ∞ is an ideI(Y ; X) = I(Y ; Z inv , Z spu ) = I(Y ; Z inv ) alization. In practice, taking λi sufficiently large (but finite) yields approximations where each reg+ I(Y ; Z spu |Z inv ) − I(Z inv ; Z spu |Y ) ularization term is bounded by a small tolerance Since I(Y ; Z spu |Z inv ) ≤ H(Z spu |Z inv ) and ϵ, and the conclusions hold up to ϵ errors. This is I(Z inv ; Z spu |Y ) ≥ 0, we have: standard in constrained optimization and does not affect the practical validity of the theorem. I(Y ; X) ≤ I(Y ; Z inv ) + H(Z spu |Z inv ) A.3 Distributionally Robust Risk Advantage = I(Y ; Z inv ) + H(X|Z inv ) of Invariant Representations ≤ I(Y ; Z inv ) + H(X) − I(X; Z inv ) Having established how to obtain causal invariant modality representations, we now prove Rearranging terms: their theoretical advantages for worst-case out-ofdistribution risk under distribution shift. I(Y ; Z inv ) ≥ I(Y ; X) − (H(X) − I(X; Z inv )) 15

Ee∈Etest [Re (hraw )] satisfies:

Setting ϵ(c) = H(X) − I(X; Z inv ), when I(X; Z inv ) > c, we have ϵ(c) < H(X) − c, ensuring that I(Y ; Z inv ) approaches I(Y ; X) as c increases. This indicates that the invariant representations Z inv retain most of the predictive information about Y that is present in the raw features X. Moreover, regrading assumption 3, in real-world multimodal learning, the Lipschitz continuity assumption is widely applicable. For classification tasks using cross-entropy loss, when input features are normalized (as is common practice), the loss becomes Lipschitz continuous. Similarly, mean squared error loss for regression tasks is inherently Lipschitz continuous. This regularity ensures stable optimization and meaningful generalization bounds. Since ℓ is assumed to be L-Lipschitz continuous, the risk gap is bounded (Song et al., 2025):

Rmax − R̄ =

R (hinv ) = R (hinv ) = C,

(Rmax − Re (hraw ))

e∈Etest

which implies the existence of environment e∗ such that: σ ∗ Re (hraw ) ≥ R̄ + |Etest | In practice, σ is often substantial when dealing with significant domain shifts, making this inequality practically relevant. Step 4 (Worst-Case Risk Comparison): This step reveals why invariant predictors typically outperform raw predictors in worst-case scenarios. The generalization errors δraw and δinv represent the gap between training and testing performance. Notably, δinv is typically smaller than δraw because invariant features generalize better across environments. In contrast, σ is often large in realworld scenarios with significant distribution shifts. For example, in cross-domain sentiment analysis, performance can vary dramatically between domains (e.g., movie reviews vs. product reviews), resulting in large σ values. From Steps 1 and 2:

where h∗inv and h∗raw are the optimal predictors using invariant and raw representations respectively. This shows that when c is sufficiently large, Z inv almost completely preserves the information needed for prediction. Step 2 (Environment Invariance Property): The invariance condition is realistic in many applications where causal factors remain stable despite environmental changes. For example, in object recognition, shape and category remain invariant while lighting, pose, and background may vary. This property ensures that hinv exhibits consistent performance across environments: e′

|Etest |

X

σ Rmax − Rmin = ≥ |Etest | |Etest |

Re (h∗inv ) ≤ Re (h∗raw ) + L · ϵ(c)

e

1

ROOD (hinv ) = C ≤ Re (h∗raw )+L·ϵ(c),

∀e ∈ Etest

By empirical risk minimization theory, standard generalization bounds apply: Re (hinv ) ≤ Re (h∗inv ) + δinv ,

∀e, e ∈ Eall

∀e ∈ Etest

where δinv can be very small by realization using an expressive and suitable neutral network for the predictor, which decreases with the number of training samples and depend on model complexity. Moreover, as h∗raw is the optimal predictor for raw features, we have:

where C is a constant. Consequently: ROOD (hinv ) = max Re (hinv ) = C e∈Etest

This consistency is a key advantage of invariant representations in distribution shift scenarios. Step 3 (Risk Variability of Raw Representations): For predictor hraw , the risk variability ′ σ = maxe,e′ ∈Etest |Re (hraw ) − Re (hraw )| is typically large in real-world applications where distribution shifts are significant. This variability stems from the model’s reliance on features that change with environment. Let Rmax = maxe∈Etest Re (hraw ) and Rmin = mine∈Etest Re (hraw ). The average risk R̄ =

Re (hraw ) ≥ Re (h∗raw ),

∀e ∈ Etest

From Step 3: ∗

ROOD (hraw ) ≥ Re (hraw ) σ ≥ R̄ + |Etest | ≥ Ee∈Etest [Re (h∗raw )] +

16

σ |Etest |

Similarly:

where PLM indicates the pre-trained language model, Ul is the input token sequence and Tl represents the sequence length. Wpro ∈ Rdl ×d and bpro ∈ R1×d are trainable parameters that map the output dimensionality of the language network to the shared feature dimensionality d. In MSA, the acoustic and visual networks employ transformer encoders (Vaswani et al., 2017) and operate according to the following steps (m ∈ {a, v}):

ROOD (hinv ) = C = Ee∈Etest [Re (hinv )] ≤ Ee∈Etest [Re (h∗raw )] + L · ϵ(c) + δinv Combining these results: ROOD (hraw ) − ROOD (hinv )   σ e ∗ ≥ Ee∈Etest [R (hraw )] + |Etest | e ∗ − (Ee∈Etest [R (hraw )] + L · ϵ(c) + δinv ) σ = −L · ϵ(c) − δinv + |Etest |

X̂m = Conv 1D(Um ; Km ) ∈ RTm ×d Xm = Transformer(X̂m ; θm ) ∈ RTm ×d

where Um ∈ RTm ×dm is the extracted raw feature sequence (see Section E for the extraction derails), Conv 1D indicates the temporal convolution whose kernel size Km is set to 3. Nota that for the CH-SIMS-v2 dataset, we use the same feature set as in previous works (Feng et al., 2024; Liu et al., 2022), which are feature vectors instead of sequences. Therefore, we simply use multi-layer perception networks as the unimodal networks for visual and acoustic modalities.

The inequality ROOD (hraw ) > ROOD (hinv ) holds when: σ > |Etest |(L · ϵ(c) + δinv ) This condition is typically satisfied in practice because: • σ tends to be large in real-world domain shifts, as environments often differ significantly in their distributional properties.

In the MHD and MSD tasks, to more effectively model humor-specific cues, we follow prior work (Hasan et al., 2021) by extracting an additional Humor-Centric Feature (HCF) from the language modality. This HCF serves as a fourth modality and is represented as Uh ∈ RTl ×dh (detailed in (Hasan et al., 2021)). In addition, each data sample in the MHD and MSD tasks includes both a target punchline segment and its preceding context. We merge the feature sequences of the punchline and context along the temporal axis to construct the unimodal input representations Um ∈ RTm ×dm (m ∈ M = {a, v, l, h}). The unimodal network for the HCF modality is analogous to the framework employed for the visual and acoustic modalities. The specific steps of the transformer-based unimodal networks for MHD and MSD are delineated as follows (m ∈ {a, v, h}):

• δinv is small because invariant features generalize well across environments and we can adopt proper realization of the predictor h to reduce δinv . • ϵ(c) can be made small by ensuring Z inv captures sufficient information from X. This proves that under realistic conditions, predictors based on invariant representations achieve strictly lower worst-case out-of-distribution risk than those using raw representations.

B

(11)

Unimodal Networks

This section outlines the architecture of our unimodal networks and elaborates on the steps for generating unimodal representations, which serve as the foundation for subsequent causal analysis. To make a fair comparison, following established practices in recent work (Xiao et al., 2024; Mai et al., 2023b), we utilize pre-trained language models (He et al., 2021; Lan et al., 2020) to derive high-quality textual features. The following steps outline the language network’s workflow for all downstream tasks X̂l = PLM(Ul ; θl ) ∈ RTl ×dl (10) Xl = (X̂l Wpro + bpro ) ∈ RTl ×d

X̂m = Transformer(Um ; θm ) ∈ RTm ×dm

(12) Xm = Conv 1D(X̂m ; Km ) ∈ RTm ×d Finally, we employ a straightforward linear layer to fuse the language and HCF modalities, thereby reducing complexity for the following model stages: Xl ←− Linear(Xl ⊕ Xh ; θlin ) ∈ RTl ×d (13) For all unimodal representations Xm ∈ RTm ×d , we perform mean pooling at the time dimension to obtain the final unimodal representations Xm ∈ R1×d . 17

C

Datasets

are treated as positive samples, while those without laughter form negative samples. It is partitioned into 7,614 training, 980 validation, and 994 test instances. (6) MUStARD (Castro et al., 2019): The MUStARD dataset is designed for multimodal sarcasm detection (MSD), comprising video segments sourced from popular TV series including Friends, The Big Bang Theory, The Golden Girls, and Sarcasmaholics. The collection contains 690 humanannotated segments, labeled as either sarcastic or non-sarcastic. Similar to UR-FUNNY, it provides contextual clues by incorporating both the target punchline and the preceding dialogue segments for each sample.

(1) CMU-MOSI (Zadeh et al., 2016): This dataset is a standard benchmark for MSA, encompassing more than 2,000 online video clips collected from the Internet. Each clip is labeled with a sentiment score on a -3 to 3 Likert scale, where 3 and -3 denote extreme positive and negative sentiments, respectively. (2) The OOD version of CMU-MOSI (Sun et al., 2022): CMU-MOSI (OOD) is built using a modified simulated annealing algorithm (Aarts et al., 1987), which iteratively adjusts the test distribution. The resulting significant shifts in word–sentiment correlations relative to the training set establish it as a challenging benchmark for evaluating model robustness to distribution shifts in MSA. (3) CMU-MOSEI (Zadeh et al., 2018b): The CMU-MOSEI dataset is a large-scale, widelyadopted benchmark for MSA, collected from online videos. Its key characteristics include: (I) Scale: over 22,000 video clips; (II) Source: more than 1,000 YouTube speakers and 250+ topics, randomly sampled; (III) Annotation: each clip has two labels—a six-class emotion category and a sentiment score ranging from -3 (strongly negative) to 3 (strongly positive). For the MSA task, our evaluation adopts the sentiment labels of CMUMOSEI, which are consistent with the scale used in CMU-MOSI. (4) CH-SIMS-v2 (Liu et al., 2022): CH-SIMSv2 serves as a Chinese MSA benchmark with the following characteristics: (I) Source: Videos collected from 11 scenarios (interviews, talk shows, films, etc.) to mimic real-world interaction; (II) Quality: Filtered to retain high-quality acoustic and visual streams; (III) Split: Partitioned into training, validation, and test sets in a 9:2:3 ratio, corresponding to 2,722, 647, and 1,034 segments respectively; (IV) Label Distribution: The training set contains 921 negative, 433 weakly negative, 232 neutral, 318 weakly positive, and 818 positive samples. (5) UR-FUNNY (Hasan et al., 2019): Derived from TED talk videos involving 1,741 speakers, UR-FUNNY serves as a benchmark for multimodal humor detection (MHD). Each data sample includes a multimodal punchline segment and its preceding context segments, the latter being provided to support contextual modeling. The dataset is built by identifying punchlines via the laughter tag in transcripts. Video segments followed by laughter

D

Evaluation Metrics

For CMU-MOSI and CMU-MOSEI datasets, we adopt the following evaluation metrics: (1) Acc7: the accuracy of classifying sentiment scores into seven discrete classes; (2) Acc2: the binary accuracy for differentiating between positive and negative sentiments; (3) F1 score: a harmonic mean that balances precision and recall for binary sentiment classification; (4) MAE: the mean absolute error between model predictions and sentiment labels; and (5) Corr: the correlation coefficient reflecting the strength and direction of the relationship between predictions and sentiment labels. For Acc7, predictions are rounded to the nearest integer within the scale from -3 to 3. When calculating Acc2 and F1 score, neutral segments are not considered. And the neutral segments are included in the calculations of MAE, Corr, and Acc7. For the CH-SIMS-v2 dataset, we use Acc5, Acc3, Acc2, F1 score, MAE, and Corr as in previous works (Liu et al., 2022; Feng et al., 2024). For the MHD and MSD tasks, we report the binary accuracy (i.e., humorous or non-humorous, sarcastic or non-sarcastic) of the model.

E

Feature Extraction Details

Regarding the visual modality, following previous approaches (Mai et al., 2023d; Xiao et al., 2024), Facet 1 is used to gather an array of visual attributes such as facial action units and facial landmarks for the MSA task. Facial feature extraction for the CH-SIMS-v2 dataset aligns with established practice (Feng et al., 2024; Liu et al., 2022), utilizing 1

18

iMotions 2017. https://imotions.com/

Table 6: The unimodal feature dimensionality of different datasets. CMU-MOSI CMU-MOSI (OOD) CMU-MOSEI CH-SIMS-v2 UR-FUNNY MUStARD

Language 768 768 768 768 768 768

Acoustic 74 74 74 25 60 60

Visual 47 47 35 177 36 36

HCF 4 4

OpenFace (Baltrusaitis et al., 2018) to obtain measures such as 68 facial landmarks, 17 action units, head pose, head orientation, and eye gaze direction. For MHD and MSD, in line with prior approaches (Hasan et al., 2021; Mai et al., 2023a), OpenFace 2 (Baltrusaitis et al., 2018) is utilized for the extraction of facial action unit features as well as rigid and non-rigid facial shape parameters. For acoustic modality, COVAREP (Degottex et al., 2014) is used for the extraction of a sequence of acoustic features, including 12 Mel-frequency cepstral coefficients, pitch tracking, speech polarity, etc. For the CH-SIMS-v2 dataset, acoustic features are represented as 25-dimensional eGeMAPS low-level descriptors (LLD), extracted via OpenSmile (Eyben, 2010) at 16 kHz. For language modality, following state-of-the-art methods (Xiao et al., 2024), DeBERTa (He et al., 2021) and BERT (Devlin et al., 2019) are employed to learn informative language representations. For the CH-SIMS-v2 dataset, we follow prior works (Feng et al., 2024; Liu et al., 2022) and utilize BERT (Devlin et al., 2019) to obtain textual features. For MHD and MSD, following prior methods (Mai et al., 2023b,a), ALBERT (Lan et al., 2020) is adopted. The input feature dimensionality for each modality is summarized in Table 6.

F

Figure 7: The structures of encoder, decoder, and pre′ dictor. d represents the hidden dimensionality.

The structures of the predictive attention weight generator and the predictor are shown in Figure 7. (2) Protocol for the evaluation of noisy modalities: To evaluate model robustness with noisy modalities, we produce corrupted features by employing the equation below: n Xm = (1 − N R) · Xm + N R · N (14) where Xm is the unimodal representation, N R is the noisy rate ranging from 0.1 to 0.7, N is the Gaussian noise data of mean 0 and variance 1, and n is the noisy unimodal representation that is Xm used to learn invariant representation. To simulate realistic noise, the described noise mixing process is applied to every modality across all samples. Perturbations are introduced at the feature level because input noise ultimately propagates to feature representations. This approach aligns with common practice in MAC, where a standardized feature set is typically used to ensure fairness and protect privacy, making feature-level noise injection both reasonable and practical. For a fair comparison, all baselines (Yuan et al., 2024; Mai et al., 2023c, 2024) are reproduced using the same training and testing protocols as our CmIR. To further assess the robustness of CmIR under out-of-distribution (OOD) noise conditions (Table 4), we evaluate its performance against two additional corruption types: Laplace noise and random erasing noise (the latter simulates data loss by randomly zeroing out a subset of features). During this evaluation, each sample has an equal 50% chance of being corrupted either by Laplacian noise via Eq. 14 or by random feature dropout (missing), where the dropout rate is controlled by the noise ratio N R.

Experimental Details

(1) Hyperparameter Setting: Our proposed CmIR is developed with PyTorch 1.13.1 on an NVIDIA RTX3090 GPU (CUDA 11.6). Training utilizes the AdamW optimizer (Loshchilov and Hutter, 2019). In line with prior work (Mai et al., 2023d), optimal hyperparameters are determined through an extensive random grid search of 50 iterations on the validation set. Using the identified best configuration, the model is retrained five times, and the final performance is averaged across these runs. The specific hyperparameter settings can be found in Table 7. Note that the modality-specific covariance matrix Σm in Eq. 3 is set as the identity matrix. 19

Table 7: Hyperparameter Settings of CmIR. MAE, MSE and BCE denote mean absolute error, mean square error and binary cross-entropy, respectively. Loss Function Batch Size Learning Rate Number of Environments K Noise Intensity α(1) Weight λ1 Weight λ2 Weight λ3 Shared Dimensionality d ′ Hidden Dimensionality d Weight α in Mutual Information Constraint

G

CMU-MOSI MSE 48 1e-5 1 0.1 0.1 0.001 0.05 150 256 1

CMU-MOSEI MSE 50 1e-5 5 1 0.1 0.1 0.01 150 128 1

Baselines

CH-SIMS-v2 MAE 50 5e-4 2 0.1 0.01 0.01 0.1 100 100 0.01

MUStARD BCE 32 2e-5 4 0.1 0.05 0.01 0.05 128 48 0.1

UR-FUNNY BCE 64 7e-6 2 0.1 0.001 0.01 0.01 120 120 0.1

CMU-MOSI (OOD) MSE 48 1e-5 4 0.1 0.01 0.01 0.01 150 150 1

module that dynamically estimates the contribution and noise degree of individual learners; (2) Complete Multimodal Information Bottleneck (CMIB) (Mai et al., 2023c): It adopts the information bottleneck principle to eliminate redundancy and noise from both unimodal and multimodal features, thus establishing it as a baseline for handling noisy modalities.; (3) Two-stage Modality Denoising and Complementation (TMDC) (Zhuang et al., 2025): Its training process involves two distinct stages. The first, the intra-modality denoising stage, aims to enhance representational robustness by using denoising modules to obtain clean modalityspecific and shared representations from complete data, thereby reducing noise interference. The second, the inter-modality complementation Stage, utilizes these representations to address modality absence through cross-modal compensation.

The causality-based baselines include: (1) Subject Causal Intervention (SuCI) (Xu et al., 2025b): It introduces a simple yet effective causal intervention module designed to decouple the influence of subjects as unobserved confounders, thus obtaining unbiased predictions through true causal effects; (2) CounterfactuaL mUltimodal sEntiment (CLUE) (Sun et al., 2022): It leverages causal inference and counterfactual reasoning to prune away spurious direct textual influences, preserving only the genuine indirect multimodal effects, thereby strengthening generalization to out-of-distribution data; (3) General dEbiAsing fRamework (GEAR) (Sun et al., 2023): To enhance out-of-distribution robustness, it distinguishes robust features from biased ones, quantifies sample-level bias, and employs inverse probability weighting to de-emphasize highly biased samples; (4) Multimodal Debiasing Framework (MulDeF) (Huan et al., 2024): It integrates causal intervention with front-door adjustment and multimodal causal attention during the training phase. At inference, it applies counterfactual reasoning to mitigate both verbal and nonverbal biases, which enhances out-of-distribution generalization.; (5) Attention-based Causality-Aware FusionAttention-based Causality-Aware Fusion (AtCAF) (Huang et al., 2025): It learns causalityaware multimodal representations for sentiment analysis through a dedicated text debiasing module and counterfactual attention across modalities. The baselines for handling noisy modality include: (1) Multimodal Boosting (Mai et al., 2024): It is built upon multiple base learners organized in a boosting-like manner, with each learner addressing different facets of the multimodal data. It is further equipped with a contribution learning

The additional baselines for MSA include: (1) Information-Theoretic Hierarchical Perception (ITHP) (Xiao et al., 2024): Its design is grounded in the information bottleneck principle, where one modality is designated as the core, while others act as detectors within the information pathway to distill and refine the information flow; (2) Self-Supervised Multi-task Multimodal sentiment analysis network (Self-MM) (Yu et al., 2021): This approach employs a self-supervised strategy to infer sentiment labels for individual modalities by leveraging the global labels of multimodal samples, thereby learning more discriminative unimodal representations; (3) DisentangledLanguage-Focused Model (DLF) (Wang et al., 2025): It introduces a feature disentanglement module to separate shared and specific information across modalities. The process is further refined by four geometric measures to reduce redundancy and prioritize language-focused features. A languagetargeted attractor is also designed to enhance lan20

Multimodal Emotion Recognition (FDMER) (Yang et al., 2022): It learns the common and private feature representations for each modality, which achieves the modality consistency and disparity constraints by designing tailored losses for modality-invariant and modality-specific subspaces. The additional baselines for MHD and MSD include: (1) Multimodal Global Contrastive Learning (MGCL) (Mai et al., 2023b): MGCL applies supervised contrastive learning to multimodal representations, employing diverse augmentation strategies to construct positive and negative sample pairs for each representation; (2) Multimodal Correlation Learning (MCL) (Mai et al., 2023a): MCL formulates a supervised correlation learning objective that preserves modality-specific characteristics while fostering a more discriminative joint embedding space; (3) Multimodal Multitask Interaction Learning (MIL) (Zhang et al., 2024): It performs joint sarcasm and sentiment detection, integrating a cross-modal target attention mechanism to align textual with visual/acoustic content and a multimodal interaction module to model the shared and distinct patterns of both tasks; (4) Decoupled Multimodal Distillation (DMD) (Li et al., 2023): DMD enables adaptive cross-modal knowledge transfer by decoupling unimodal features into modality-irrelevant and modality-exclusive subspaces, followed by a specialized graph distillation unit to handle each subspace effectively; (5) Humor Knowledge Enriched Transformer (HKT) (Hasan et al., 2021): HKT incorporates humor-centric features as external knowledge to resolve the ambiguity and leverage the subtle sentiment cues within the language modality; (6) Multimodal Transformer (MulT) (Tsai et al., 2019a): MulT leverages stacked cross-modal transformers to project and align source modalities to a target modality, effectively mitigating the inter-modal gap.

guage representations using complementary information from other modalities; (4) Gradient and Structure Consistency (GSCon) (Shi et al., 2025): It proposes a balanced gradient direction that aligns each modality’s optimization direction to ensure unbiased convergence, and aligns the spatial structure of samples in different modalities to avoid the interaction noise caused by multimodal alignment; (5) Diffusion Bridge (Lee et al., 2025): It directly reduces the modality gap by leveraging denoising diffusion probabilistic models; (6) Enhanced Dynamic Emotion Experts (EMOE) (Fang et al., 2025): The framework comprises a mixture of modality experts that dynamically adjust modality importance based on input features, coupled with a unimodal distillation mechanism to preserve the predictive capacity of individual modalities within the fused representation; (7) Contrastive FEature DEcomposition (ConFEDE) (Yang et al., 2023): It enhances multimodal representations by jointly performing contrastive learning and contrastive decomposition of features; (8) Kolmogorov–Arnold Network with Multimodal Clean Pareto KANMCP (Luo et al., 2025a): It combines interpretable cross-modal modeling via KANs with feature denoising and compression based on DRDMIB, yielding discriminative multimodal inputs while alleviating modality imbalance; (9) Multimodal Adaptation Gate BERT/ALBERT (MAGBERT/MAG-ALBERT) (Rahman et al., 2020): It incorporates visual and acoustic information into BERT/ALBERT through a dedicated multimodal adaptation gate; (10) Acoustic Visual Mixup Consistent (AV-MC) (Liu et al., 2022): This method utilizes modality mix-up to augment visual and acoustic representations, thereby strengthening their contribution to sentiment analysis; (11) Knowledge-Guided Dynamic Modality Attention Fusion (KUDA) (Feng et al., 2024): It guides the multimodal fusion process with external emotional knowledge, dynamically selecting the dominant modality and adjusting the weighting of all modalities; (12) MISA (Hazarika et al., 2020): It decomposes each unimodal input into modalityinvariant and modality-specific components, which are subsequently fused for final prediction; (13) MultiModal InfoMax (MMIM) (Han et al., 2021): In MMIM, representation learning is enhanced by maximizing the mutual information both between unimodal features and between multimodal representations at different levels and their unimodal counterparts; (14) Feature-Disentangled

H

Comparison between Different Invariance Constraints

In Section 4.2, we propose two different methods to realize invariance constraint. The first one is a standard way to compute and minimize the conditional distribution distance (we directly minimize unimodal predictions from different environments for regression tasks, and use KL-divergence to min21

Table 8: Further discussions on different invariance constraints.

Acc7 55.2 55.1

imize conditional distribution distance for classification tasks), which needs to implement a unimodal predictor for each modality and generate unimodal prediction. The second one is to directly minimize the distance between different causal invariant representations extracted under different environments, which is adopted in our CmIR. Here we provide a comparison between these two strategies, and the results are shown in Table 8. From Table 8 we can infer that these two strategies actually yield comparable outcomes, likely because although the first method can more directly evaluate the differences between conditional distributions, it introduces a unimodal predictor. However, when the unimodal predictor is insufficiently trained, it may introduce evaluation errors. Therefore, we adopt the second approach, which enables a more rigorous constraint without introducing a unimodal predictor, thereby reducing model complexity.

J

I

Acc2 90.1 89.6

CMU-MOSI F1 score MAE 90.1 0.616 89.5 0.616

Corr 0.855 0.853

Standard Ours

Acc7 48.4 49.8

Acc2 87.4 87.8

CMU-MOSEI F1 score MAE 87.3 0.523 87.7 0.513

Corr 0.781 0.793

MUStARD Acc 79.7 80.0

Significant Test

In this section, we provide the results of significant test between the proposed CmIR and a strong baseline ITHP (Xiao et al., 2024) using t-test on the CMU-MOSI and CMU-MOSEI datasets. We run each model for ten times under different random seeds to gather the results. As shown in Table 10, the t-test results indicate that CmIR has a statistically significant difference with ITHP (Xiao et al., 2024) for all evaluation metrics except the Corr metric (p < 0.05 denotes a statistically significant difference), suggesting that the improvement of CmIR is significant.

K Probing Experiments for Z inv and Z spu Orthogonality is only a practical proxy for mutual information minimization. To strengthen the evidence that Z inv and Z spu are well separated and contain different information, we have added two probing experiments: (a) training a predictor to predict the environment label from Z inv or Z spu alone. Since the difference between different environments lies in the noise level (noise coefficient α(e) ), we directly train an additional fusion network and predictor that use Z inv / Z spu to predict the noise coefficient α(e) . Finally, we use MSE to evaluate the effectiveness of the prediction (lower accuracy from Z inv indicates better invariance); (b) Predicting the sentiment label from each component separately. As shown in Table 11, Z inv yields high MSE in environment prediction but high accuracy in sentiment prediction, while Z spu does the opposite, confirming semantic separation. Notably, Z spu only relies on label bias for prediction (predicting all samples as positive yields an accuracy of 62.8% which is exactly the accuracy of the prediction from Z spu ), suggesting it contains non-causal information.

Model Complexity Analysis

The proposed CmIR does not design complex multimodal fusion to explore sufficient inter-modal interactions, but merely introduces an encoder and a decoder to learn invariant modality representations for each modality. Each encoder/decoder only has a few layers. Therefore, the model complexity of the proposed CmIR is acceptable. To verify our claim, we compare the model complexity of CmIR with competitive baselines on the MUStARD (Castro et al., 2019) and CMU-MOSI (Zadeh et al., 2016) datasets. As shown in Table 9, on MUStARD, CmIR outperforms competitive baselines with minimal number of parameters, demonstrating the effectiveness of the proposed framework and indicating that the performance improvement of CmIR dose not results from the increase in the number of parameters. On CMU-MOSI, we evaluate the FLOPs and number of parameters of our CmIR. As shown in Table 9, CmIR has significantly fewer parameters than EMOE and GSCon, slightly more than Diffusion Bridge, while its FLOPs are lower than all baselines, indicating the high efficiency of our method.

L

Sensitivity Analysis of Noise Intensity

In this section, we have conducted new experiments on CMU-MOSI (OOD) and provided the sensitivity analysis of noise intensity α(1) in Table 12. As shown in Table 12, when the value of α(1) is small, 22

Table 9: Model complexity analysis on the MUStARD and CMU-MOSI datasets. MUStARD Model Acc2 HKT (Hasan et al., 2021) 76.47 MCL (Mai et al., 2023a) 77.94 MGCL (Mai et al., 2023b) 77.94 CmIR 80.00

Parameters (M) 17.10M 13.83M 14.28M 13.44

CMU-MOSI Model Acc2 Parameters (M) EMOE (Fang et al., 2025) 84.8 317.30M GSCon (Shi et al., 2025) 88.1 248.52M Diffusion Bridge (Lee et al., 2025) 86.9 185.46M CmIR 89.6 187.51M

Table 10: Paired t-test analysis between CmIR and ITHP (Xiao et al., 2024).

ITHP

Acc7 0.002

Acc2 0.004

CMU-MOSI F1 MAE 0.004 4.98e-4

Corr 0.439

CmIR performs poorly. This is because when the imposed noise is small, its impact on modality representations is limited, making it difficult to simulate diverse environments for learning effective causal invariant representations. Moreover, when α(1) fluctuates within a large range (from 1e-7 to 1), the model maintains good performance, demonstrating the robustness of CmIR. Additionally, when α(1) = 1, the model achieves stronger performance than the default setting (α(1) = 0.1), indicating the performance potential of CmIR (better results can be obtained through more careful hyperparameter tuning). Compared with the number of environments K (see Figure 5 (d)), α(1) has a greater impact and should be set to a relatively large value.

23

Acc7 3.65e-5

CMU-MOSEI Acc2 F1 MAE 5.73e-4 7.78e-4 1.12e-4

Corr 0.566

FLOPs (G) 10.02 12.22 11.52 8.95

Table 11: Probing results on CMU-MOSEI.

Component Z inv Z spu

Environment Prediction MSE 2.04 0.87

Acc7 55.1 41.4

Sentiment prediction Acc2 F1 MAE 87.8 87.7 0.513 62.8 48.5 0.841

Corr 0.793 0.001

Table 12: Sensitivity Analysis of Noise Intensity α(1) on CMU-MOSI (OOD).

α(1) Acc7 Acc2 F1 score

1e-8 43.5 79.7 79.7

1e-7 45.8 82.8 82.7

1e-6 47.5 84.1 83.9

1e-5 47.0 83.6 83.5

24

1e-4 47.3 84.0 84.1

1e-3 48.0 84.4 84.3

1e-2 48.8 83.9 83.9

1e-1 48.5 84.4 84.4

1 49.0 84.6 84.6

Record · ID 120523 · SHA-256 dcfa25e5579368ce
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.