arXiv:2606.27114v1 [cs.LG] 25 Jun 2026
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding Haoran Zhang
Chuanpu Li
Yuxin Fu
Center for Applied Statistics and School of Statistics, Renmin University of China Beijing, China [email protected]
Alibaba Group Beijing, China
Alibaba Group Beijing, China
Bin Tong
Guan Wang
Bo Zheng
Alibaba Group Beijing, China
Alibaba Group Beijing, China
Alibaba Group Beijing, China
Feng Zhou∗ Center for Applied Statistics and School of Statistics, Renmin University of China Beijing, China [email protected]
Abstract
CCS Concepts
Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios. In this paper, we propose the Cross-Head Attention Uplift Network (CHAUN) and Robust Adversarial Inverse Propensity Score (RA-IPS) method to address these limitations. CHAUN employs shared feature embeddings and crosshead attention mechanisms to dynamically integrate treatmentspecific and control-specific representations, enhancing inter-group correlation modeling. Theoretically, we prove that access to the true propensity scores ensures ITE identifiability even with unobserved confounders. For practical scenarios lacking true propensity scores, RA-IPS adversarially optimizes propensity weights within constrained uncertainty sets to mitigate bias from unobserved variables. Experiments on public datasets (CRITEO-UPLIFT, LAZADA) and a production e-commerce dataset demonstrate CHAUN’s superiority over state-of-the-art uplift models, achieving a relative improvements of up to 25.6% in QINI scores. RA-IPS further enhances robustness, outperforming standard IPS by 5.4% under unobserved confounding. The results validate the effectiveness of our proposed methods in real-world causal inference tasks.
• Information systems → Information systems applications.
∗ Corresponding author.
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. Conference acronym ’XX, Woodstock, NY © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/XXXXXXX.XXXXXXX
Keywords Causal Inference, Uplift Modeling, Unobserved Confounders, Inverse Propensity Score ACM Reference Format: Haoran Zhang, Chuanpu Li, Yuxin Fu, Bin Tong, Guan Wang, Bo Zheng, and Feng Zhou∗ . 2026. Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’XX). ACM, New York, NY, USA, 12 pages. https://doi. org/XXXXXXX.XXXXXXX
1
Introduction
Uplift modeling, a causal inference technique for quantifying the marginal effect of treatments on individual behavior, plays a pivotal role in real-world applications such as advertisement [15, 36], user growth [2, 7], and online marketing [18, 42]. The objective of uplift models is to estimate the causal effect of administering a treatment to an individual unit, formally defined as the Individual Treatment Effect (ITE), or uplift. Uplift modeling enables prediction of ITE at the unit level, thereby facilitating the implementation of precision intervention strategies through data-driven decision frameworks. We focus on the canonical scenario of binary treatment allocation, whereby observational units are exclusively assigned to either the treatment group or control group. Uplift modeling differs from conventional supervised learning approaches due to the inherent unobservability of ITE labels. For each observational unit, researchers can only empirically measure the factual outcome under either the treatment condition or the control condition, while the corresponding counterfactual outcome remains fundamentally
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
inaccessible. Consequently, the accurate prediction of ITE requires valid counterfactual estimation frameworks for observational units. One of the essential characteristics of uplift modeling lies in the inherent inter-group similarity between the prediction tasks for treatment group outcomes and control group outcomes. Although the supervision of the treatment group outcomes and control group outcomes is conducted independently, the two prediction tasks remain inherently correlated. Existing methods that incorporate treatment as an input feature [11, 18] effectively leverage similarity patterns but fundamentally lack explicit mechanisms to sufficiently capture treatment-induced variation. In contrast, multi-head architectures [34, 35] based on multi-task learning frameworks lack flexible utilization of such inter-group similarity. Currently, no flexible architecture simultaneously emphasizes treatment-induced discrimination and exploits inter-group similarity. Another issue to be addressed in uplift modeling is selection bias [12], which indicates the presence of confounding variables that simultaneously influence both treatment assignment and outcome responses across observational units. Without proper adjustment for confounders, it becomes challenging to isolate the treatment effect from the compounded variance attributable to multifactorial influences in observational causal inference. Traditional statistical research conventionally employs propensity score-based methodologies [1, 27, 28] for sample debiasing. Recent uplift models based on neural network address selection bias through specialized architectures and tailored loss functions [9, 34, 42]. However, these methods fundamentally rely on the strong ignorability assumption [25] (a.k.a. unconfoundedness), which posits that all confounding variables are captured within the observed covariates. The strong ignorability assumption serves as a sufficient condition for the identification of ITE [12, 26], and typically holds when selection bias originates from observed covariates. In complex systems such as bidding advertising systems, the treatment assignment mechanism operates through a multi-stage causal pathway: advertisers’ targeting policies determine eligibility for ad exposure, while the impression delivery is mediated by the bidding platform’s allocation algorithms. This hierarchical selection process introduces latent confounders that remain unobservable to advertisers. In scenarios with unobserved confounders, identifying the ITE and debiasing samples remains challenging. In this paper, we address the two aforementioned challenges. To flexibly exploit inter-group similarity, we propose the Cross-Head Attention Uplift Network (CHAUN). The CHAUN framework first learns shared feature embeddings across users through representation alignment, then generates treatment-specific and controlspecific deep latent representations via parallel encoding pathways. These dual-branch representations are adaptively fused through attention-weighted integration, where attention mechanisms dynamically determine cross-representation interaction weights. The synthesized representation is ultimately utilized for uplift prediction, enabling counterfactual inference by contrasting treatmenteffect responses under treatment and control conditions. To mitigate selection bias, we incorporate inverse propensity score (IPS) weighting into the loss function and impose a global regularization constraint on the propensity scores to stabilize the weights. To address debiasing in the presence of unobserved confounders, we establish that access to the true propensity scores enables the
Trovato et al.
identification of ITE and unbiased IPS estimator. In practical operational settings where only nominal propensity scores estimated from observed features are accessible, we develop Robust Adversarial Inverse Propensity Scores (RA-IPS). We assume that the true propensity score is constrained to vary within a neighborhood of the nominal propensity score, and we optimize for the worst-case scenario to enhance the model’s robustness against unobserved confounding. Specifically, our contributions are summarized as follows: • We propose CHAUN, a generalized uplift modeling framework that adaptively mediates cross-head interaction through attention gating within a dual-head architecture. • We demonstrate that under unobserved confounding, access to the true propensity scores is sufficient to ensure both the identifiability of ITE and the unbiasedness of IPS estimator. • We propose RA-IPS, a robust inverse propensity score method that addresses unobserved confounders without requiring true propensity scores, leveraging adversarial weighting to enhance stability under unobserved confounding. • We conduct extensive experiments on two public datasets and a production dataset to validate the effectiveness of the proposed CHAUN and RA-IPS methods.
2
Related Work
In this section, we review existing approaches that leverage intergroup similarity inherent in uplift modeling frameworks, address selection bias, mitigate the impact of unobserved confounders.
2.1
Exploiting Inter-group Similarity
Empirical studies [4, 17] suggest that potential outcomes under treatment and control conditions share similar functional structures or model parameters, the ITE as their difference inherently exhibits relatively lower complexity. The S-Learner [17] and frameworks like EFIN [18], ECUP [11] that encode treatment indicators as model covariates, naturally exploit this prior. EUEN [15] explicitly models the control-group outcome and ITE, then constructs the treatmentgroup outcome through additive integration of them. FlexTENet [4] implements partial neuron sharing between layers of the treatment head and control head within its dual-head multi-task architecture. Unlike these methods, CHAUN employs cross-head attention over treatment/control representations to exploit inter-group similarity.
2.2
Addressing Selection Bias
In statistical research, practitioners typically employ propensity scores to perform matching [28] or inverse probability weighting [10], thereby constructing pseudo-RCT samples that approximate randomized experimental conditions. In neural network-based approaches, a multitude of methods leverage architectural designs and regularization constraints to mitigate selection bias. BNN [13] learns a shared feature representation network for both treatment and control groups, while employing Integral Probability Metrics loss to constrain the distributional discrepancy between group representations. CFRNet [34] retains the shared feature representation, while employing a multi-head architecture with treatment-specific heads to learn group-specific outcome predictions. Building upon CFRNet, DR-CFR [9] introduces confounder disentanglement and
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
designs propensity score-based weights that account for both factual and counterfactual occurrence probabilities. DESCN [42] and GNUM [43] enable unified parameter learning across all samples by constructing shared labels for both treatment and control groups. We adopt a weighting approach but introduce a global regularization constraint on the propensity scores to stabilize the weights.
2.3
3
Preliminaries
In this section, we extend the Neyman-Rubin potential outcomes framework [29] to incorporate unobserved confounders, to formalize our uplift modeling problem. We then propose solutions under the requirement of leveraging true propensity scores without direct observation of unobserved confounders.
3.1
confounders. Notably, when unobserved confounders exist, the overlap assumption generalizes to require 0 < 𝜋 (𝑥, 𝑢) < 1, ∀𝑥, 𝑢.
3.2
Problem Definition
Assuming we have a dataset D consisting of 𝑁 samples (𝑥𝑖 , 𝑢𝑖 , 𝑡𝑖 , 𝑦𝑖 ) where 𝑥𝑖 denotes observed features, 𝑢𝑖 denotes unobserved confounders, 𝑡𝑖 ∈ {0, 1} denotes the binary treatment, and 𝑦𝑖 denotes the outcome response. Each instance of D is independently sampled from the joint distribution 𝑝 (𝑥, 𝑢, 𝑡, 𝑦(0), 𝑦(1)). Let 𝜋 (𝑥, 𝑢) = 𝑃 (𝑡 = 1|𝑥, 𝑢) denote the true propensity score, which governs the treatment assignment. The unobservability of 𝑢 renders the true propensity score inaccessible. In practice, we can obtain or estimate the nominal propensity scores 𝜋˜ (𝑥) = 𝑃 (𝑡 = 1|𝑥) based on observed features. The ITE to be estimated is expressed as: 𝐼𝑇 𝐸 (𝑥) = E(𝑦(1) − 𝑦(0)|𝑥). (1) Recent works [18, 37, 42] in uplift modeling under the NeymanRubin potential outcomes framework [29] universally incorporate three foundational assumptions: 1) Ignorability: treatment assignment is conditionally independent of potential outcomes given observed covariates, i.e., 𝑦 (1), 𝑦 (0) ⊥⊥ 𝑡 |𝑥. 2) Consistency: 𝑦𝑖 corresponds to the potential outcome under the administered treatment, i.e., 𝑦𝑖 = 𝑡𝑖 𝑦𝑖 (1) + (1 − 𝑡𝑖 )𝑦𝑖 (0). 3) Overlap: each unit has a non-zero probability to be assigned to each treatment, i.e., 0 < 𝜋˜ (𝑥) < 1, ∀𝑥. The ignorability assumption entails the absence of unobserved confounders. In the subsequent analysis, we maintain the Consistency and Overlap assumptions throughout, while systematically examining scenarios with and without the presence of unobserved
Identifying the ITE
The ignorability assumption ensures the identifiability of ITE by 𝐼𝑇 𝐸 (𝑥) = E(𝑦 (1) − 𝑦(0)|𝑥) = E(𝑦(1)|𝑥) − E(𝑦(0)|𝑥) = E(𝑦(1)|𝑥, 𝑡 = 1) − E(𝑦 (0)|𝑥, 𝑡 = 0)
Handling Unobserved Confounders
When a small amount of RCT data is accessible, it can be utilized to correct the bias in large scale observational data [3, 14, 46]. When RCT data are not possible, causal identification necessitates incorporating structural assumptions or auxiliary information. Common strategies include employing instrumental variables to constrain treatment assignment mechanisms [8, 30, 40], or identifying proxy variables that correlate with unobserved confounders to enable partial recovery of their latent distributional properties [22, 38]. In the absence of supplementary information, methodologies leveraging sensitivity analyses [6, 41] enhance robustness against unobserved confounders through adversarial learning. Our method is based on propensity score sensitivity analysis, but we rigorously derive a more reasonable range for propensity score perturbations.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
(2)
= E(𝑦|𝑥, 𝑡 = 1) − E(𝑦|𝑥, 𝑡 = 0), where E(𝑦|𝑥, 𝑡 = 1) and E(𝑦|𝑥, 𝑡 = 0) can be estimated from observed data. In (2), the penultimate equality is derived from the Ignorability assumption, and the final equality follows from the Consistency assumption. Note that when the ignorability assumption fails, E(𝑦 (𝑘)|𝑥) ≠ E(𝑦(𝑘)|𝑥, 𝑡 = 1) since 𝑝 (𝑦 (𝑘)|𝑥) and 𝑝 (𝑦 (𝑘)|𝑥, 𝑡 = 𝑘) are not equivalent in distribution, 𝑘 = 0, 1. Therefore, under such circumstances, we require additional information about the distribution of 𝑢 to resolve the identifiability of ITE. Diverging from existing approaches that introduce instrumental variables [8, 30, 40] or proxies [22, 38], we draw inspiration from the sufficiency of propensity scores in [35], demonstrating that precise knowledge of the true propensity scores 𝜋 (𝑥, 𝑢) suffices to guarantee ITE identifiability. Theorem 1. Under the assumption of the presence of unobserved confounders u with known true propensity scores 𝜋 (𝑥, 𝑢), the ITE can be identified through two approaches: 1) 𝑡𝑦 (1 − 𝑡)𝑦 𝐼𝑇 𝐸 (𝑥) = E − 𝑥 (3) 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢) 2) If the nominal propensity score 𝜋˜ (𝑥) is known: 1 − 𝜋˜ (𝑥) 𝜋˜ (𝑥) 𝑦 𝑥, 𝑡 = 1 − E 𝑦 𝑥, 𝑡 = 0 . (4) 𝐼𝑇 𝐸 (𝑥) = E 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢) Theorem 1 provides a statistical perspective demonstrating that for identifying the ITE in the presence of unobserved confounders 𝑢, it suffices to understand the joint influence mechanism of 𝑢 with observed covariates 𝑥 on treatment assignment, rather than necessitating complete distributional characterization of 𝑢. Notably, a similar identification principle can be extended to the estimation of the Average Treatment Effect (ATE). Crucially, when true propensity scores 𝜋 (𝑥, 𝑢) are accessible, inverse weighting through them enables the construction of a pseudo-population that effectively blocks backdoor paths from unobserved confounders to the treatment. We next explain how to integrate this concept with weighting methods commonly used in machine learning.
3.3
Inverse Propensity Score
We now analyze selection bias in uplift modeling through a machine learning lens, focusing on how covariate shift between treatment groups impacts model generalization. Let 𝜏 (𝑥), 𝜇 0 (𝑥), 𝜇 1 (𝑥) denote the 𝐼𝑇 𝐸 (𝑥), E(𝑦|𝑥, 𝑡 = 0) and E(𝑦|𝑥, 𝑡 = 1) fitted by the model respectively, 𝑇 = {𝑖 : 𝑡𝑖 = 1}, 𝐶 = {𝑖 : 𝑡𝑖 = 0} are treatment group and control group. Selection bias in uplift modeling refers to the phenomenon where the model-learned 𝜇 0 (𝑥), 𝜇 1 (𝑥) are trained exclusively on the treatment and control group samples respectively, leading to poor generalization to the overall population distribution.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Analogous issues are prevalent in recommendation systems [31, 32] and CTR prediction [21, 44]. Formally, the bias arises from the nonequivalence between the following two loss functions: 𝑁
1 ∑︁ (𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 (0)) + 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 (1))), 𝑁 𝑖=1 ∑︁ 2 ∑︁ L 𝑓 𝑎𝑐𝑡 = ( 𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 (0)) + 𝑙 (𝜇 1 (𝑥𝑖 ), 𝑦𝑖 (1))), 𝑁 𝑖 ∈𝐶 𝑖 ∈𝑇 L𝑖𝑑𝑒𝑎𝑙 =
(5) (6)
where 𝑙 (·, ·) denotes the per-sample loss function, typically specified as mean squared error or cross-entropy. While our objective is to optimize the ideal loss L𝑖𝑑𝑒𝑎𝑙 , the fundamental unobservability of counterfactual outcomes necessitates optimizing the factual loss L 𝑓 𝑎𝑐𝑡 in practice, where E𝑡 (L 𝑓 𝑎𝑐𝑡 ) ≠ L𝑖𝑑𝑒𝑎𝑙 holds in general due to selection bias [32]. To tackle this issue, the established methodology applies IPS [12, 39] as a reweighting mechanism for the loss function, where propensity weights compensate for selection bias through importance sampling: 𝑁 1 ∑︁ (1 − 𝑡𝑖 )𝑙 (𝜇0 (𝑥𝑖 ), 𝑦𝑖 (0)) 𝑡𝑖 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 (1)) + L𝐼 𝑃𝑆 = . 𝑁 𝑖=1 1 − 𝜋˜ (𝑥𝑖 ) 𝜋˜ (𝑥𝑖 ) The statistical validity of the IPS method originates from its construction of an unbiased estimator for L𝑖𝑑𝑒𝑎𝑙 under the ignorability assumption, i.e., E𝑡 (L𝐼 𝑃𝑆 ) = L𝑖𝑑𝑒𝑎𝑙 . (7) However, in the presence of unobserved confounders, Theorem 2 below no longer holds universally due to the discrepancy between nominal and true propensity scores, i.e., 𝜋 (𝑥𝑖 , 𝑢𝑖 ) ≠ 𝜋˜ (𝑥𝑖 ). Unbiasedness can be preserved if and only if the weighting scheme employs the true propensity score [6]. Theorem 2. Under the assumption of the presence of unobserved confounders u, let L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 denote the IPS estimator using the true propensity score 𝜋 (𝑥𝑖 , 𝑢𝑖 ), formally expressed as: 𝑁 1 ∑︁ (1 − 𝑡𝑖 )𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 (0)) 𝑡𝑖 𝑙 (𝜇 1 (𝑥𝑖 ), 𝑦𝑖 (1)) L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 = + . 𝑁 𝑖=1 1 − 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝜋 (𝑥𝑖 , 𝑢𝑖 ) Then it is an unbiased estimator of L𝑖𝑑𝑒𝑎𝑙 .
4
Methodology
In this section, we elaborate on the proposed method consisting of two distinct methodological components: (1) CHAUN, a generalpurpose uplift network architecture for binary treatment settings, and (2) RA-IPS, a novel approach that enhances the robustness of IPS estimation under unobserved confounding. CHAUN predicts nominal propensity scores and uses them for weighting; this approach is fully valid in the absence of unobserved confounders, but must be combined with RA-IPS when such confounders are present.
4.1
CHAUN
The architecture of the proposed CHAUN is illustrated in Fig. 1. The architecture primarily comprises three components: a Shared Feature Embedding Layer, a Propensity Learner, and an Outcome Learner. In the subsequent subsections, we systematically elaborate on the design principles and implementation details of each module.
4.1.1 Shared Feature Embedding. For an instance with observed non-treatment features 𝑥𝑖 , assumed to consist of both continuous and sparse features, we process them separately. For continuous features, we apply a dimension-preserving projection with learnable affine transformations, while for sparse features, we initialize dedicated embedding tables and retrieve the corresponding embeddings via lookup operations based on their categorical values. The processed feature embeddings are concatenated into a unified vector 𝑥𝑟𝑒𝑝 , serving as the shared input to both the Propensity Learner and Outcome Learner. 4.1.2 Propensity Learner. Upon obtaining the instance representation 𝑥𝑟𝑒𝑝 , the Propensity Learner processes it through a multi-layer perceptron (MLP) architecture comprising stacked fully-connected layers interleaved with activation functions. The network culminates in a linear projection layer that reduces dimensionality to 1, followed by a sigmoid activation function to yield the final propensity score prediction. Let the MLP contain 𝑛 hidden layers, the Propensity Learner is formulated as: (𝑘 ) (𝑘 −1) 𝑥ℎ𝑖𝑑𝑑𝑒𝑛 = ReLU(Linear(𝑥ℎ𝑖𝑑𝑑𝑒𝑛 )) ∈ R𝑑 , 𝑘 = 1, · · · , 𝑛, (𝑛) 𝑝ˆ = Sigmoid(Linear(𝑥ℎ𝑖𝑑𝑑𝑒𝑛 )) ∈ (0, 1), (𝑘 ) where 𝑥ℎ𝑖𝑑𝑑𝑒𝑛 denotes the output of the 𝑘-th layer, with the initial (0) hidden state defined as 𝑥ℎ𝑖𝑑𝑑𝑒𝑛 = 𝑥𝑟𝑒𝑝 .
4.1.3 Outcome Learner. In contrast to conventional dual-head architectures that either independently predict potential outcomes of treatment and control groups from shared feature representations or impose regularization constraints to differentiate the heads, we propose an adaptive interaction mechanism that establishes dynamic coupling between the dual learners through attentionbased gating. The Outcome Learner initially processes 𝑥𝑟𝑒𝑝 through two MLPs to derive potential outcome representations under both treated and control conditions, formally expressed as: 𝑐 𝑥𝑟𝑒𝑝 = MLP(𝑥𝑟𝑒𝑝 ) ∈ R𝑑 , 𝑡 𝑥𝑟𝑒𝑝 = MLP(𝑥𝑟𝑒𝑝 ) ∈ R𝑑 ,
where the MLP is composed of multiple fully-connected layers interleaved with nonlinear activation functions. Rather than directly utilizing the two representations to generate corresponding outputs, we implement an attention-gated fusion mechanism to produce the final prediction through adaptive weighted combination. The fused representation is then processed by a lightweight output layer to generate the final prediction. For binary response variables, this process can be formally expressed as: exp(𝑞𝑐 𝑘 𝑐 )𝑣 𝑐 + exp(𝑞𝑐 𝑘 𝑡 )𝑣 𝑡 ∈ R𝑑 , exp(𝑞𝑐 𝑘 𝑐 ) + exp(𝑞𝑐 𝑘 𝑡 ) exp(𝑞𝑡 𝑘 𝑐 )𝑣 𝑐 + exp(𝑞𝑡 𝑘 𝑡 )𝑣 𝑡 𝑡 𝑥𝑎𝑡𝑡𝑛 = ∈ R𝑑 , exp(𝑞𝑡 𝑘 𝑐 ) + exp(𝑞𝑡 𝑘 𝑡 ) 𝑐 𝜇 0 = Sigmoid(Linear(𝑥𝑎𝑡𝑡𝑛 )) ∈ (0, 1),
𝑐 𝑥𝑎𝑡𝑡𝑛 =
𝑡 𝜇 1 = Sigmoid(Linear(𝑥𝑎𝑡𝑡𝑛 )) ∈ (0, 1), 𝑐 𝑡 where 𝑞𝑐 , 𝑘 𝑐 , 𝑣 𝑐 and 𝑞𝑡 , 𝑘 𝑡 , 𝑣 𝑡 are derived from 𝑥𝑟𝑒𝑝 and 𝑥𝑟𝑒𝑝 via learnable linear projections respectively.
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Propensity Learner
Shared Feature Embedding
Sigmoid MLP Continuous Projection Features Concatenation
𝒑
𝓛𝒑𝒓𝒐𝒑_𝒓𝒆𝒈
Propensity Score
𝓛𝑰𝑷𝑺
𝑲𝒕 𝒙𝒓𝒆𝒑 Shared Representation
MLP Treatment Rep. MLP
𝑲𝒄 𝑽𝒄
𝒙𝒕𝒓𝒆𝒑
Sigmoid
𝑽𝒕
𝑸
Control Rep. MLP
𝝁𝟏
𝒙𝒕𝒂𝒕𝒕𝒏
𝑲𝒕
Embedding Lookup
𝑽𝒕
𝑸𝒕
MLP Sparse Features
𝓛𝒑𝒓𝒐𝒑
𝝁𝟎
𝒙𝒄𝒂𝒕𝒕𝒏
𝑲𝒄 𝑽𝒄
𝒙𝒄𝒓𝒆𝒑
Sigmoid
Outcome Learner
Figure 1: The architecture overview of CHAUN. 4.1.4 Loss. To address sample selection bias, we incorporate IPS into the training loss computation, which requires accurate estimation of propensity scores. However, naively learning these scores through binary classification of treatment assignment labels may yield pathological solutions in which 𝑝ˆ (𝑥𝑖 ) = 𝑡𝑖 , a seemingly perfect but overlap-violating estimator. For stabilizing the inverse probability weights, a global regularization constraint is imposed on the propensity score estimation, grounded in the theoretical framework established by the following proposition. Proposition 1. Assuming the covariates x and treatment assignment t are governed by the joint distribution p(x, t) (whether the Ignorability assumption holds or not) with the overlap condition satisfied, the following holds: 1−𝑡 𝑡 =E = 1. (8) E 𝜋˜ (𝑥) 1 − 𝜋˜ (𝑥) Building upon Proposition 1, the supervised loss function for propensity score estimation comprises two key components: 𝑁
1 ∑︁ 𝑙 (𝑝ˆ (𝑥𝑖 ), 𝑡𝑖 ), 𝑁 𝑖=1 !2 !2 𝑁 𝑁 1 ∑︁ 𝑡𝑖 1 ∑︁ 1 − 𝑡𝑖 L𝑝𝑟𝑜𝑝_𝑟𝑒𝑔 = −1 + −1 , 𝑁 𝑖=1 𝑝ˆ (𝑥𝑖 ) 𝑁 𝑖=1 1 − 𝑝ˆ (𝑥𝑖 )
4.2
RA-IPS
As demonstrated in Section 3.3, the IPS estimator based on nominal propensity scores 𝜋˜ (𝑥) fails to achieve effective bias correction under the presence of unobserved confounders. In contrast, with access to true propensity scores 𝜋 (𝑥, 𝑢), debiasing can be achieved even without accurate information about the unobserved confounders. However, true propensity scores are generally unobservable in practice. To address this fundamental limitation, we develop a robust IPS estimation framework that incorporates sensitivity analysis principles, specifically designed to improve robustness to unobserved confounding via adversarial propensity. The nominal propensity score 𝜋˜ (𝑥) can be formally characterized as the conditional expectation of the true propensity score 𝜋 (𝑥, 𝑢) given the observed covariates 𝑥, ∫ 𝜋˜ (𝑥) =
𝜋 (𝑥, 𝑢)𝑝 (𝑢|𝑥)𝑑𝑢.
While nominal propensity scores cannot identify the true latent propensity scores, they constrain the feasible domain of the true propensity score.
L𝑝𝑟𝑜𝑝 =
Proposition 2. Suppose that 𝑎(𝑥) < 𝜋 (𝑥, 𝑢) < 1 − 𝑎(𝑥), then the following inequalities each hold with probability at most 𝜂, 1 1 1 − 2𝑎(𝑥) − |≥ √ , 𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥) 2𝑎(𝑥) 𝜋˜ (𝑥) 𝜂 1 1 1 − 2𝑎(𝑥) | − |≥ √ . 1 − 𝜋 (𝑥, 𝑢) 1 − 𝜋˜ (𝑥) 2(1 − 𝑎(𝑥)) 𝜋˜ (𝑥) 𝜂 |
where 𝑙 (·, ·) denotes the cross-entropy loss function. Computed propensity scores are gradient-detached and used to reweight the potential outcome prediction loss: 𝑁 1 ∑︁ 𝑡𝑖 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 ) (1 − 𝑡𝑖 )𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 ) L𝐼 𝑃𝑆 = + . 𝑁 𝑖=1 𝑝ˆ (𝑥𝑖 ) 1 − 𝑝ˆ (𝑥𝑖 ) The final optimization objective function is expressed as: L = L𝐼 𝑃𝑆 + 𝛼 L𝑝𝑟𝑜𝑝 + 𝛽L𝑝𝑟𝑜𝑝_𝑟𝑒𝑔 , where 𝛼 and 𝛽 are hyperparameters.
Under the assurance of Proposition 2, it is statistically valid to posit that the true propensity score resides within a neighborhood of the nominal propensity score with high probability. Therefore, a reasonable assumption is that the true propensity score is obtained by perturbing the nominal propensity score. We adhere to the theoretical framework of RD-IPS [6] by introducing a hyperparameter to regulate the impact of unobserved confounders on the logit-transformed propensity scores. Formally, the true propensity
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
score can be mathematically expressed as follows:
established in (11), we propose the following rigorously justified uncertainty set:
𝜋˜ (𝑥) = Sigmoid(𝑓 (𝑥)), 𝜋 (𝑥, 𝑢) = Sigmoid(𝑓 (𝑥) + 𝑔(𝑢)).
𝑁
Assuming 𝑔(𝑢) is bounded by |𝑔(𝑢)| < log(Γ), where Γ ≥ 1 is a hyperparameter corresponding to the strength of unobserved confounders. This leads to the constraints on the true weights: 𝑎(𝑥, 𝑡) ≤ 𝑤 (𝑥, 𝑢, 𝑡) ≤ 𝑏 (𝑥, 𝑡), where 𝑎(𝑥, 𝑡) = 1 + (𝑤˜ (𝑥, 𝑡) − 1) Γ1 , 𝑏 (𝑥, 𝑡) = 1 + (𝑤˜ (𝑥, 𝑡) − 1)Γ, and 𝑡 1−𝑡 𝑡 1−𝑡 𝑤˜ (𝑥, 𝑡) = 𝜋˜ (𝑥 ) + 1−𝜋˜ (𝑥 ) , 𝑤 (𝑥, 𝑢, 𝑡) = 𝜋 (𝑥,𝑢 ) + 1−𝜋 (𝑥,𝑢 ) denote the nominal weight and true weight respectively. Therefore, the RD-IPS method formulates the following uncertainty set:
W ∩ {𝑊 ∈ R+𝑁 :
1 ∑︁ (𝑤𝑖 − 𝑤˜ 𝑖 ) ≤ 𝜖𝑁 }, 𝑁 𝑖=1
(12)
where 𝜖𝑁 > 0, 𝜖𝑁 → 0 as 𝑁 → ∞. Therefore, in our adversarial learning framework, we incorporate such constraints as a regularization term to ensure feasible weight configurations during robust optimization. The final objective function to be optimized is formulated as follows: 𝑁 𝑁 1 ∑︁ 1 ∑︁ 𝑊ˆ = argmax 𝑤𝑖 𝑙 (𝜇𝑡𝑖 (𝑥𝑖 ), 𝑦𝑖 (𝑡𝑖 )) + 𝜆( (𝑤𝑖 − 𝑤˜ 𝑖 )) 2, 𝑁 𝑖=1 𝑊 ∈ W 𝑁 𝑖=1
W = {𝑊 ∈ R+𝑁 : 𝑎(𝑥𝑖 , 𝑡𝑖 ) ≤ 𝑤𝑖 ≤ 𝑏 (𝑥𝑖 , 𝑡𝑖 )}, 𝑁
and maximize the L𝐼 𝑃𝑆 objective over all possible configurations of true weights to optimize for the worst-case scenario, thereby enhancing the robustness. The optimization objective is:
L𝑅𝐴−𝐼 𝑃𝑆 =
1 ∑︁ 𝑤ˆ 𝑖 𝑙 (𝜇𝑡𝑖 (𝑥𝑖 ), 𝑦𝑖 (𝑡𝑖 )). 𝑁 𝑖=1
(13)
Proposition 3. Assuming the covariates 𝑥, unobserved confounders 𝑢 and treatment assignment 𝑡 are governed by the joint distribution 𝑝 (𝑥, 𝑢, 𝑡) with the overlap condition satisfied, the following holds: 𝑡 1−𝑡 E =E = 1. (10) 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢)
By empirically constraining the relationship between true weights and nominal weights in large-scale observational data, we establish a necessary condition for 𝑊 being the true weight configurations. Within a further constrained uncertainty set, we optimize for the worst-case scenario among all admissible 𝑊 , thereby ensuring that the model’s learned potential outcomes remain robust when the impact of unobserved confounders on treatment assignment is bounded within a specified range. To analyze the generalization error, we define a unified hypothesis class F consisting of joint prediction functions 𝑓𝜓 : X × {0, 1} → R, where each 𝑓𝜓 ∈ F is given by 𝑓 (𝑥, 𝑡) = 𝑡 · 𝜇1 (𝑥) + (1 − 𝑡) · 𝜇0 (𝑥), and the pair (𝜇0, 𝜇1 ) is parameterized by the model architecture. The empirical Rademacher complexity is then defined as # " 𝑁 1 ∑︁ 𝜎𝑖 𝑙 (𝑓 (𝑥𝑖 , 𝑡𝑖 ), 𝑦𝑖 ) , R (F ) = E𝜎∼{ −1,1} 𝑁 sup 𝑓 ∈ F 𝑁 𝑖=1
In large-scale observational datasets, a properly specified nominal propensity score and true propensity score should conform to the theoretical guarantees established in Proposition 1 and Proposition 3. According to the weak law of large numbers, we have
where 𝜎𝑖 are i.i.d. Rademacher random variables (P(𝜎𝑖 = ±1) = 12 ). Let 𝑙𝑖 , 𝑤𝑖 denote the observable loss and true weight for the 𝑖-th 𝑖 sample, i.e., 𝑙𝑖 = 𝑙 (𝜇𝑡𝑖 (𝑥𝑖 ), 𝑦𝑖 (𝑡𝑖 )), 𝑤𝑖 = 𝜋 (𝑥𝑡𝑖𝑖,𝑢𝑖 ) + 1−𝜋1−𝑡 (𝑥𝑖 ,𝑢𝑖 ) .
𝑁
1 ∑︁ L𝑅𝐷 −𝐼 𝑃𝑆 = max 𝑤𝑖 𝑙 (𝜇𝑡𝑖 (𝑥𝑖 ), 𝑦𝑖 (𝑡𝑖 )). 𝑊 ∈W 𝑁 𝑖=1
(9)
However, since the loss terms for all samples are non-negative, the maximization operation trivially leads to each weight attaining its upper bound within the uncertainty set. Such weight configurations remain structurally infeasible under real-world data distributions, a limitation arising from the absence of global constraints on weight combinations in W.
Theorem 3. (Generalization Bound) Suppose that the true weight configurations lie within the uncertainty set in (12), and 𝑙𝑖 ≤ 𝐶 1, 𝑤𝑖 ≤ 𝐶 2, 𝑤ˆ 𝑖 ≤ 𝐶 2, ∀𝑖. For any 𝑓𝜙 ∈ F , we have the following with probability at least 1 − 𝜂: √︂ log(2/𝜂) |L𝑅𝐴−𝐼 𝑃𝑆 (𝜙) − L𝑖𝑑𝑒𝑎𝑙 (𝜙)| ≤ 2𝐶 2 R (F ) + 𝐶 1𝐶 2 (1 + ). 2𝑁
𝑁 𝑝 1 ∑︁ lim 𝑤 (𝑥𝑖 , 𝑢𝑖 , 𝑡𝑖 ) → − 2, 𝑁 →∞ 𝑁 𝑖=1 𝑁 𝑝 1 ∑︁ 𝑤˜ (𝑥𝑖 , 𝑡𝑖 ) → − 2. 𝑁 →∞ 𝑁 𝑖=1
lim
Thus,
5
𝑁 𝑝 1 ∑︁ (𝑤 (𝑥𝑖 , 𝑢𝑖 , 𝑡𝑖 ) − 𝑤˜ (𝑥𝑖 , 𝑡𝑖 )) → − 0. 𝑁 →∞ 𝑁 𝑖=1
lim
Experiments
(11)
However, the solution for 𝑊 derived from (9) would result in
In this section, we conduct extensive experiments to answer the following research questions:
• RQ1: How does our CHAUN outperform different baseline uplift
! 𝑁 𝑁 1 ∑︁ 1 ∑︁ (𝑤 (𝑥𝑖 , 𝑢𝑖 , 𝑡𝑖 ) − 𝑤˜ (𝑥𝑖 , 𝑡𝑖 )) = (Γ − 1) 𝑤˜ (𝑥𝑖 , 𝑡𝑖 ) − 1 , 𝑁 𝑖=1 𝑁 𝑖=1
• RQ2: How does ablation or substitution of CHAUN’s core archi-
which fundamentally violates (11) if Γ > 1 and 𝑁 is large enough. Building upon the constraints on true weights and nominal weights
• RQ3: Does RA-IPS outperform the standard IPS estimator under
models? tectural components impact model performance? unobserved confounding?
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
5.1
Experimental Setup
5.1.1 Datasets. We conduct experiments on two widely used largescale binary-treatment uplift datasets, CRITEO-UPLIFT [5] and LAZADA [42] as well as a proprietary off-site advertising campaign dataset (Production). For CRITEO-UPLIFT, we randomly split it into training and evaluation sets with an 8/2 ratio and select visit as the target label. Since treatment assignment is randomized, exposure within the treatment group exhibits selection bias. Therefore, we use exposure rather than treatment assignment as the treatment label. Subsequently, we implement propensity score matching (PSM) to construct a pseudo-RCT evaluation set. In the Production dataset, treatment assignment is influenced by decisions made by external media platforms, making it infeasible to collect RCT data as an evaluation set. Therefore, we construct the evaluation set using propensity score matching (PSM). These external media decisions are driven by a set of features that cannot be fully obtained due to cost, risk control, permission constraints, and other practical limitations. We thus regard these features as unobserved confounders. Nevertheless, we were able to obtain these features for a subset of samples and used them in the PSM procedure to ensure that the resulting evaluation set is a fully pseudo-RCT. Key statistics of these three datasets used in our experiments are presented in Table 1. 5.1.2 Metrics. We employ three metrics commonly used in evaluating uplift modeling ranking performance, uplift score at first ℎ percentile (LIFT@ℎ, we set ℎ to 30), normalized area under the uplift curve (AUUC), normalized area under the QINI curve (QINI) and normalized principled uplift curve (PUC) [45] . The AUUC and QINI are calculated using functions from python package scikit-uplift. 5.1.3 Baselines. We compare our CHAUN with many neural network based uplift models including S-Learner [17], T-Learner [17], TARNet [34], CFRNet [34], DragonNet [35], CEVAE [19], FlexTENet [4], EUEN [15], DESCN [42], EFIN [18]. And we compare our proposed RA-IPS with RD-IPS [6]. 5.1.4 Implementation Details. All experiments are implemented using PyTorch 2.1.0 [24] and conducted on a single Tesla V100 GPU with 32GB memory. We use Adam [16] as the optimizer and a maximum iteration count of 50. To prevent overfitting, we implement an L2 weight decay of 1𝑒 −4 along with an early stopping strategy with a patience of 10. For each input batch, continuous features are first processed through batch normalization. Using the QINI metric as the criterion, we search for the optimal primary model and training hyperparameters within the parameter ranges described in Table 2. The code of experiments will be provided after the paper is accepted.
5.2
Overall Performance Comparison
We report the main experimental results of CHAUN and other baseline models on three real-world datasets in Table 3, and visualize the uplift curves in Fig. 2. According to these results, we have the following observations: 1) AUUC and QINI generally exhibit consistent monotonicity, whereas PUC demonstrates divergent patterns of similarity under specific conditions. This characteristic enables them to evaluate uplift models from complementary dimensions. 2) On two public datasets, S-Learner and T-Learner outperform some
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
methods with complex architectures or extra regularization. 3) TARNet and DragonNet achieve solid performance across all datasets, demonstrating that shared feature embedding layers, incorporating propensity score prediction as an auxiliary task, and properly calibrated weighting mechanisms can reliably improve model effectiveness. 4) Our proposed CHAUN demonstrates consistently superior performance across all three evaluated datasets. Specifically, it secures top-1 rankings on 9 out of 12 evaluation metrics and top-2 positions on 3 metrics. Notably, on two public benchmark datasets, CHAUN achieves state-of-the-art results across nearly all measurement dimensions. As evidenced by Figure 3, CHAUN’s uplift curve demonstrates strong monotonicity, achieving nearly the highest uplift in the top 20% of samples and the lowest uplift in the bottom 20% of samples, which validates its exceptional ranking capability for treatment effect stratification. On the Production dataset, only DESCN achieves comparable performance to CHAUN through domain-specific assumptions and customized loss designs, with gains being highly dataset-sensitive. Compared with the baseline DragonNet, CHAUN achieves QINI improvements ranging from 3% to 25.6% across three distinct datasets, demonstrating robust performance gains under varying data conditions. For the CRITEO-UPLIFT and LAZADA datasets, where treatment assignment is fully determined by observed covariates, we randomly selected subsets of features and employed permutation tests to verify that they simultaneously affected both treatment assignment and potential outcomes, thereby ensuring that they indeed constituted confounders. We then masked these features to simulate unobserved confounding scenarios. Subsequently, we evaluated the performance of IPS, RD-IPS, and RA-IPS using two architectures: 1) a network with shared embedding layers and separate heads for propensity score, treated outcome, and control outcome prediction (architecturally analogous to DragonNet but distinguished by a multi-layer propensity score head, we formally denote this framework as BaseNet); 2) the proposed CHAUN. As analytically demonstrated in Section 4.2, the weight configurations constructed in the RD-IPS method do not align with the worst-case scenario within the theoretical bounds of the confounding assumptions controlled by Γ. Instead, they amplify individual sample weights based on their nominal propensity scores. This approach effectively amplifies the scale of the IPS loss, which may empirically enhance practical optimization dynamics, but fundamentally deviates from the theoretically grounded adversarial robustness framework. Therefore, as demonstrated by the average metrics computed through repeated experiments in Table 4, RD-IPS exhibits no significant improvement over IPS across all benchmarks. In contrast, our proposed RA-IPS method rigorously identifies and optimizes against the worst-case scenarios induced by unobserved confounders affecting treatment assignment. Under most evaluation metrics, RA-IPS achieves significant performance improvements during causal effect estimation, attaining up to a 5.4% enhancement in QINI compared to IPS.
5.3
In-Depth Analysis
5.3.1 Ablation Study. We conduct ablation studies to validate the efficacy of the novel cross-head attention module within CHAUN. A comparative baseline involves removing the attention module, which corresponds to the aforementioned BaseNet architecture. In
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Table 1: The statistics of datasets. Dataset
CRITEO-UPLIFT Test
Train
Test
Train
Test
Size Features Treatment Ratio Average Conversion Rate Relative Average Uplift Average Uplift
11183673 12 3.07% 4.7% 1070.04% 37.88%
112788 12 50.00% 24.12% 180.02% 22.85%
926669 83 22.17% 2.00% 502.13% 4.72%
181669 83 52.12% 3.52% 11.11% 0.37%
21350813 93 53.78% 30.01% 109.27% 20.66%
11963551 93 49.99% 40.60% 2.86% 1.15%
𝑑 𝑛 𝑏𝑠 𝑙𝑟 𝛼, 𝛽, 𝜆
{26, 27, 28, 29 } {1, 2, 3, 4} {28, 29, 210, 211 } {5𝑒 −3, 1𝑒 −3, 5𝑒 −4, 1𝑒 −4 } {10, 1, 1𝑒 −1, 1𝑒 −2 }
hidden dimensions number of layers batch size learning rate weight of multi-task loss
0.4 0.3 0.2 0.1 10% 20% 30% 40% 50% 60% 70% 80% 90%100%
TARNet FlexTENet
0.8
LAZADA
n Weig ht A0, 1 Predic
0.4
0.2
0.0
0.2
0.4
0.6
Attention Weigh 0.8 t A1, 1
1.0
Attenti o
0.0
1.0 0.8 0.6 0.4 0.2 0.0
(a) Predicted outcome.
0.14
1
0 (ITE)
0.2
0.4
0.12 0.10
Predicted ITE
Functionality
0.0150 0.0125 0.0100 0.0075 0.0050 0.0025 0.0000 0.0025
0.16
ted Outcome
Range
CRITEO-UPLIFT
1 (treated outcome) 0 (control outcome)
0.6
Name
CHAUN EUEN
Production
Train
Table 2: Main model and training hyperparameters and their value range.
0.0
LAZADA
Split
0.08 0.06 0.04 0.02 0.00 0.0
0.6
0.8
Attention Weight A1, 1
1.0
(b) Predicted ITE.
Figure 3: Scatter plots illustrating the relationship between predicted outcomes/predicted ITEs and attention weights. A1,1 denotes the attention weights where treatment group representations serve as both queries and keys, while A0,1 represent the weights where control group representations act as queries and treatment group representations as keys. 10% 20% 30% 40% 50% 60% 70% 80% 90%100%
CFRNet DESCN
DragonNet EFIN
CEVAE
Figure 2: The uplift curve on CRITEO-UPLIFT and LAZADA datasets. Samples are sorted in descending order based on their predicted uplift scores and divided into 10 decile groups. Then we calculate the actual uplift within each group. An ideal uplift curve exhibits strict monotonicity (steadily decreasing uplift across ranked deciles) and high discriminatory power, maximizing treatment benefits in top-ranked subgroups while minimizing harm in low-response groups.
addition, we consider another method analogous to the weighted summation in the attention module as a baseline, Multi-gate Mixtureof-Experts (MMoE) [20]. MMoE utilizes multiple expert networks and trains task-specific gating functions to compute weighted combinations of their outputs. The key distinction of CHAUN’s attention mechanism lies in its explicit exploitation of inter-group similarity between the treatment and control groups, where attention weights are derived through inner product similarity computations. To ensure a fair comparison, we configure MMoE with 2 experts for potential outcome prediction and calibrate the number of layers across models to maintain comparable parameter scales. The
results of the ablation experiments are presented in Table 5. As evidenced by the results, MMoE consistently demonstrates superior performance compared to BaseNet across the majority of experimental scenarios. Notably, CHAUN achieves the best performance across all three datasets. On the CRITEO-UPLIFT dataset with lowdimensional features and pronounced treatment effects, CHAUN performs comparably to MMoE. However, on datasets featuring high-dimensional and complex feature spaces (LAZADA and Production), CHAUN demonstrates significantly stronger advantages, achieving QINI coefficient improvements of up to 19.8%. 5.3.2 Visualization of Prediction. To clarify how the core attention module of CHAUN leverages inter-group similarity to enhance the discriminative power of predicted ITEs, we rigorously investigate the relationship between predicted outcomes, ITEs and computed attention weights on the Production dataset, with results visualized in Fig. 3. As shown in Fig. 3a, the attention weights computed when using treatment group representations versus control group representations as queries (on x-axis and y-axis, respectively) exhibit close alignment. Furthermore, both predicted potential outcomes 𝜇0 and 𝜇 1 increase monotonically with higher attention weights assigned to treatment group representations, i.e., greater focus on treatment group representations leads to larger predicted potential outcomes. Combined with Fig. 3b, we observe that when attention weights heavily focus on control group representations, both 𝜇 0
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Table 3: Overall performance comparison between our CHAUN and other baseline models. "Random" refers to the calculated expected metric value after random scoring and ranking. All experiments are conducted with 5 random seeds, with metric averages computed for robustness evaluation. Higher LIFT@30, AUUC, QINI, PUC indicate superior uplift modeling capability. The best results of each benchmark are in bold and the second best are underlined. CRITEO-UPLIFT
Method Random S-Learner T-Learner TARNet CFRNet DragonNet CEVAE EUEN FlexTENet DESCN EFIN CHAUN
LAZADA
Production
LIFT@30
AUUC
QINI
PUC
LIFT@30
AUUC
QINI
PUC
LIFT@30
AUUC
QINI
PUC
0.2285 0.3820 0.3844 0.3865 0.3683 0.3846 0.3605 0.3665 0.3648 0.3841 0.3516 0.3904
0.0000 0.1541 0.1551 0.1701 0.1463 0.1595 0.1308 0.1349 0.1418 0.1516 0.1372 0.1620
0.0000 0.1848 0.1859 0.1890 0.1754 0.1898 0.1585 0.1616 0.1708 0.1820 0.1651 0.1955
0.0000 0.1585 0.1591 0.1625 0.1505 0.1631 0.1392 0.1395 0.1465 0.1560 0.1411 0.1664
0.0037 0.0080 0.0078 0.0084 0.0064 0.0082 0.0072 0.0081 0.0080 0.0076 0.0077 0.0087
0.0000 0.0033 0.0024 0.0042 0.0025 0.0041 0.0029 0.0031 0.0036 0.0028 0.0032 0.0044
0.0000 0.0236 0.0172 0.0302 0.0176 0.0250 0.0200 0.0223 0.0253 0.0190 0.0230 0.0314
0.0000 0.0060 0.0099 0.0064 0.0052 0.0067 0.0032 0.0030 0.0054 0.0031 0.0059 0.0091
0.0114 0.0188 0.0181 0.0234 0.0236 0.0228 0.0256 0.0245 0.0166 0.0286 0.0221 0.0261
0.0000 0.0083 0.0099 0.0127 0.0131 0.0118 0.0129 0.0111 0.0133 0.0137 0.0109 0.0144
0.0000 0.0074 0.0086 0.0113 0.0116 0.0106 0.0118 0.0097 0.0117 0.0122 0.0092 0.0130
0.0000 0.0073 0.0131 0.0120 0.0122 0.0113 0.0124 0.0101 0.0091 0.0150 0.0099 0.0144
Table 4: Performance comparison of IPS, RD-IPS, and RA-IPS methods. All experiments were conducted with 5 random seeds, with metric averages computed for robustness evaluation. Best results of each benchmark are in bold. CRITEO-UPLIFT-Masked
Method
LAZADA-Masked
Production
LIFT@30
AUUC
QINI
PUC
LIFT@30
AUUC
QINI
PUC
LIFT@30
AUUC
QINI
PUC
BaseNet+IPS BaseNet+RD-IPS BaseNet+RA-IPS
0.3564 0.3568 0.3580
0.1332 0.1335 0.1356
0.1592 0.1597 0.1621
0.1376 0.1378 0.1401
0.0081 0.0082 0.0083
0.0030 0.0030 0.0031
0.0215 0.0215 0.0217
0.0044 0.0044 0.0044
0.0230 0.0227 0.0229
0.0119 0.0117 0.0121
0.0107 0.0106 0.0109
0.0113 0.0110 0.0114
CHAUN+IPS CHAUN+RD-IPS CHAUN+RA-IPS
0.3612 0.3619 0.3637
0.1390 0.1398 0.1430
0.1658 0.1668 0.1704
0.1432 0.1444 0.1474
0.0083 0.0083 0.0084
0.0031 0.0031 0.0033
0.0221 0.0219 0.0230
0.0047 0.0044 0.0045
0.0257 0.0256 0.0263
0.0144 0.0142 0.0151
0.0130 0.0129 0.0136
0.0144 0.0141 0.0156
Table 5: Ablation study of CHAUN. BaseNet denotes a simplified variant of CHAUN with the attention module removed, while MMoE replaces the attention weighting mechanism with task-specific gating operations. All experiments were conducted with 5 random seeds, with metric averages computed for robustness evaluation. The best results of each benchmark are in bold. CRITEO-UPLIFT
Method
Production
QINI
PUC
LIFT@30 AUUC
QINI
PUC
LIFT@30 AUUC
QINI
PUC
0.3861 0.3891 0.3904
0.1915 0.1940 0.1955
0.1641 0.1664 0.1664
0.0083 0.0084 0.0087
0.0236 0.0262 0.0314
0.0065 0.0057 0.0091
0.0230 0.0249 0.0261
0.0119 0.0124 0.0130
0.0113 0.0131 0.0144
and 𝜇 1 collapse near zero, leading to nearly all predicted ITEs being trivial. Conversely, when attention predominantly emphasizes treatment group representations, 𝜇 0 and 𝜇 1 exhibit large values with divergent ITEs, resembling predictions from independent model heads without inter-group interaction. Notably, the most significant ITE estimates emerge within an intermediate range of attention weights, where the model balances inter-group feature sharing and treatment-specific heterogeneity, thereby amplifying contrastive signals between counterfactual outcomes. Based on these observations, we conclude that representations from different treatment groups are mapped to shared query spaces but distinct key and value spaces, preserving treatment-specific patterns. The attention weights dynamically calibrate inter-group interactions, enhancing ITE discriminability through amplified counterfactual contrasts.
0.0041 0.0037 0.0044
Production
Mean of QINI Std of QINI
0.0240
Mean of QINI Std of QINI
0.0150
0.0230 0.0220 0.0210 0.0200 0.0190
0.0107 0.0137 0.0144
LAZADA-Masked
0.0250
QINI
0.1597 0.1616 0.1620
QINI
BaseNet (removed attention) MMoE (replaced attention) CHAUN
LAZADA
LIFT@30 AUUC
0.0140 0.0130 0.0120
1
1.01 1.05 1.1 Value of
1.2
1.5
1
1.01 1.05 1.1 Value of
1.2
1.5
Figure 4: Performance (QINI) of RA-IPS as the hyperparameter Γ varies. The mean and standard deviation are computed from results obtained with five different random seeds.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
5.3.3 Hyperparameter Sensitivity. The RA-IPS method requires selecting an appropriate hyperparameter Γ corresponding to the tolerable strength of unobserved confounders. We investigate how the performance of RA-IPS varies under different values of Γ, and present the results on LAZADA-Masked and Production datasets in Fig. 4. Γ = 1 corresponds to the standard IPS. It can be observed that when Γ varies within a small range, RA-IPS consistently achieves stable improvements over IPS. However, as the value of Γ gradually increases, the performance begins to fluctuate noticeably, making it no longer guaranteed to outperform IPS, while also exhibiting a larger standard deviation. This suggests that when applying RA-IPS, if the strength of unobserved confounders is completely unknown, it is advisable to start with a relatively small Γ.
Trovato et al.
Proof. 1) By the Law of Total Expectation, ∫ E(𝑦(1)|𝑥) = E(𝑦 (1)|𝑥, 𝑢)𝑝 (𝑢|𝑥) 𝑑𝑢 ∫ 1 = 𝑡E(𝑦 (1)|𝑥, 𝑢)𝑝 (𝑢, 𝑡 |𝑥) 𝑑𝑢 𝑑𝑡 𝜋 (𝑥, 𝑢) 𝑡𝑦(1) 𝑡𝑦 =E 𝑥 =E 𝑥 , 𝜋 (𝑥, 𝑢) 𝜋 (𝑥, 𝑢) analogously, we obtain: E(𝑦 (0)|𝑥) = E thus,
6
𝐼𝑇 𝐸 (𝑥) = E
Conclusions
In this paper, we address the lack of flexible integration of intergroup correlations in uplift modeling by proposing CHAUN, a concise yet effective framework that dynamically computes inter-group attention scores to model such associations. Furthermore, for uplift scenarios with unobserved confounders, we establish that knowing the true propensity score for each unit serves as a sufficient condition to eliminate confounding bias. For practical settings where true propensity scores are unavailable, we introduce RA-IPS, which enhances robustness against unobserved confounders through adversarial learning over plausible propensity score spaces. Extensive comparative and ablation experiments demonstrate CHAUN’s effectiveness as a general-purpose uplift model and RA-IPS’s consistent improvements over conventional IPS methods in the presence of unobserved confounders.
Acknowledgements This work was supported by the NSFC Project (No.62576346), the MOE Project of Key Research Institute of Humanities and Social Sciences (22JJD110001), the fundamental research funds for the central universities, and the research funds of Renmin University of China (24XNKJ13), and Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing.
A
Theoretical Proof
(Theorem 1). Under the assumption of the presence of unobserved confounders u with known true propensity scores 𝜋 (𝑥, 𝑢), the ITE can be identified through two approaches: 1) 𝐼𝑇 𝐸 (𝑥) = E
𝑡𝑦 (1 − 𝑡)𝑦 − 𝑥 . 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢)
2) If the nominal propensity score 𝜋˜ (𝑥) is known:
(1 − 𝑡)𝑦 𝑥 , 1 − 𝜋 (𝑥, 𝑢)
𝑡𝑦 (1 − 𝑡)𝑦 𝑥 . − 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢)
∫ 2) E(𝑦 (1)|𝑥) =
E(𝑦(1)|𝑥, 𝑢)𝑝 (𝑢|𝑥)𝑑𝑢 ∫
𝜋˜ (𝑥) 𝑝 (𝑢|𝑥, 𝑡 = 1)𝑑𝑢 𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥) =E 𝑦 𝑥, 𝑡 = 1 , 𝜋 (𝑥, 𝑢)
=
E(𝑦|𝑥, 𝑢, 𝑡 = 1)
analogously, we obtain: E(𝑦 (0)|𝑥) = E
1 − 𝜋˜ (𝑥) 𝑦 𝑥, 𝑡 = 0 , 1 − 𝜋 (𝑥, 𝑢)
thus, 𝐼𝑇 𝐸 (𝑥) = E
𝜋˜ (𝑥) 1 − 𝜋˜ (𝑥) 𝑦 𝑥, 𝑡 = 1 − E 𝑦 𝑥, 𝑡 = 0 . □ 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢)
(Theorem 2). Under the assumption of the presence of unobserved confounders u, let L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 denote the IPS estimator using the true propensity score 𝜋 (𝑥𝑖 , 𝑢𝑖 ), formally expressed as 𝑁 1 ∑︁ (1 − 𝑡𝑖 )𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 (0)) 𝑡𝑖 𝑙 (𝜇 1 (𝑥𝑖 ), 𝑦𝑖 (1)) L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 = + . 𝑁 𝑖=1 1 − 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝜋 (𝑥𝑖 , 𝑢𝑖 ) Then it is an unbiased estimator of L𝑖𝑑𝑒𝑎𝑙 . Proof. For each sample, (1 − 𝑡𝑖 )𝑙 (𝜇0 (𝑥𝑖 ), 𝑦𝑖 (0)) 𝑡𝑖 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 (1)) E𝑡𝑖 + 1 − 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝑙 (𝜇 1 (𝑥𝑖 ), 𝑦𝑖 (1)) 𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 (0)) + 𝜋 (𝑥𝑖 , 𝑢𝑖 ) = (1 − 𝜋 (𝑥𝑖 , 𝑢𝑖 )) 1 − 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝜋 (𝑥𝑖 , 𝑢𝑖 ) = 𝑙 (𝜇 0 (𝑥𝑖 ), 𝑦𝑖 (0)) + 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 (1)) Since treatment assignments are independent across all samples, we have E𝑡 (L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 ) 𝑁 (1 − 𝑡𝑖 )𝑙 (𝜇0 (𝑥𝑖 ), 𝑦𝑖 (0)) 𝑡𝑖 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 (1)) 1 ∑︁ = E𝑡𝑖 + 𝑁 𝑖=1 1 − 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝜋 (𝑥𝑖 , 𝑢𝑖 ) 𝑁
𝐼𝑇 𝐸 (𝑥) = E
𝜋˜ (𝑥) 1 − 𝜋˜ (𝑥) 𝑦 𝑥, 𝑡 = 1 − E 𝑦 𝑥, 𝑡 = 0 . 𝜋 (𝑥, 𝑢) 1 − 𝜋 (𝑥, 𝑢)
=
1 ∑︁ (𝑙 (𝜇0 (𝑥𝑖 ), 𝑦𝑖 (0)) + 𝑙 (𝜇1 (𝑥𝑖 ), 𝑦𝑖 (1))) = L𝑖𝑑𝑒𝑎𝑙 . 𝑁 𝑖=1
□
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
(Proposition 2). For a fixed 𝑥, suppose that 𝑎(𝑥) < 𝜋 (𝑥, 𝑢) < 1−𝑎(𝑥), ∀𝑢, then the following inequalities each hold with probability at most 𝜂, 1 1 1 − 2𝑎(𝑥) − |≥ √ , 𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥) 2𝑎(𝑥) 𝜋˜ (𝑥) 𝜂 1 1 1 − 2𝑎(𝑥) | − |≥ √ . 1 − 𝜋 (𝑥, 𝑢) 1 − 𝜋˜ (𝑥) 2(1 − 𝑎(𝑥)) 𝜋˜ (𝑥) 𝜂 |
Proof. Since 𝑎(𝑥) < 𝜋 (𝑥, 𝑢) < 1 − 𝑎(𝑥), we have
bounded difference property holds for the supremum Φ(D). By McDiarmid’s Inequality [23], with probability at least 1 − 𝜂, √︂ log(2/𝜂) . Φ(D) ≤ E D [Φ(D)] + 𝑀 2𝑁 By Symmetrization Lemma [33], the expected uniform deviation is bounded by twice the Rademacher complexity of the weighted loss class. Since the nominal weights are bounded by 𝐶 2 , the contraction property of Rademacher complexity implies that this quantity is at most 2𝐶 2 R (F ). Hence,
(𝜋 (𝑥, 𝑢) − 𝑎(𝑥))(𝜋 (𝑥, 𝑢) − (1 − 𝑎(𝑥))) ≤ 0 ⇔ 𝜋 (𝑥, 𝑢) 2 − 𝜋 (𝑥, 𝑢) + 𝑎(𝑥)(1 − 𝑎(𝑥)) ≤ 0
E D [Φ(D)] ≤ 2𝐶 2 R (F ). (14)
Taking the conditional expectation of (14) with respect to 𝑥, we obtain: E𝑢 |𝑥 (𝜋 (𝑥, 𝑢) 2 ) ≤ 𝜋˜ (𝑥) − 𝑎(𝑥)(1 − 𝑎(𝑥)) ⇔
Var𝑢 |𝑥 (𝜋 (𝑥, 𝑢)) ≤ −𝜋˜ (𝑥) 2 + 𝜋˜ (𝑥) − 𝑎(𝑥)(1 − 𝑎(𝑥)) (15)
The right-hand side of (15) attains its maximum (𝑎(𝑥) − 12 ) 2 at 𝜋˜ (𝑥) = 12 . Notice that
⇔
1 1 | − | ≥𝜖 𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥) |𝜋 (𝑥, 𝑢) − 𝜋˜ (𝑥)| ≥ 𝜖𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥)
⇒
|𝜋 (𝑥, 𝑢) − 𝜋˜ (𝑥)| ≥ 𝜖𝑎(𝑥) 𝜋˜ (𝑥)
By Chebyshev’s Inequality, we have Pr(|
1 1 − | ≥ 𝜖) ≤ Pr(|𝜋 (𝑥, 𝑢) − 𝜋˜ (𝑥)| ≥ 𝜖𝑎(𝑥) 𝜋˜ (𝑥)) 𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥) 1 2 Var𝑢 |𝑥 (𝜋 (𝑥, 𝑢)) 2 − 𝑎(𝑥) ≤ ≤ . (𝜖𝑎(𝑥) 𝜋˜ (𝑥))) 2 𝜖𝑎(𝑥) 𝜋˜ (𝑥)) 1 −𝑎 (𝑥 )
Letting 𝜂 = ( 𝜖𝑎2(𝑥 ) 𝜋˜ (𝑥 ) ) ) 2 , then we have with probability at most 𝜂, |
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
1 1 − 2𝑎(𝑥) 1 |≥ − √ . 𝜋 (𝑥, 𝑢) 𝜋˜ (𝑥) 2𝑎(𝑥) 𝜋˜ (𝑥) 𝜂
The other inequality can be analogously proved following the same methodology. □ (Theorem 3). Suppose that the true weight configurations lie within the uncertainty set in (12), and 𝑙𝑖 < 𝐶 1, 𝑤𝑖 < 𝐶 2 . For any 𝑓𝜙 ∈ F , we have that with probabilty at least 1 − 𝜂, √︂ log(2/𝜂) ) |L𝑅𝐴−𝐼 𝑃𝑆 (𝜙) − L𝑖𝑑𝑒𝑎𝑙 (𝜙)| ≤ 2𝐶 2 R (F ) + 𝐶 1𝐶 2 (1 + 2𝑁 Proof. |L𝑅𝐴−𝐼 𝑃𝑆 (𝜙)−L𝑖𝑑𝑒𝑎𝑙 (𝜙)| ≤ |L𝑅𝐴−𝐼 𝑃𝑆 (𝜙)−L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 (𝜙)| + |L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 (𝜙) − L𝑖𝑑𝑒𝑎𝑙 (𝜙)|. Next, we derive upper bounds for each of these two terms individually. Since the true weighting combination lies within our predefined uncertainty set, for the first term we have |L𝑅𝐴−𝐼 𝑃𝑆 (𝜙) − L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 (𝜙)| = L𝑅𝐴−𝐼 𝑃𝑆 (𝜙) − Í𝑁 L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 (𝜙) = 𝑁1 𝑖=1 (𝑤ˆ 𝑖 − 𝑤𝑖 )𝑙𝑖 ≤ 𝐶 1𝐶 2 We now bound the second term |Ltrue-IPS (𝜙) − Lideal (𝜙)|. By our assumption, the persample loss and the true weights are bounded: 𝑙𝑖 ≤ 𝐶 1 and 𝑤𝑖 ≤ 𝐶 2 . Therefore, each weighted loss term 𝑤𝑖 𝑙𝑖 is bounded by 𝑀 = 𝐶 1𝐶 2 . Consider the random variable Φ(D) = sup 𝑓𝜓 ∈ F |L𝑖𝑑𝑒𝑎𝑙 (𝑓𝜓 ) − L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 (𝑓𝜓 )|. Replacing any single data point (𝑥𝑖 , 𝑡𝑖 , 𝑦𝑖 ) in the sample set D can change L𝑡𝑟𝑢𝑒 −𝐼 𝑃𝑆 (𝑓𝜓 ) by at most 𝑀/𝑁 , and this
Combining these results, we conclude that with probability at least 1 − 𝜂, √︂ log(2/𝜂) |L𝑅𝐴−𝐼 𝑃𝑆 (𝜙) − L𝑖𝑑𝑒𝑎𝑙 (𝜙)| ≤ 2𝐶 2 R (F ) + 𝐶 1𝐶 2 (1 + ). 2𝑁 □
References [1] Heejung Bang and James M. Robins. 2005. Doubly Robust Estimation in Missing Data and Causal Inference Models. Biometrics 61 (2005). [2] Min Cheng, Xinru Liao, Quanlian Liu, Bin Ma, Jian Xu, and Bo Zheng. 2022. Learning Disentangled Representations for Counterfactual Regression via Mutual Information Minimization. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (2022). [3] Bénédicte Colnet, Imke Mayer, Guanhua Chen, Awa Dieng, Ruohong Li, Gaël Varoquaux, Jean-Philippe Vert, Julie Josse, and Shu Yang. 2024. Causal Inference Methods for Combining Randomized Trials and Observational Studies: A Review. Statist. Sci. 39, 1 (2024), 165 – 191. doi:10.1214/23-STS889 [4] Alicia Curth and Mihaela van der Schaar. 2021. On inductive biases for heterogeneous treatment effect estimation (NIPS ’21). Curran Associates Inc., Red Hook, NY, USA, Article 1215, 12 pages. [5] Eustache Diemert, Artem Betlei, Christophe Renaudin, Massih-Reza Amini, Théophane Gregoir, and Thibaud Rahier. 2021. A Large Scale Benchmark for Individual Treatment Effect Prediction and Uplift Modeling. arXiv:2111.10106 [stat.ML] [6] Sihao Ding, Peng Wu, Fuli Feng, Yitong Wang, Xiangnan He, Yong Liao, and Yongdong Zhang. 2022. Addressing Unmeasured Confounder for Recommendation with Sensitivity Analysis. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Washington DC, USA) (KDD ’22). Association for Computing Machinery, New York, NY, USA, 305–315. [7] Shuyang Du, James Lee, and Farzin Ghaffarizadeh. 2019. Improve User Retention with Causal Learning. In CD@KDD. [8] Jason Hartford, Greg Lewis, Kevin Leyton-Brown, and Matt Taddy. 2017. Deep IV: a flexible approach for counterfactual prediction (ICML’17). JMLR.org, 1414–1423. [9] Negar Hassanpour and Russell Greiner. 2020. Learning Disentangled Representations for CounterFactual Regression. In International Conference on Learning Representations. [10] Miguel A. Hernan. 2024. Causal Inference: What If. Taylor & Francis, Boca Raton. [11] Yinqiu Huang, Shuli Wang, Min Gao, Xue Wei, Changhao Li, Chuan Luo, Yinhua Zhu, Xiong Xiao, and Yi Luo. 2024. Entire Chain Uplift Modeling with ContextEnhanced Learning for Intelligent Marketing. In Companion Proceedings of the ACM Web Conference 2024 (Singapore, Singapore) (WWW ’24). Association for Computing Machinery, New York, NY, USA, 226–234. [12] Guido W. Imbens and Donald B. Rubin. 2015. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press. [13] Fredrik D. Johansson, Uri Shalit, and David Sontag. 2016. Learning representations for counterfactual inference. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 (New York, NY, USA) (ICML’16). JMLR.org, 3020–3029. [14] Nathan Kallus, Aahlad Manas Puli, and Uri Shalit. 2018. Removing hidden confounding by experimental grounding. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18). Curran Associates Inc., Red Hook, NY, USA, 10911–10920. [15] Wenwei Ke, Chuanren Liu, Xiangfu Shi, Yiqiao Dai, Philip S. Yu, and Xiaoqiang Zhu. 2021. Addressing Exposure Bias in Uplift Modeling for Large-scale Online Advertising. In 2021 IEEE International Conference on Data Mining (ICDM). 1156– 1161. [16] Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
[17] Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu. 2017. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the National Academy of Sciences of the United States of America 116 (2017), 4156 – 4165. [18] Dugang Liu, Xing Tang, Han Gao, Fuyuan Lyu, and Xiuqiang He. 2023. Explicit Feature Interaction-aware Uplift Network for Online Marketing. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4507–4515. [19] Christos Louizos, Uri Shalit, Joris Mooij, David Sontag, Richard Zemel, and Max Welling. 2017. Causal Effect Inference with Deep Latent-Variable Models. arXiv:1705.08821 [stat.ML] [20] Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018. Modeling Task Relationships in Multi-task Learning with Multi-gate Mixtureof-Experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (London, United Kingdom) (KDD ’18). Association for Computing Machinery, 1930–1939. [21] Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu, Xiaoqiang Zhu, and Kun Gai. 2018. Entire Space Multi-Task Model: An Effective Approach for Estimating Post-Click Conversion Rate (SIGIR ’18). Association for Computing Machinery, New York, NY, USA, 1137–1140. [22] Wang Miao, Zhi Geng, and Eric J. Tchetgen Tchetgen. 2016. Identifying Causal Effects With Proxy Variables of an Unmeasured Confounder. Biometrika 1054 (2016), 987–993. [23] Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. 2012. Foundations of Machine Learning. The MIT Press. [24] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024–8035. [25] Judea Pearl. 2009. Causality: Models, Reasoning and Inference. Cambridge University Press, USA. [26] Judea Pearl. 2022. Detecting Latent Heterogeneity. Association for Computing Machinery, New York, NY, USA. [27] James M. Robins, Miguel A. Hernán, and Babette A. Brumback. 2000. Marginal Structural Models and Causal Inference in Epidemiology. Epidemiology 11 (2000), 550–560. [28] Paul R. Rosenbaum and Donald B. Rubin. 1983. The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika 70 (1983). [29] Donald B Rubin. 2005. Causal Inference Using Potential Outcomes. J. Amer. Statist. Assoc. 100, 469 (2005), 322–331. [30] Kara Rudolph, Nicholas Williams, and Ivan Diaz. 2024. Using instrumental variables to address unmeasured confounding in causal mediation analysis. Biometrics 80 (01 2024). [31] Yuta Saito, Suguru Yaginuma, Yuta Nishino, Hayato Sakata, and Kazuhide Nakata. 2020. Unbiased Recommender Learning from Missing-Not-At-Random Implicit Feedback. In Proceedings of the 13th International Conference on Web Search and Data Mining (Houston, TX, USA) (WSDM ’20). Association for Computing Machinery, New York, NY, USA, 501–509. [32] Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims. 2016. Recommendations as treatments: debiasing learning and evaluation (ICML’16). JMLR.org, 1670–1679. [33] Shai Shalev-Shwartz and Shai Ben-David. 2013. Understanding Machine Learning: From Theory to Algorithms. Understanding Machine Learning: From Theory to Algorithms (01 2013). doi:10.1017/CBO9781107298019 [34] Uri Shalit, Fredrik D. Johansson, and David A. Sontag. 2016. Estimating individual treatment effect: generalization bounds and algorithms. In International Conference on Machine Learning. [35] Claudia Shi, David M. Blei, and Victor Veitch. 2019. Adapting neural networks for the estimation of treatment effects. Curran Associates Inc., Red Hook, NY, USA. [36] Wei Sun, Pengyuan Wang, Dawei Yin, Jian Yang, and Yi Chang. 2015. Causal inference via sparse additive models with application to online advertising (AAAI’15). 297–303. [37] Zexu Sun, Qiyu Han, Minqin Zhu, Hao Gong, Dugang Liu, and Chen Ma. 2025. Robust Uplift Modeling with Large-Scale Contexts for Real-time Marketing (KDD ’25). Association for Computing Machinery, New York, NY, USA, 1325–1336. [38] Eric J Tchetgen Tchetgen, Andrew Ying, Yifan Cui, Xu Shi, and Wang Miao. 2020. An Introduction to Proximal Causal Learning. arXiv:2009.10982 [stat.ME] [39] Steven K. Thompson. 2012. Sampling. Wiley, Hoboken, N.J. [40] Anpeng Wu, Kun Kuang, Bo Li, and Fei Wu. 2022. Instrumental Variable Regression with Confounder Balancing. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.). PMLR, 24056–24075.
Trovato et al.
[41] Zhiheng Zhang, Quanyu Dai, Xu Chen, Zhenhua Dong, and Ruiming Tang. 2023. Robust Causal Inference for Recommender System to Overcome Noisy Confounders. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for Computing Machinery, New York, NY, USA, 2349–2353. [42] Kailiang Zhong, Fengtong Xiao, Yan Ren, Yaorong Liang, Wenqing Yao, Xiaofeng Yang, and Ling Cen. 2022. DESCN: Deep Entire Space Cross Networks for Individual Treatment Effect Estimation. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4612–4620. [43] Dingyuan Zhu, Daixin Wang, Zhiqiang Zhang, Kun Kuang, Yan Zhang, Yulin Kang, and Jun Zhou. 2023. Graph Neural Network with Two Uplift Estimators for Label-Scarcity Individual Uplift Modeling. In Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (WWW ’23). Association for Computing Machinery, New York, NY, USA, 395–405. [44] Feng Zhu, Mingjie Zhong, Xinxing Yang, Longfei Li, Lu Yu, Tiehua Zhang, Jun Zhou, Chaochao Chen, Fei Wu, Guanfeng Liu, and Yan Wang. 2023. DCMT: A Direct Entire-Space Causal Multi-Task Framework for Post-Click Conversion Estimation. 2023 IEEE 39th International Conference on Data Engineering (ICDE) (2023), 3113–3125. [45] Minqin Zhu, Zexu Sun, Ruoxuan Xiong, Anpeng Wu, Baohong Li, Caizhi Tang, Jun Zhou, Fei Wu, and Kun Kuang. 2025. Rethinking Causal Ranking: A Balanced Perspective on Uplift Model Evaluation. In Proceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu (Eds.). PMLR, 80134–80154. [46] Yaochen Zhu, Yinhan He, Jing Ma, Mengxuan Hu, Sheng Li, and Jundong Li. 2024. Causal Inference with Latent Variables: Recent Advances and Future Prospectives. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain) (KDD ’24). Association for Computing Machinery, New York, NY, USA, 6677–6687.
Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009