ConceptioArchivearXiv CS
arXiv CSopen access

CanniUplift: A Holistic Framework for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

arXiv:2607.05242v1 [cs.LG] 6 Jul 2026

CanniUplift: A Holistic Framework for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling Zuwang He∗

Shihao Shu∗

Yuli Qu∗

[email protected] Taobao & Tmall Group of Alibaba Beijing, China

[email protected] Taobao & Tmall Group of Alibaba Beijing, China

[email protected] Taobao & Tmall Group of Alibaba Hangzhou, Zhejiang, China

Hanyu Gao

Ziliang Zhang

Diwei Chen

[email protected] Taobao & Tmall Group of Alibaba Beijing, China

[email protected] Taobao & Tmall Group of Alibaba Beijing, China

[email protected] Taobao & Tmall Group of Alibaba Hangzhou, Zhejiang, China

Xiangda Yan

Buyu Gao

Tanchao Zhu

[email protected] Taobao & Tmall Group of Alibaba Hangzhou, Zhejiang, China

[email protected] Taobao & Tmall Group of Alibaba Hangzhou, Zhejiang, China

[email protected] Taobao & Tmall Group of Alibaba Hangzhou, Zhejiang, China

Yumeng Li†

Junxiong Zhu

[email protected] Taobao & Tmall Group of Alibaba Beijing, China

[email protected] Taobao & Tmall Group of Alibaba Hangzhou, Zhejiang, China

Abstract Personalized incentive allocation is vital for e-commerce, where uplift modeling is the standard for estimating Individual Treatment Effects (ITE). However, traditional models often fail in complex multiseller environments with violations of the Stable Unit Treatment Value Assumption (SUTVA). We identify two critical challenges: Seller-level Cannibalization, where incentives shift expenditure between shops without growing the platform, and Incentive-level Cannibalization, where organic conversions or alternative rewards introduce significant noise into incrementality estimation. In this paper, we propose CanniUplift, a unified framework to mitigate these dual-source cannibalization effects. Specifically, we design Platform-level Global Alignment (PGA) to capture cross-shop substitution through global GMV consistency constraints. To tackle incentive-driven noise, we introduce Redemption-based Decomposition Denoising (RDD), which uses redemption behavior to decompose treated outcomes and reduce attribution noise within an entire-space framework. Furthermore, a Treat-Attention mechanism is designed to model intricate interactions between users’ historical behaviors and current treatment options. Extensive experiments on both synthetic and large-scale industrial datasets demonstrate that CanniUplift significantly outperforms state-ofthe-art baselines. Ablation studies confirm that the integration of ∗ These authors contributed equally to this work. † Corresponding author.

This work is licensed under a Creative Commons Attribution 4.0 International License. KDD ’26, Jeju Island, Republic of Korea © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2259-2/2026/08 https://doi.org/10.1145/3770855.3818334

PGA and RDD consistently improves wAUUC and wQINI. Successfully deployed online, our framework achieved a 4.08% relative increase in platform-wide incremental GMV (ΔGMV) over the production baseline and improved ROI in online A/B tests, proving effective in driving global platform growth.

CCS Concepts • Information systems → Computational advertising.

Keywords Uplift modeling, Individual Treatment Effect, Cannibalization, Personalized Incentive Allocation, E-commerce ACM Reference Format: Zuwang He, Shihao Shu, Yuli Qu, Hanyu Gao, Ziliang Zhang, Diwei Chen, Xiangda Yan, Buyu Gao, Tanchao Zhu, Yumeng Li, and Junxiong Zhu. 2026. CanniUplift: A Holistic Framework for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD ’26), August 09–13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3770855.3818334

1

Introduction

With the rapid proliferation of digital commerce, personalized marketing has become a pivotal strategy for e-commerce platforms to drive user engagement and revenue growth. At the heart of these strategies lies the allocation of incentives, such as digital coupons and subsidies. Unlike traditional conversion rate (CVR) models that focus on predicting user behavior, Uplift Modeling has emerged as the standard for estimating the Individual Treatment Effect (ITE)—the true incremental value driven by a specific

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Browse Seq

User

See

Outcome

Parallel World Control (no coupon)

Green Shirt ($90) $90 Blue Shirt ($120)

Zuwang He et al.

learns cross-seller substitution effects, ensuring optimization aligns with platform-wide growth. • We propose a redemption-based decomposition module. By separating conversions with and without redemption, the model reduces attribution noise from organic demand and concurrent incentives, leading to more robust uplift estimation. • We design a treat-attention mechanism that models the intricate interactions between a user’s historical item-level engagement and their responsiveness to different incentive types, providing a more granular representation for personalized allocation. • Extensive evaluations on industrial and synthetic datasets show CanniUplift consistently outperforms state-of-the-art baselines. Successfully deployed on a real-world e-commerce system, the framework delivers marked gains in platformwide incremental GMV and ROI.

Seller GMV uplift: $0 Platform GMV uplift: $0

Buy it (GMV $90) Upsell Coupon

Seller GMV uplift: ▲ $120

lower pCVR $20 Off

Yellow Shirt ($100)

Platform GMV uplift: ▲ $30

Buy it (GMV $120) Downsell Coupon

Teal Shirt ($30)

Seller GMV uplift: ▲ $30

higher pCVR $5 Off

Full Wallet Existing

$50 Off

Existing

20% Off

Empty Wallet No other Coupons

Platform GMV uplift: ▼ $60

Buy it (GMV $30)

Conflict: Seller wins, platform loses

Compare

Receive test coupon

$50 vs $5 ?

$5 Off

Ignore $5, use $50

Receive test coupon

Only have $5

$5 Off

Use $5

Observed Redemption: No Observed Conversion: Yes ✓ True Uplift: Lower than pred Buy

(Would buy anyway with existing better coupons)

Observed Redemption: Yes ✓ Observed Conversion: Yes ✓ True Uplift: Higher than pred Buy

(Purchase driven by this allocated coupon)

Figure 1: Two critical violations of SUTVA in multi-seller e-commerce ecosystems. Top: Seller-level cannibalization. Bottom: Incentive-level cannibalization.

intervention. By identifying “persuadable” users, platforms can optimize resource allocation and maximize overall returns. However, applying uplift modeling in a complex, multi-seller e-commerce ecosystem faces significant theoretical and practical hurdles. Most existing uplift frameworks, such as DragonNet and EUEN, implicitly rely on the Stable Unit Treatment Value Assumption (SUTVA). This assumption posits that the treatment assigned to one unit does not affect the outcomes of other units. In a large-scale marketplace, this assumption is frequently violated due to what we term Multi-source Cannibalization Effects, as illustrated in Fig. 1, which manifest in two critical dimensions: First, Seller-level Cannibalization. Traditional uplift models optimize for individual seller increments, which inadvertently triggers “downsell effects”. As shown in Fig. 1 (top), a user might originally intend to purchase a $90 item, but a $5 coupon allocated to a different seller nudges them toward a $30 alternative. While the individual seller records a gain, the platform GMV drops. Such models fail to capture this “zero-sum” substitution, where seller uplift is achieved at the expense of global platform growth. Second, Incentive-level Cannibalization. In complex e-commerce environments, users are often exposed to multiple concurrent rewards. As shown in Fig. 1 (bottom), users with existing higher-value coupons may ignore the assigned test coupon but still complete a transaction. This “conversion without redemption” behavior leads to biased uplift estimation and wastes allocation opportunity costs, as the model misattributes organic motivation or alternative incentives to the assigned treatment. We propose CanniUplift, a holistic framework to mitigate dualsource cannibalization and boost platform-wide incremental GMV. Our contributions are four-fold: • We relax the SUTVA constraint via a multi-head architecture with a global GMV objective. By aggregating shop-level predictions under holistic constraints, the model implicitly

2

Preliminary

In this section, we establish the mathematical notation and foundational framework for uplift modeling within the context of platformwide e-commerce marketing. Unlike traditional supervised learning, uplift modeling aims to estimate the causal effect of a specific intervention.

2.1

Problem Setup and Notations

Let U be the set of users and S be the set of sellers (shops) on an e-commerce platform. For a marketing campaign, we observe a dataset: 𝑁 D = {(𝑥𝑖 , 𝑡𝑖,𝑠 , 𝑟𝑖,𝑠 , 𝑦𝑖,𝑠 )}𝑖=1 , (1) where 𝑥𝑖 ∈ R𝑑 represents the feature vector of user 𝑖, encompassing historical behaviors and context, 𝑡𝑖,𝑠 ∈ {0, 1} is the treatment indicator, where 𝑡𝑖,𝑠 = 1 if user 𝑖 is assigned an incentive (e.g., a shop-specific coupon) for seller 𝑠, and 𝑡𝑖,𝑠 = 0 otherwise, 𝑟𝑖,𝑠 ∈ {0, 1} is the redemption indicator, representing whether user 𝑖 actually utilized the assigned incentive for the transaction, and 𝑦𝑖,𝑠 ∈ R ≥0 denotes the observed outcome, such as the Gross Merchandise Volume (GMV) contributed by user 𝑖 to seller 𝑠.

2.2

The Potential Outcome Framework

Following the Neyman-Rubin Potential Outcome Framework [22], each user 𝑖 has two potential outcomes for a given seller 𝑠: 𝑦𝑖,𝑠 (1) and 𝑦𝑖,𝑠 (0), representing the outcome if the user were treated or not, respectively. The Individual Treatment Effect (ITE), also known as Conditional Average Treatment Effect (CATE), is defined as: 𝜏𝑖,𝑠 (𝑥𝑖 ) = E[𝑦𝑖,𝑠 (1) − 𝑦𝑖,𝑠 (0)|𝑋 = 𝑥𝑖 ].

(2)

The fundamental challenge of causal inference is that only one potential outcome is observed for any individual, i.e., 𝑦𝑖,𝑠 = 𝑡𝑖,𝑠 𝑦𝑖,𝑠 (1) + (1 − 𝑡𝑖,𝑠 )𝑦𝑖,𝑠 (0).

2.3

SUTVA and the Cannibalization Challenge

Standard uplift models typically rely on the Stable Unit Treatment Value Assumption (SUTVA) [9]. In practice, SUTVA is commonly interpreted as:

CanniUplift for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

(1) No Interference: the potential outcome of a unit is unaffected by the treatment assignments of other units. (2) No Hidden Versions of Treatment: the treatment is welldefined, i.e., there are no multiple versions of the “same” treatment that would lead to different potential outcomes. In our scenario, SUTVA is frequently violated due to Seller Cannibalization. Specifically, an incentive for seller 𝐴 may divert a user’s organic expenditure from seller 𝐵. Thus, the potential outcome 𝑦𝑖,𝑠 is not independent of treatments assigned to other sellers {𝑡𝑖,𝑠 ′ }𝑠 ′ ≠𝑠 . To address this, we define the Global Platform GMV for user 𝑖 as: ∑︁ 𝑌𝑖 = 𝑦𝑖,𝑠 . (3)

• Seller-level Cannibalization: This occurs when the Cross Effect term in the above equation is negative. It indicates that the observed growth in shop 𝑠 𝑗 is partially or fully offset by a decrease in other shops 𝑠𝑘 (𝑘 ≠ 𝑗), as the user merely reallocates their limited budget instead of generating new demand. • Incentive-level Cannibalization: This reflects the interaction between different reward types or organic motivation. When a user who would have converted organically (or via a platform-wide reward) utilizes the assigned shop coupon 𝑡𝑖,𝑠 𝑗 = 1, the Apparent Increment term becomes a “pseudo-incrementality”, effectively cannibalizing the platform’s potential profit or organic GMV that would have occurred without the specific treatment.

𝑠∈S

Rather than merely optimizing shop-level metrics, our objective is to maximize the global incrementality Δ𝑌𝑖 = E[𝑌𝑖 |Treat] − E[𝑌𝑖 |Control].

2.4

Incentive-level Noise

In real-world marketing, a user may be assigned a coupon (𝑡𝑖,𝑠 = 1) but purchase without redeeming it (𝑟𝑖,𝑠 = 0), a noncompliance phenomenon where the observed outcome is a mixture of latent compliance types. Formally, we regard 𝑟 as a post-treatment variable, decomposing the observed outcome distribution as: 𝑝 (𝑦 | 𝑡 = 1) = 𝑝 (𝑟 = 1 | 𝑡 = 1) 𝑝 (𝑦 | 𝑡 = 1, 𝑟 = 1) + 𝑝 (𝑟 = 0 | 𝑡 = 1) 𝑝 (𝑦 | 𝑡 = 1, 𝑟 = 0).

(4)

The second component (𝑡 = 1, 𝑟 = 0) corresponds to purchases that happen without redeeming the assigned coupon, which may be driven by organic demand or other concurrent incentives. This component contaminates uplift estimation if one naively treats all 𝑡 = 1 conversions as incremental. Therefore, it is crucial to distinguish the “conversion with redemption” and “conversion without redemption” mechanisms to reduce attribution bias in estimating the incremental value of the assigned incentive.

2.5

Formal Definition of Cannibalization Effects

To mathematically characterize the impact of interventions in a non-SUTVA environment, we define the Platform Net Increment Δ𝑌𝑖 (𝑠 𝑗 ) as the total change in user 𝑖’s platform-wide expenditure 𝑌𝑖 after assigning treatment 𝑡𝑖,𝑠 𝑗 = 1 (e.g., a specific shop-coupon) to seller 𝑠 𝑗 : Δ𝑌𝑖 (𝑠 𝑗 ) = E[𝑌𝑖 |𝑡𝑖,𝑠 𝑗 = 1] − E[𝑌𝑖 |𝑡𝑖,𝑠 𝑗 = 0]  ∑︁  = E[𝑦𝑖,𝑠𝑘 |𝑡𝑖,𝑠 𝑗 = 1] − E[𝑦𝑖,𝑠𝑘 |𝑡𝑖,𝑠 𝑗 = 0] 𝑠𝑘 ∈ S

  = E[𝑦𝑖,𝑠 𝑗 |𝑡𝑖,𝑠 𝑗 = 1] − E[𝑦𝑖,𝑠 𝑗 |𝑡𝑖,𝑠 𝑗 = 0] | {z }

(5)

Apparent Increment (Local Lift)

+

∑︁ 

E[𝑦𝑖,𝑠𝑘 |𝑡𝑖,𝑠 𝑗 = 1] − E[𝑦𝑖,𝑠𝑘 |𝑡𝑖,𝑠 𝑗 = 0]



𝑠𝑘 ≠𝑠 𝑗

|

{z

}

Cross Effect: Cannibalization (-) or Spillover (+)

where 𝑦𝑖,𝑠𝑘 denotes the GMV contributed by user 𝑖 to seller 𝑠𝑘 , as defined in Eq. (3). Based on this formulation, we generalize the cannibalization effect into two dimensions:

3

Methodology

Fig. 2 illustrates the overall framework. It is important to emphasize that the green module on the right corresponds to the baseline structure—a seller-view local uplift learning paradigm, whose core is to model the treated/control outcomes for each candidate seller independently and estimate uplift. In contrast, the left blue PGA module addresses Seller Cannibalization through platform-level global alignment, while the middle purple RDD module addresses Incentive Cannibalization through redemption-based decomposition. All three parts share the same bottom representation and candidate-interaction encoder designed in this work, while we redesign the task heads and training objectives to upgrade uplift learning from local incremental estimation to a unified framework of denoising + global alignment, which better matches real-world e-commerce environments with multi-seller competition and concurrent incentives.

3.1

Representation Learning and Candidate Interaction

In e-commerce incentive delivery, whether a user will be triggered to generate incremental value by a coupon or a shop largely depends on recent behaviors and historical coupon usage patterns, rather than static profiles alone. More importantly, the heterogeneity of uplift mainly arises from the user–candidate interaction: the same user may respond very differently to different shops, coupon amounts, thresholds, or validity periods. Therefore, instead of compressing all signals into a single user vector, CanniUplift adopts a two-stage design—multi-source sequence encoding + candidate treat-attention—to explicitly generate a candidate-specific interaction representation for each candidate. Multi-source sequence encoding (Shared Encoder). We construct three types of inputs, apply embedding, and feed them into a Transformer encoder: Candidate sets (used by the subsequent treat-attention): (1) treatment candidates: coupon/red-packet attributes including threshold, discount type (discount vs. full-reduction), face value, etc.; (2) seller candidates: the set of candidate shops that can be delivered/recommended in the current decision.

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

Zuwang He et al.

Figure 2: The architecture of the CanniUplift framework. The model consists of three core components: (Left) The Platform-level Global Alignment (PGA) module, which applies a global GMV consistency constraint to capture cross-shop substitution effects. (Middle) The Redemption-based Decomposition Denoising (RDD) module, which uses redemption behavior to decompose treated outcomes into redemption and non-redemption paths and reduce incentive-level attribution noise. (Right) The Sellerview Local Uplift module, which provides baseline ITE estimation. All modules are supported by a shared Multi-source Sequence Encoder and a Treat-Attention mechanism that captures fine-grained user-candidate interactions. Statistical static features (encoded by embedding layers): e.g., shop-side category information and user-side profile features. Sequential features (encoded by self-attention):

Here, 𝐻𝑢 retains fine-grained temporal signals (e.g., “whether the last purchase redeemed a coupon” and “recently preferred shops or categories”), which serve as the Key/Value for candidate interaction.

(1) user behavior sequence: time-ordered item/shop behaviors (impression/click/purchase, etc.), capturing short-term purchase intent and preference drift; (2) treatment sequence: historical incentive exposure /claim /redemption, capturing sensitivity to incentive mechanisms and redemption propensity.

Candidate interaction encoding: Treat-Attention (Candidate Encoder). For each user 𝑢, the system typically provides a set of candidates (candidate shops, candidate coupons, or their combinations). We denote each candidate by 𝑐, whose feature vector 𝐸𝑐 is formed by concatenating the candidate shop embedding and coupon attributes (face value, threshold, validity period, applicable categories, etc.). To capture heterogeneous responses across candidates, we adopt Treat-Attention: using the candidate as Query and the multi-source sequence representation as Key/Value to obtain a candidate-level interaction representation:

Multi-source inputs are encoded by a shared embedding layer and a Transformer encoder, producing the sequence representation 𝐻𝑢 and the pooled representation ℎ𝑢 for downstream use:  𝐻𝑢 = Transformer [𝐸 beh, 𝐸 ben, 𝐸 oth ] ,

ℎ𝑢 = Pool(𝐻𝑢 ).

(6)

 𝑧𝑢,𝑐 = Attn 𝑞(𝐸𝑐 ), 𝐾 (𝐻𝑢 ), 𝑉 (𝐻𝑢 ) .

(7)

CanniUplift for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

This mechanism explicitly captures which historical fragments are most relevant to the current candidate. For price-sensitive users, attention may focus more on historical coupon claiming and redemption behaviors; for brand-loyal users, it may concentrate on interactions with the current seller or similar sellers; for highly substitutable products, it may highlight same-category crossseller browsing, which provides informative signals for learning seller cannibalization. Finally, each candidate obtains an interaction vector 𝑧𝑢,𝑐 , which is fed into the seller-level, platform-level, and redemption-related heads to predict seller GMV, enforce platformlevel GMV consistency, and construct the RDD-based denoised GMV prediction.

3.2

Baseline Structure: Seller-view Local Uplift Learning

The green module on the right represents a common industrial uplift/ITE learning paradigm: it predicts GMV under treatment/control (or equivalently uplift) for each candidate seller and optimizes with shop-level supervision: Control Net outputs 𝑌0 (expected GMV without coupon) and Uplift Net outputs 𝑌1 (expected GMV with coupon), and computes uplift = 𝑌1 − 𝑌0 . For each candidate seller, its observed GMV is used to supervise the corresponding output. Such seller-view methods implicitly rely on SUTVA (no interference). In platform marketing, this assumption is often violated due to cross-seller substitution and incentive noncompliance/credit contamination, which motivates our two corresponding modules introduced below.

3.3

Platform-level Global Alignment

To directly optimize platform-wide net incrementality rather than single-shop growth, we introduce a Platform head (Head1 in Fig. 2) on top of the shared representation and candidate encoder. It aggregates predictions over candidate sellers and aligns them with the user’s total platform GMV, encouraging the model to capture cross-seller cannibalization effects from the training objective. Given the candidate seller set S𝑢 for user 𝑢, the platform head pla outputs an aggregatable prediction 𝑝𝐺𝑀𝑉𝑠𝑖 for each 𝑠𝑖 ∈ S𝑢 , and computes: ∑︁ pla pla 𝑝𝐺𝑀𝑉𝑢 = 𝑝𝐺𝑀𝑉𝑠𝑖 . (8) 𝑠𝑖 ∈ S𝑢

From logs we also obtain the user’s total platform GMV within the observation window (summed across shops). During training, we minimize their discrepancy:  ∑︁  pla pla Lpla = L 𝑝𝐺𝑀𝑉𝑢 , 𝐺𝑀𝑉𝑢 . (9) 𝑢

Under this constraint, if the model predicts high uplift for multiple shops simultaneously from the seller view, their aggregation will pla systematically inflate 𝑝𝐺𝑀𝑉𝑢 and be penalized by the platform loss. Conversely, when a shop’s “lift” mainly comes from other shops’ “loss”, the total platform pie does not grow, and the platform constraint suppresses such spurious uplift. Therefore, the platform head provides an additive consistency global training signal, enabling the model to implicitly learn substitution relations without explicitly modeling all pairwise seller interactions.

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

While training aggregates all sellers per user to enforce global consistency, inference only requires single-seller predictions to estimate platform-level uplift. The detailed derivation is provided in Appendix C.

3.4

Redemption-based Decomposition Denoising

In practice, treatment assignment (issuing a coupon) does not necessarily imply treatment receipt (redeeming it). A user may purchase under assignment but not redeem the assigned coupon (𝑟 = 0), either due to purely organic motivation or because other concurrent incentives “hijack” the conversion credit. This noncompliance phenomenon makes 𝑡=1 outcomes a mixture of heterogeneous mechanisms, which can severely contaminate uplift attribution if one naively treats all treated conversions as incremental. To explicitly model this incentive cannibalization, we introduce Redemption-based Decomposition Denoising (RDD), which factorizes the treated outcome into two disjoint paths: purchase with redemption and purchase without redemption. For each user– candidate seller pair (𝑢, 𝑠𝑖 ), the Redem Net predicts the redemption probability: 𝑝𝑟𝑠𝑖 ≜ 𝑝 (𝑟 = 1 | 𝑧𝑢,𝑠𝑖 ), (10) where 𝑧𝑢,𝑠𝑖 is the candidate-specific interaction representation produced by Treat-Attention. The treatment-side head outputs two 𝑠𝑖 path-specific GMV increments, denoted as 𝑝Δ𝐺𝑀𝑉𝑟𝑠𝑖 and 𝑝Δ𝐺𝑀𝑉1−𝑟 , corresponding to the redemption path (𝑟 =1) and the non-redemption path (𝑟 =0), respectively. The control-side head outputs the baseline GMV under no treatment, denoted as 𝑝𝐺𝑀𝑉𝑐𝑠𝑖 . Following the law of total expectation, the final seller-level GMV prediction used for supervision is constructed as a redemptionweighted mixture: š 𝑢,𝑠𝑖 = 𝑝𝐺𝑀𝑉𝑐𝑠𝑖 + 𝑝𝑟𝑠𝑖 · 𝑝Δ𝐺𝑀𝑉𝑟𝑠𝑖 + (1 − 𝑝𝑟𝑠𝑖 ) · 𝑝Δ𝐺𝑀𝑉 𝑠𝑖 . (11) 𝐺𝑀𝑉 1−𝑟 This design encourages the model to explain treated outcomes through two behaviorally distinct paths, so that conversions without redemption are not forced to share the same uplift pattern as conversions with redemption. Unlike a pure auxiliary redemption prediction task, RDD embeds redemption directly into the main treated-outcome path, which follows the spirit of Entire-Space[18] modeling by explicitly accounting for the post-treatment redempš 𝑢,𝑠𝑖 is supervised by the observed sellertion dimension. Here, 𝐺𝑀𝑉 level GMV through the seller-level Tweedie loss, while 𝑝𝑟𝑠𝑖 is supervised by the observed redemption label through the redemption classification loss. The formal definitions of these loss terms, together with the platform-level aggregation loss, are given in Sec. 3.5.

3.5

Loss Function

GMV prediction in e-commerce typically exhibits zero inflation (many users do not purchase, yielding 𝑦 = 0) and a long-tail distribution (a small fraction of large orders dominate). To better fit both properties, improve training stability, and enhance tail modeling, we adopt the Tweedie loss for all GMV-related outputs (seller local GMV, platform aggregated GMV, and GMV heads in the denoising paths): ∑︁ Tw ˆ LGMV = LTw (𝑦, 𝑦), (Tweedie family, 1 < 𝑝 < 2). (12)

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

Zuwang He et al.

The Tweedie family can model the compound structure of “a point mass at zero + continuous positive values” under a unified distributional assumption, which better matches the real GMV generation process. Meanwhile, the redemption sub-task (Redem head) uses cross-entropy loss. The overall objective consists of three terms: seller-level GMV supervision, redemption classificaš 𝑢,𝑠𝑖 denotes the tion, and platform-level aggregation. Here, 𝐺𝑀𝑉 seller-level GMV prediction constructed by the seller-view or RDD pla š 𝑢,𝑠 branch, 𝐺𝑀𝑉 denotes the platform-view prediction produced 𝑖 by the Platform head, and 𝑟ˆ𝑢,𝑠𝑖 denotes the redemption probability predicted by the Redem head: Lseller =

∑︁ ∑︁

  š 𝑢,𝑠𝑖 , LTw 𝐺𝑀𝑉𝑢,𝑠𝑖 , 𝐺𝑀𝑉

(13)

𝑢 𝑠𝑖 ∈ S𝑢

∑︁

(14) (15)

𝑢 𝑠𝑖 ∈ S𝑢

The final training objective is: Ltotal = Lseller + 𝜆pla Lpla + 𝜆redem Lredem .

(16)

In our experiments, 𝜆pla and 𝜆redem are set to 1 unless otherwise specified. The overall training procedure is summarized in Appendix B.

4 Experiment 4.1 Datasets 4.1.1 Industrial Dataset. Our industrial data are collected from the online logs of a large-scale e-commerce system. This dataset constitutes a highly challenging task: when aggregating the Individual Treatment Effects (ITE) for all associated sellers of a single user using baseline models, the estimated total gain overestimates the actual online observed effect by approximately 35%. Furthermore, among the behaviors where coupons are issued and conversions occur, approximately 25% ultimately fail to complete redemption, indicating significant “pseudo-conversion” noise that further interferes with the estimation of true incremental effects. 4.1.2 Synthetic Dataset. We construct a synthetic dataset comprising user features x𝑢 , seller features x𝑠 , treatment variable 𝑡, and potential outcomes 𝑦0 and 𝑦1 . To simulate seller cannibalization, we introduce a decay mechanism where uplift diminishes based on the number of same-category sellers the user has sequentially encountered. Due to the complexity of multi-incentive interactions and compliance heterogeneity, we focus solely on seller-level cannibalization in synthetic data, while RDD’s effectiveness on incentive cannibalization is validated through industrial data ablations. Complete generation details are provided in Appendix A.

4.2

where: 𝜙 represents the population proportion selected after ranking by the model’s uplift scores from high to low (e.g., 𝜙 = 0.1 represents the top 10% with the highest scores); 𝑁 𝑇 (𝜙) and 𝑅𝑇 (𝜙) are the total number of users and the total GMV generated in the treatment group within that population, respectively; 𝑁 𝐶 (𝜙) and 𝑅𝐶 (𝜙) correspond to the control group. AUUC measures the area under the Uplift curve: ∫ 1 AUUC = Uplift(𝜙) 𝑑𝜙, (18) 0

pla ª © pla š 𝑢,𝑠 LTw ­𝐺𝑀𝑉𝑢 , 𝐺𝑀𝑉 ®, 𝑖 𝑢 𝑠 ∈ S 𝑢 𝑖 « ¬ ∑︁ ∑︁  Lredem = CE 𝑟𝑢,𝑠𝑖 , 𝑟ˆ𝑢,𝑠𝑖 .

Lpla =

∑︁

treatment group relative to the control group within that subset[7]. The value at any point Uplift(𝜙) on the curve is computed as:   𝑇 𝑅 (𝜙) 𝑅𝐶 (𝜙) − × (𝑁 𝑇 (𝜙) + 𝑁 𝐶 (𝜙)), (17) Uplift(𝜙) = 𝑁 𝑇 (𝜙) 𝑁 𝐶 (𝜙)

Evaluation Metrics

4.2.1 AUUC. AUUC is computed based on the Uplift curve, which depicts: when we rank the population from high to low according to the model’s predicted uplift scores and select different proportions of the population for intervention, the cumulative total uplift of the

where a larger value indicates stronger model capability in identifying high-uplift users. 4.2.2 QINI. The QINI curve depicts the cumulative net incremental gain at each population proportion 𝜙. The value at any point 𝑄 (𝜙) is computed as: © ∑︁ ª |𝑇 (𝜙)| . (19) 𝑦𝑖 − ­ 𝑦𝑗 ® · |𝐶 (𝜙)| 𝑖 ∈𝑇 (𝜙 ) « 𝑗 ∈𝐶 (𝜙 ) ¬ The core idea is to calibrate for differences in treatment and control group sizes, thereby fairly evaluating net gain. The QINI coefficient is defined as the area between the model’s QINI curve and the random selection baseline: ∫ 1 QINI = (𝑄 model (𝜙) − 𝑄 random (𝜙)) 𝑑𝜙 . (20) ∑︁

𝑄 (𝜙) =

0

A higher QINI value indicates better model performance in identifying individuals with high treatment effects. 4.2.3 Weighted AUUC/QINI. In the industrial dataset with multiple treatment types (e.g., coupons with different values), we adopt weighted AUUC and QINI (wAUUC and wQINI). Continuous treatment values are discretized into bins, with AUUC/QINI computed per bin and aggregated via sample-size weighting. We report AUUC/QINI at two granularities. Seller-level metrics evaluates uplift ranking for each user–seller candidate pair, reflecting the model’s ability to identify high-increment seller-level allocation opportunities. User-level metrics first aggregates the predicted uplift over candidate sellers associated with the same user and then evaluates the ranking at the user/platform level, reflecting whether the model can identify users with higher platform-wide incremental potential. For the industrial dataset, both seller-level and user-level metrics are computed in their weighted forms due to multiple treatment values.

4.3

Experimental Setup

Hardware. All experiments on the industrial dataset are conducted on the internal production system. Experiments on the synthetic dataset are also performed on the internal computing platform with abundant computational resources. Optimization and Training. We employ Optuna[2] for automated hyperparameter tuning with 40 trials per experiment. Each

CanniUplift for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

Table 1: Main experimental results on both synthetic and industrial datasets. Model DragonNet CFRNet TARNet TLearner EUEN X-net M3TN UMLC HUM Ours

Industrial Dataset

Synthetic Dataset

Seller AUUC

Seller QINI

User AUUC

User QINI

Seller AUUC

Seller QINI

User AUUC

User QINI

0.650 0.694 0.729 0.730 0.744 0.713 0.709 0.681 0.701 0.769

0.248 0.261 0.284 0.289 0.297 0.273 0.281 0.255 0.266 0.314

0.705 0.748 0.771 0.782 0.794 0.786 0.775 0.742 0.759 0.849

0.238 0.276 0.294 0.301 0.319 0.307 0.313 0.271 0.289 0.348

1.0475 1.0084 1.0074 0.8308 1.0940 1.0612 1.1382 1.0291 0.9714 1.1596

0.5432 0.4966 0.5030 0.3305 0.5837 0.5589 0.6273 0.5263 0.4728 0.6380

0.9438 0.9960 1.1052 0.9548 1.1245 0.9481 0.9714 0.9165 0.9612 1.1474

0.4430 0.4839 0.5972 0.4512 0.6024 0.4553 0.4591 0.4238 0.4613 0.6237

trial trains for a maximum of 30 epochs with early stopping if validation metrics do not improve for 5 consecutive epochs. The batch size is set to 4096 with L2 regularization coefficient 𝜆 = 1×10−4 . The search space includes learning rate 𝜂 ∈ {5 × 10−4, 1 × 10−4, 5 × 10−5 } and hidden dimension 𝑑ℎ ∈ {128, 256, 512}.

4.4

Main Experimental Results

We select DragonNet[26], CFRNet[25], TARNet[25], TLearner[15], EUEN[14], M3TN[27], UMLC[28], HUM[36] as baseline methods, with evaluation metrics using the commonly adopted AUUC and QINI in uplift modeling. As shown in Table 1, on the industrial dataset, DragonNet, CFRNet, and TARNet, which adopt shared network architectures with multi-head structures, perform poorly. In contrast, TLearner and EUEN, which decouple the control network independently, demonstrate clear advantages. This is primarily because, unlike simple binary treatment problems, our business scenario contains very rich treatment-related features, and the control group prediction should shield this information to avoid bias; this decoupling is more easily achieved in the EUEN and TLearner structures. Since EUEN performs best among all baseline models, we select it as our baseline. In the industrial dataset, “Ours” denotes the addition of Platformlevel Global Alignment and Redemption-based Decomposition Denoising to the EUEN baseline, abbreviated as PGA and RDD, respectively. In the synthetic dataset, since only seller-level cannibalization is simulated, “Ours” denotes the addition of PGA to EUEN. Across both datasets and all metrics, our method achieves the best performance. The synthetic results verify the effectiveness of PGA under controlled seller-level cannibalization, while the industrial results further demonstrate the benefit of jointly mitigating sellerlevel and incentive-level cannibalization in real-world deployment.

4.5

Ablation Studies

To verify the effectiveness of each module, we conduct ablation experiments. The baseline uses EUEN to estimate GMV-related uplift. As shown in Table 2, adding PGA improves performance. Since RDD uses redemption labels as supervision, we introduce a control variant, denoted as “+Redem”, which adds redemption prediction as an auxiliary task but does not decompose the treated outcome

into redemption and non-redemption paths. This comparison isolates the benefit of the proposed path-level decomposition from the mere use of redemption supervision. Under the same redemptionlabel supervision, RDD still improves over +Redem, showing that the decomposition structure contributes beyond auxiliary redemption prediction. When PGA and RDD are combined, they mutually reinforce each other, yielding the best performance. It is noteworthy that PGA brings a particularly large improvement on User AUUC, suggesting that platform-level alignment helps rank users by their platform-wide incremental potential. In contrast, RDD yields consistent gains over the +Redem variant, especially on QINI-related metrics, indicating that explicitly separating redemption and non-redemption paths helps reduce overattribution in cumulative uplift estimation. AUUC mainly reflects ranking quality, while QINI emphasizes cumulative net gain. These results show that PGA and RDD improve uplift estimation from complementary perspectives: PGA aligns seller-level predictions with platform-level growth, while RDD reduces attribution noise from mixed conversion paths. Table 2: Ablation study results on the industrial dataset. We compare the base EUEN model with different combinations of Platform-level Global Alignment (PGA) and Redemptionbased Decomposition Denoising (RDD). Metric Seller AUUC Seller QINI User AUUC User QINI

4.6

Baseline

+PGA

+Redem

+RDD

Ours

0.744 0.297 0.794 0.319

0.751 0.302 0.826 0.337

0.749 0.299 0.803 0.326

0.757 0.308 0.818 0.332

0.769 0.314 0.849 0.348

Seller Cannibalization Analysis

To verify whether the model effectively learns cannibalization effects, we define the individual estimated cannibalization rate as: 𝑝 seller − 𝑝 platform 𝑔= , (21) 𝑝 seller where 𝑝 seller is the predicted seller-level incremental GMV and 𝑝 platform is the predicted platform-level incremental GMV.

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

Figure 3: Correlation between estimated cannibalization rate and user behavioral features. X-axis: model-predicted cannibalization rate 𝑔. Blue curve: ratio of recently browsed samecategory item prices to current item price. Orange curve: number of same-category sellers visited in the past 10 days. Both features show monotonic positive correlation with cannibalization rate.

Since individual-level ground truth cannibalization rates are unavailable, we validate the model’s effectiveness by analyzing two typical cannibalization scenarios: (1) users who frequently visit same-category sellers recently, indicating strong organic purchase intent where coupons may fail to generate true incrementality; (2) users whose recently browsed same-category items have significantly higher average prices than the current item, where coupons may lead users to purchase cheaper items, resulting in negative incrementality. Based on these hypotheses, we conduct validation using users’ click sequences from the past 10 days. As shown in Fig. 3, both behavioral features demonstrate strong positive correlations with estimated cannibalization rates, validating our modeling assumptions. The blue curve shows that users who recently browsed higher-priced same-category items exhibit higher cannibalization rates, confirming that coupons for lowerpriced items are more likely to divert purchases from higher-value alternatives. The orange curve reveals that users who visited more same-category sellers also show elevated cannibalization rates, indicating that cross-shop browsing behavior signals stronger organic intent and substitution effects. These consistent patterns verify that the PGA module effectively captures seller-level cannibalization signals during training.

4.7

Incentive Cannibalization Analysis

To validate the effectiveness of the Redemption-based Decomposition Denoising (RDD) module in mitigating incentive cannibalization, we analyze the predicted uplift for users who converted without redeeming their assigned coupons. These “conversion without redemption” cases represent a key source of noise, as they may be driven by organic demand or alternative incentives rather than the assigned treatment, leading to overestimated uplift. Fig. 4 compares the predicted uplift between the baseline model (without RDD) and our proposed method (with RDD) across two user groups segmented by redemption behavior (r). For users who

Zuwang He et al.

Figure 4: Comparison of predicted GMV uplift between baseline (without RDD) and proposed method (with RDD), segmented by redemption behavior. Left (r=0): users who converted without redeeming the assigned coupon. Right (r=1): users who converted with redemption. Numbers indicate average predicted uplift in each segment. The baseline overestimates uplift in the r=0 group by 7.6% (3.435 vs 3.192), while both models align in the r=1 group. redeemed their coupons (𝑟 = 1), both models produce similar predictions, since redemption provides a direct observable signal that the assigned coupon participated in the conversion path. However, for users who converted without redemption (r=0), the baseline model predicts substantially higher uplift compared to our RDDequipped model. This discrepancy reveals the presence of incentive cannibalization: the r=0 group likely includes users with strong organic intent or those influenced by higher-value concurrent incentives, leading the baseline to overattribute incrementality to the assigned (but unused) coupon. By explicitly decomposing the two conversion paths, RDD assigns lower uplift to such non-redemption conversions, reducing overestimation by approximately 7% in the r=0 segment. This denoising pattern suggests that RDD mitigates over-attribution in non-redemption conversions and improves the robustness of uplift estimation.

4.8

Online A/B Test Results Table 3: Online A/B test results Marketing Cost

Platform ΔGMV

Platform ROI

-2.45%

+4.08%

+6.69%

To validate the effectiveness of the proposed method in realworld business scenarios, we deployed an online A/B test from October 5, 2025 to October 12, 2025. ROI is defined as: ΔGMV ROI = , (22) Cost where ΔGMV represents the incremental GMV of the treatment group relative to the control group.

CanniUplift for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

As shown in Table 3, compared to the production baseline model, our approach reduces marketing costs by 2.45% and improves platform-wide ΔGMV by 4.08%, resulting in a 6.69% increase in platform-wide ROI. These results confirm that our framework effectively denoises cannibalization effects in ITE estimation and reduces opportunity cost waste, demonstrating significant practical value in real-world business applications.

5 Related Work 5.1 Foundations of Uplift Modeling Uplift modeling, or Individual Treatment Effect estimation, is rooted in the Potential Outcomes Framework[9, 22] and Causal Diagram theory[20]. Early research focused on Tree-based methods, which adapt splitting criteria to maximize the difference in treatment effects[4, 8, 23]. Subsequently, Meta-learners gained popularity for their flexibility in using any supervised learning algorithm as a base. This category includes S-Learner[15], T-Learner[15], X-Learner[15], and R-Learner[19], which utilize different objective functions to minimize the mean squared error of the estimated uplift. While foundational, these methods often struggle with high-dimensional features and selection bias in observational data[10, 31].

5.2

Deep Representation Learning for ITE

With the success of deep learning, neural network-based approaches have been proposed to learn balanced representations to mitigate selection bias[12, 16, 34]. A seminal work, CFRNet[25], introduced the Integral Probability Metric (IPM) to regularize the distance between treatment and control group distributions[12, 24]. To further enhance estimation accuracy, DragonNet[26] incorporated propensity score estimation into representation learning, leveraging the sufficiency of the propensity score for unbiasedness[3, 21]. Recent advances also explore information-theoretic bounds[5] and generative adversarial networks (GANs)[35] to synthesize counterfactual outcomes. However, these models predominantly operate under the SUTVA assumption, which ignores the complex interdependencies in competitive marketplaces.

5.3

Multi-treatment and Multi-task Learning in Marketing

E-commerce platforms often involve multiple treatment options (e.g., various coupon types or discount levels), leading to the development of multi-treatment uplift models[1, 37]. Multi-task learning (MTL) architectures, such as MMoE[17] and PLE[29], have been adapted to jointly model conversion and redemption[11, 38]. For instance, EUEN[14] and DESCN[39] modeling the entire user space to alleviate data sparsity. Despite these efforts, existing MTL-based uplift models typically treat individual seller increments as independent goals, failing to account for the Seller-level Cannibalization where gains in one shop are offset by losses in another.

5.4

Causal Inference under Interference and Noise

The violation of SUTVA, known as interference or spillover effects, has been studied in social networks and marketplace experiments[13,

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

30]. In e-commerce, this manifests as the cannibalization of organic traffic or cross-seller competition. Furthermore, the gap between treatment assignment and actual redemption—what we term Incentive Cannibalization—introduces significant measurement noise[33]. While recent works have explored denoising techniques for recommender systems[6, 32], their application in uplift modeling remains sparse. Our work, CanniUplift, bridges this gap by proposing a holistic framework that explicitly models multi-source cannibalization through platform-level alignment and redemptionbased denoising.

6

Conclusion

In this paper, we address the critical yet often overlooked challenge of multi-source cannibalization in e-commerce uplift modeling. By relaxing the SUTVA assumption, we identify two primary forms of interference: Seller Cannibalization, where incentives merely shift expenditure between shops, and Incentive Cannibalization, where rewards are redeemed by organically motivated users. To this end, we propose CanniUplift, which effectively filters out “pseudoincrementality” by enforcing global consistency and explicitly modeling the redemption mechanism. Extensive experiments on both synthetic and large-scale industrial datasets demonstrate that CanniUplift significantly outperforms state-of-the-art uplift models. More importantly, online A/B tests in a real-world production environment confirm that our framework can reduce marketing costs while simultaneously driving higher platform-wide incremental GMV and ROI. These results underscore the importance of accounting for cross-unit interference in complex marketplace ecosystems.

7

Limitation

CanniUplift explicitly models seller-level and incentive-level cannibalization but does not address temporal cannibalization, where promotions cause users to shift future purchases forward and inflate short-term uplift at the cost of long-term revenue. Additionally, because CanniUplift relies on historical user behavior to estimate cannibalization effects, it suffers from latency in capturing realtime substitution. Future work will explore sequential and forwardlooking uplift modeling, such as predicting future seller visits and coupon redemptions, to better capture delayed purchase shifts and dynamic substitution patterns.

References [1] Naoufal Acharki, Ramiro Lugo, Antoine Bertoncello, and Josselin Garnier. 2023. Comparison of meta-learners for estimating multi-valued treatment heterogeneous effects. In International conference on machine learning. PMLR, 91–132. [2] Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 2623–2631. [3] Serge Assaad, Shuxi Zeng, Chenyang Tao, Shounak Datta, Nikhil Mehta, Ricardo Henao, Fan Li, and Lawrence Carin. 2021. Counterfactual representation learning with balancing weights. In International Conference on Artificial Intelligence and Statistics. PMLR, 1972–1980. [4] Susan Athey and Guido Imbens. 2016. Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences 113, 27 (2016), 7353–7360. [5] Alicia Curth and Mihaela Van der Schaar. 2021. On inductive biases for heterogeneous treatment effect estimation. Advances in Neural Information Processing Systems 34 (2021), 15883–15894. [6] Zhihao Guo, Peng Song, Chenjiao Feng, Kaixuan Yao, Chuangyin Dang, and Jiye Liang. 2025. Causal intervention for knowledge graph denoising in recommender

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

systems. International Journal of Machine Learning and Cybernetics 16, 11 (2025), 8551–8567. [7] Pierre Gutierrez and Jean-Yves Gérardy. 2017. Causal inference and uplift modelling: A review of the literature. In International conference on predictive applications and APIs. PMLR, 1–13. [8] Behram Hansotia and Brad Rukstales. 2002. Incremental value modeling. Journal of Interactive Marketing 16, 3 (2002), 35–46. [9] Paul W Holland. 1986. Statistics and causal inference. Journal of the American statistical Association 81, 396 (1986), 945–960. [10] Guido W Imbens and Donald B Rubin. 2015. Causal inference in statistics, social, and biomedical sciences. Cambridge university press. [11] Jiarui Jin, Xianyu Chen, Weinan Zhang, Yuanbo Chen, Zaifan Jiang, Zekun Zhu, Zhewen Su, and Yong Yu. 2022. Multi-scale user behavior network for entire space multi-task learning. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 874–883. [12] Fredrik Johansson, Uri Shalit, and David Sontag. 2016. Learning representations for counterfactual inference. In International conference on machine learning. PMLR, 3020–3029. [13] Ramesh Johari, Hannah Li, Inessa Liskovich, and Gabriel Y Weintraub. 2022. Experimental design in two-sided platforms: An analysis of bias. Management Science 68, 10 (2022), 7069–7089. [14] Wenwei Ke, Chuanren Liu, Xiangfu Shi, Yiqiao Dai, Philip S Yu, and Xiaoqiang Zhu. 2021. Addressing exposure bias in uplift modeling for large-scale online advertising. In 2021 IEEE International Conference on Data Mining (ICDM). IEEE, 1156–1161. [15] Sören R Künzel, Jasjeet S Sekhon, Peter J Bickel, and Bin Yu. 2019. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences 116, 10 (2019), 4156–4165. [16] Christos Louizos, Uri Shalit, Joris M Mooij, David Sontag, Richard Zemel, and Max Welling. 2017. Causal effect inference with deep latent-variable models. Advances in neural information processing systems 30 (2017). [17] Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-ofexperts. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1930–1939. [18] Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu, Xiaoqiang Zhu, and Kun Gai. 2018. Entire space multi-task model: An effective approach for estimating post-click conversion rate. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 1137–1140. [19] Xinkun Nie and Stefan Wager. 2021. Quasi-oracle estimation of heterogeneous treatment effects. Biometrika 108, 2 (2021), 299–319. [20] Judea Pearl. 2009. Causality. Cambridge university press. [21] Paul R Rosenbaum and Donald B Rubin. 1983. The central role of the propensity score in observational studies for causal effects. Biometrika 70, 1 (1983), 41–55. [22] Donald B Rubin. 1974. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology 66, 5 (1974), 688. [23] Piotr Rzepakowski and Szymon Jaroszewicz. 2010. Decision trees for uplift modeling. In 2010 IEEE International Conference on Data Mining. IEEE, 441–450. [24] Patrick Schwab, Lorenz Linhardt, and Walter Karlen. 2018. Perfect match: A simple method for learning representations for counterfactual inference with neural networks. arXiv preprint arXiv:1810.00656 (2018). [25] Uri Shalit, Fredrik D Johansson, and David Sontag. 2017. Estimating individual treatment effect: generalization bounds and algorithms. In International conference on machine learning. PMLR, 3076–3085. [26] Claudia Shi, David Blei, and Victor Veitch. 2019. Adapting neural networks for the estimation of treatment effects. Advances in neural information processing systems 32 (2019). [27] Zexu Sun and Xu Chen. 2024. M3TN: Multi-Gate Mixture-of-Experts Based Multi-Valued Treatment Network for Uplift Modeling. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 5065–5069. doi:10.1109/ICASSP48485.2024.10446323 [28] Zexu Sun, Qiyu Han, Minqin Zhu, Hao Gong, Dugang Liu, and Chen Ma. 2025. Robust uplift modeling with large-scale contexts for real-time marketing. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 1325–1336. [29] Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong. 2020. Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations. In Proceedings of the 14th ACM conference on recommender systems. 269–278. [30] Eric J Tchetgen Tchetgen and Tyler J VanderWeele. 2012. On causal inference in the presence of interference. Statistical methods in medical research 21, 1 (2012), 55–75. [31] Stefan Wager and Susan Athey. 2018. Estimation and inference of heterogeneous treatment effects using random forests. J. Amer. Statist. Assoc. 113, 523 (2018), 1228–1242. [32] Wenjie Wang, Fuli Feng, Xiangnan He, Liqiang Nie, and Tat-Seng Chua. 2021. Denoising implicit feedback for recommendation. In Proceedings of the 14th ACM international conference on web search and data mining. 373–381.

Zuwang He et al.

[33] Hong Wen, Jing Zhang, Yuan Wang, Fuyu Lv, Wentian Bao, Quan Lin, and Keping Yang. 2020. Entire space multi-task modeling via post-click behavior decomposition for conversion rate prediction. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 2377–2386. [34] Liuyi Yao, Sheng Li, Yaliang Li, Mengdi Huai, Jing Gao, and Aidong Zhang. 2018. Representation learning for treatment effect estimation from observational data. Advances in neural information processing systems 31 (2018). [35] Jinsung Yoon, James Jordon, and Mihaela Van Der Schaar. 2018. GANITE: Estimation of individualized treatment effects using generative adversarial nets. In International conference on learning representations. [36] Chenhao Zhai, Chang Meng, Xueliang Wang, Shuchang Liu, Xiaolong Hu, Shisong Tang, Xiaoqiang Feng, and Xiu Li. 2026. Heterogeneous Multi-treatment Uplift Modeling for Trade-off Optimization in Short-Video Recommendation. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 2562–2572. [37] Yan Zhao, Xiao Fang, and David Simchi-Levi. 2017. Uplift modeling with multiple treatments and general response types. In Proceedings of the 2017 SIAM International Conference on Data Mining. SIAM, 588–596. [38] Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed Chi. 2019. Recommending what video to watch next: a multitask ranking system. In Proceedings of the 13th ACM conference on recommender systems. 43–51. [39] Kailiang Zhong, Fengtong Xiao, Yan Ren, Yaorong Liang, Wenqing Yao, Xiaofeng Yang, and Ling Cen. 2022. Descn: Deep entire space cross networks for individual treatment effect estimation. In Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 4612–4620.

A Synthetic Dataset Generation A.1 Data Configuration The synthetic dataset includes user features x𝑢 ∈ R106 , seller features x𝑠 ∈ R107 , treatment variable 𝑡 ∈ {0, 1}, and potential outcomes {𝑦0, 𝑦1 }. We explicitly model seller-level cannibalization effects to simulate diminishing marginal returns under repeated exposure across similar seller segments.

A.2

Feature Generation

User Features. User features comprise 𝑝 = 106 dimensions: 𝑝𝑏 = 23 binary features sampled from Bernoulli(0.5), 𝑝𝑐 = 83 continuous features from N (0, 1), and one normalized categorical attribute derived from 10 discrete categories, represented as a continuous value in [0, 1). Seller Features. Seller features comprise 𝑞 = 107 dimensions: 𝑞𝑐 = 106 continuous features from N (0, 1), and one categorical feature with 300 original categories (normalized to [0, 1)). This categorical feature is mapped to 𝐾 = 10 semantic segments via a Gaussian kernel-based soft clustering mechanism to model heterogeneous treatment effects. An auxiliary cluster label    1 1 (23) 𝑥 cluster ∼ Multinomial 6; , . . . , 6 6 is introduced for response generation only.

A.3

User–Seller Interaction Structure

We generate 500,000 unique sellers. For each of the 20,000 users, we randomly assign exactly 𝑁 = 50 sellers with replacement. Samples from the same user share identical x𝑢 and treatment assignment 𝑡.

A.4

Treatment Assignment

Following the Randomized Controlled Trial (RCT) paradigm, treatment is assigned at the user level: 𝑡 ∼ Bernoulli(0.5)

(24)

CanniUplift for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling

C

Inference with Platform-level Global Alignment C.1 From Global Training to Marginal Inference

All seller samples for a given user inherit the same 𝑡 value.

A.5

Response Generation

Control Outcome. The control outcome is defined as: 𝑦0 = 𝑎 · ∥x𝑢 ∥ 22 + 𝑏 · ∥x𝑠 ∥ 22 + 𝑐 · 𝜙 (x𝑢⊤ x𝑠 ) + 𝑑 · 𝑥𝑠class + 𝑧 0 + 𝜖0 where 𝜙 (𝑠) is a truncated quadratic function: ( 𝑠 2, |𝑠 | ≤ 𝛿 𝜙 (𝑠) = (𝛿 = 15), 2𝛿 |𝑠 | − 𝛿 2, otherwise

(25)

(26)

(27)

where (𝑒, 𝑓 , 𝑔) = (2, 2, 0.2), 𝜖1 ∼ N (0, 10), 𝑧 1 follows the same distribution as 𝑧 0 , and Δ ≥ 0 is the gain decay term reflecting diminishing marginal returns from repeated exposure.

Seller-Level Cannibalization Mechanism

Existing synthetic data generation methods for uplift modeling, such as UMLC[28], typically assume independent seller effects and ignore cross-seller interference. In contrast, we propose a novel cannibalization-aware simulation mechanism that explicitly models diminishing marginal returns when users are repeatedly exposed to sellers within the same semantic segment. Soft Clustering. The 300 original seller categories are soft-clustered into 𝐾 = 10 semantic segments via Gaussian kernels. Given category 𝑐 ∈ [0, 299], the probability of mapping to segment 𝑠 ∈ [0, 9] is:   (𝑐 −𝜇 ) 2 exp − 2𝜎 2𝑠   𝑃 (𝑠 |𝑐) = Í (28) (𝑐 −𝜇𝑘 ) 2 9 𝑘=0 exp − 2𝜎 2 where 𝜇𝑠 = 15 + 30𝑠 and 𝜎 = 8. Uplift Decay Rule. For the 𝑖-th seller of user 𝑢 with segment 𝑠𝑖 , let 𝑛𝑠<𝑖 denote the number of occurrences of segment 𝑠 among the first 𝑖 − 1 sellers. The uplift decay is:  uplift𝑖new = uplift𝑖raw − min 𝑑 · 𝑛𝑠<𝑖 , 0.5 · |uplift𝑖raw | · sgn(uplift𝑖raw ) (29) where 𝑑 = 4 + 2(𝑛𝑠<𝑖 − 1) when 𝑛𝑠<𝑖 ≥ 1, otherwise 𝑑 = 0. This design ensures that repeated exposure to the same segment leads to cumulative decay, capped at 50% of the original uplift magnitude.

B

Training Procedure of CanniUplift

Table 4 summarizes the training procedure of CanniUplift, where candidate-specific representations are computed through TreatAttention and the seller-level, platform-level, and redemption-related heads are jointly optimized using the objective defined in Section 3.5.

∑︁

© ∑︁ pla pla ª L­ 𝑝𝐺𝑀𝑉𝑠𝑖 , 𝐺𝑀𝑉𝑢 ® 𝑢 «𝑠𝑖 ∈ S𝑢 ¬

(30)

However, at inference time, our objective is to evaluate the marginal impact on platform-wide GMV when changing the intervention status of a single seller 𝑠 𝑗 from control to treatment, while keeping all other sellers’ states unchanged.

C.2

Treatment Outcome. The treatment outcome is:

A.6

During training, the Platform head enforces global consistency through an aggregated loss over all candidate sellers:

Lpla =

𝑥𝑠class is the normalized seller categorical attribute, 𝜖0 ∼ N (0, 1), and 𝑧 0 is a cluster-specific offset sampled based on 𝑥 cluster from: N (0, 1), N (2, 0.5), N (−1, 2), N (3, 1.5), N (−2, 0.8), or N (1, 2). Coefficients are (𝑎, 𝑏, 𝑐, 𝑑) = (1, 1, 0.03, 0.5).

𝑦1 = 𝑦0 + 𝑒 · 1⊤ x𝑢 + 𝑓 · 1⊤ x𝑠 + 𝑔 · (x𝑢⊤ x𝑠 ) + 𝑧 1 + 𝜖1 − Δ

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

Practical Approximation for Marginal Scoring

Consider a user 𝑢 with candidate seller set S𝑢 . The platform-level marginal increment induced by treating seller 𝑠 𝑗 is defined as: pla

pla

ΔPlatform (𝑢, 𝑠 𝑗 ) = E[𝑌𝑢 | 𝑡 𝑗 =1] − E[𝑌𝑢 | 𝑡 𝑗 =0] ∑︁ ∑︁ = E[𝑦𝑢,𝑠 | 𝑡 𝑗 =1] − E[𝑦𝑢,𝑠 | 𝑡 𝑗 =0] 𝑠 ∈ S𝑢

(31)

𝑠 ∈ S𝑢

For practical deployment, we need to estimate the marginal platform-level effect of assigning a coupon to a target seller 𝑠 𝑗 for user 𝑢. A full counterfactual comparison would require recomputing the platform-wide GMV under two worlds, i.e., assigning and not assigning the treatment to 𝑠 𝑗 . For sellers 𝑠𝑘 ≠ 𝑠 𝑗 , their treatment states are fixed as background interventions during this comparison. However, their outcomes may still be affected by changing 𝑡 𝑗 , which is exactly the seller-level interference effect considered in this work. Exhaustively recomputing the counterfactual outcomes of all other sellers is computationally expensive at serving time. Therefore, we adopt a practical marginal scoring approximation: the Platform head is trained under the global aggregation constraint above, so its candidate-specific platform-view output is encouraged to absorb the expected net impact of treating 𝑠 𝑗 on the user’s total platform GMV. Concretely, we compute two platform-view predictions for the target seller 𝑠 𝑗 : pla

𝑝 treat = 𝑝𝐺𝑀𝑉𝑠 𝑗 (𝑢, 𝑠 𝑗 , 𝑡 𝑗 =1), pla

(32)

𝑝 control = 𝑝𝐺𝑀𝑉𝑠 𝑗 (𝑢, 𝑠 𝑗 , 𝑡 𝑗 =0). The platform-level marginal uplift is then approximated by their difference: Δpla (𝑢, 𝑠 𝑗 ) ≈ 𝑝 treat − 𝑝 control = 𝛿 platform (𝑢, 𝑠 𝑗 ).

(33)

This approximation should not be interpreted as assuming that other sellers’ outcomes are unaffected. Instead, cross-seller substitution effects are implicitly captured through the parameters learned from the platform-level aggregation loss and the candidate-specific user–seller representation.

KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea

Zuwang He et al.

Table 4: Training Procedure of CanniUplift Input: Training data D with user features, candidate sellers, seller-level GMV/redemption labels, and platform-level GMV. 1. For each user 𝑢, encode multi-source behaviors: 𝐻𝑢 , ℎ𝑢 ← Encoder(𝑥𝑢 ). 2. For each candidate seller 𝑠 ∈ S𝑢 of user 𝑢, compute candidate interaction: 𝑧𝑢,𝑠 ← TreatAttn(𝑠, 𝐻𝑢 ). 3. Predict control-side GMV: 𝑝𝐺𝑀𝑉𝑐 (𝑠) ← CtrlHead(𝑧𝑢,𝑠 ). pla 4. Predict platform-view GMV: 𝑝𝐺𝑀𝑉𝑠 ← PlaHead(𝑧𝑢,𝑠 ). 5. Predict redemption probability: 𝑝𝑟 (𝑠) ← RedemHead(𝑧𝑢,𝑠 ). 6. Predict path-specific increments: 𝑝Δ𝐺𝑀𝑉𝑟 (𝑠), 𝑝Δ𝐺𝑀𝑉1−𝑟 (𝑠) ← RDDHeads(𝑧𝑢,𝑠 ). 7. Combine the RDD prediction: 𝑝𝐺𝑀𝑉𝑠 ← 𝑝𝐺𝑀𝑉𝑐 (𝑠) + 𝑝𝑟 (𝑠)𝑝Δ𝐺𝑀𝑉𝑟 (𝑠) + (1 − 𝑝𝑟 (𝑠))𝑝Δ𝐺𝑀𝑉1−𝑟 (𝑠). Í pla pla 8. Aggregate platform prediction: 𝑝𝐺𝑀𝑉𝑢 ← 𝑠 ∈ S𝑢 𝑝𝐺𝑀𝑉𝑠 . 9. Compute Lseller , Lpla , and Lredem . 10. Update parameters by minimizing Ltotal .

C.3

Why Single-Point Predictions Capture Global Effects

The key insight is that the global aggregation constraint during training encourages the model to encode cross-seller substitution effects into the model parameters. Specifically: • Global loss propagates cross-seller dependencies: When pla the model overestimates 𝑝𝐺𝑀𝑉𝑠𝑖 for one seller during trainÍ pla ing, causing 𝑠 ∈ S𝑢 𝑝𝐺𝑀𝑉𝑠 to deviate from the true platform GMV, the gradient simultaneously corrects predictions for all related sellers. • User representation captures substitution patterns: The Treat-Attention mechanism introduced in Section 3.1 enables the user behavior sequence 𝐻𝑢 to capture cross-shop browsing patterns (e.g., “user alternates between similar-category sellers”). The model learns to infer global consumption elasticity from ⟨𝑢𝑠𝑒𝑟, 𝑠𝑒𝑙𝑙𝑒𝑟 ⟩ features. • Inference as a counterfactual query: The single-point score 𝛿 platform (𝑢, 𝑠 𝑗 ) effectively answers: “If we assign a coupon

to user 𝑢 at seller 𝑠 𝑗 , how will their total platform spending change?” Since the model has learned global dependencies, this score naturally reflects the net increment accounting for cannibalization.

C.4

Practical Inference Procedure

At inference time, for each user 𝑢 and candidate seller 𝑠 𝑗 : (1) Compute platform-view prediction with treatment: 𝑝 treat = PlatformHead(𝑧𝑢,𝑠 𝑗 , 𝑡 𝑗 =1) (2) Compute platform-view prediction without treatment: 𝑝 control = PlatformHead(𝑧𝑢,𝑠 𝑗 , 𝑡 𝑗 =0) (3) Calculate marginal platform uplift: Δplatform (𝑢, 𝑠 𝑗 ) = 𝑝 treat − 𝑝 control (4) Rank sellers by Δplatform and allocate top-K coupons Crucially, inference requires only single-point predictions for the target seller 𝑠 𝑗 , without aggregating over other sellers in S𝑢 . This design significantly reduces computational cost while preserving cannibalization-aware estimation.

Record · ID 343482 · SHA-256 f9d18dc0dea138bf
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.