Conceptio › Archive › arXiv CS
arXiv CSopen access

Conformal Policy Learning with Distribution-Free Safety Guarantees

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Conformal Policy Learning with Distribution-Free Safety Guarantees Ying Jin∗1 and Naoki Egami2 1

Department of Statistics and Data Science, University of Pennsylvania Department of Political Science & Statistics and Data Science Center, Massachusetts Institute of Technology

arXiv:2609.17296v1 [stat.ME] 15 Sep 2026

2

Abstract Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of “do no harm.” In this paper, we propose conformal policy learning (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs. Keywords: AI safety, Causal inference, Conformal prediction, Policy learning

1

Introduction

Policy learning, also known as the treatment choice problem, aims to learn a rule that automatically assigns future treatment options based on individual characteristics (Manski, 2004; Hirano and Porter, 2009; Kitagawa and Tetenov, 2018; Athey and Wager, 2021). It has been the foundation for data-driven decision-making in various domains spanning precision medicine (Murphy, 2003; Qian and Murphy, 2011; Zhao et al., 2012), online advertising and recommendation systems (Li et al., 2010; Dudı́k et al., 2011), political campaigns (Imai and Strauss, 2011), and criminal justice (Kleinberg et al., 2018), among others. Policy learning methods often aim to maximize the average welfare (expectation of the realized outcome) within a policy class (Manski, 2004; Kitagawa and Tetenov, 2018; Athey and Wager, 2021). While welfare maximization is widely useful, it can be insufficient in high-stakes domains such as medicine and public policy where individual safety is of concern. In such applications, decision makers may seek not only on-average improvement of the outcomes, but also controlling the number of individuals harmed by the intervention (Gadbury et al., 2004; Kallus, 2022; Richens et al., 2022; Ben-Michael et al., 2025), i.e., “do no harm.” As an example, suppose a new intervention benefits 60% of a group while harming the remaining ∗ The reproduction code is in the GitHub repository https://github.com/ying531/conformal-policy-learning. Email: [email protected]

1

40% by the same magnitude. A decision-maker who maximizes average welfare would decide to treat the group; this would harm 40% of the population, which can be unacceptable if individual safety is of primary concern. In this paper, we study policy learning with a distribution-free safety guarantee. Formally, assume access to observed data {(Xi , Ti , Yi )}ni=1 where Yi ∈ {0, 1} is a binary outcome, Ti ∈ {0, 1} is the binary treatment, and Xi ∈ X is the observed features for each unit i. Under the potential outcome framework with SUTVA (formalized in Section 2.1), the outcome is Yi = Yi (Ti ), where (Yi (1), Yi (0)) are the potential outcomes under treatment and control, respectively. For a new test point with observed features Xn+1 and unknown potential outcomes (Yn+1 (0), Yn+1 (1)), our goal is to learn a policy π̂ : X → {0, 1} that maps features to treatment assignments with the following safety guarantee: for a pre-specified level α ∈ (0, 1),  P Yn+1 (π̂(Xn+1 )) < Yn+1 (0) ≤ α. (1.1) This safety guarantee (1.1) means the probability of the realized outcome being worse than the “status-quo” outcome under control is no greater than α, thereby limiting the risk of harming the new individual. When this policy is implemented on m new individuals, (1.1) implies that the expected number of units harmed by the treatment is no greater than mα. Such guarantees are important for two related reasons. First, negative treatment effects are substantively costly in many settings. Second, treatment rules are increasingly embedded in (semi-)automated decision systems: once a learned rule is deployed, it may assign interventions repeatedly and at scale. In such settings, trust in the system requires explicit control of the probability that it harms the individuals it chooses to treat. We highlight three applications that motivate such explicit harm-rate control. Example 1 (AI-powered Interventions): Generative AI is emerging as a new class of intervention in the social sciences, with applications designed to change attitudes and behaviors through scalable, personalized interactions (e.g., Costello et al., 2024; Bai et al., 2025). At the same time, recent empirical studies highlight an important risk: while such AI interventions may benefit many individuals and tasks, they may also harm others. For example, a randomized experiment in a global consulting firm (Dell’Acqua et al., 2026) found that access to AI can harm the productivity of high-skilled consultants when working on challenging intellectual tasks. Similarly, Bastani et al. (2025) found that generative AI without guardrails can harm the learning of high school math students. As interventions are powered by black-box AI, controlling the harm is fundamental for safety, public trust, and efficient deployment of AI-powered treatment. Example 2 (Precision Medicine): Precision medicine—prevention and treatment strategies that account for individual variability—is central in modern medicine (Kosorok and Laber, 2019). The problem of controlling individual risk while maximizing benefit has long been recognized. It is especially relevant when efficacious medications may also lead to a higher risk for certain individuals, such as opioid treatment of chronic pain (Laber et al., 2018) and type-2 diabetes with insulin therapies (Wang et al., 2018). Example 3 (Criminal Justice): In the US criminal justice system, how to safely use risk scores to evaluate an individual’s likelihood of reoffending and identify their criminogenic needs is a major topic of interest (Skeem et al., 2020). For example, Ben-Michael et al. (2025) analyze a field experiment to estimate the causal effect of algorithmic recommendations on judges’ decisions at a criminal first appearance hearing. Here, a treatment rule with the safety guarantee can prevent arrestees from committing a new crime or failing to appear in court, while avoiding unnecessarily harsh decisions. We aim to achieve the safety guarantee in a model-agnostic fashion—meaning that it holds without strong modeling assumptions on the data distribution or the learning algorithms—and tightly in finite samples, so this framework is widely applicable to various high-stakes settings.

1.1

Overview of contributions

We develop conformal policy learning (CPL) to achieve the safety guarantee (1.1). Distinct from standard policy learning methods that maximizes empirical welfare within a policy class, our starting point is to view 2

Hidden safety target

Observable proxy calibration

Labeled data (X! , %! , &! ) Test unit ("#$ Harm &"#$ 1 < &"#$ 0 ?

Conformal p-value & Safe treatment 0 (! , &!%

0 ("#$ , 0

-$ = 1

Treat iff $!"# ≤ #

selected calib data Proxy &!% = %! &! + (1-%! )(1-&! )

!!"# =

1 + ∑ &$ '$ 1{) *$ , ,$%

,$% = 0: possible harm

< ) *!"# , 0 } 1 + ∑&$ '$

P(treated and harmed) ≤ #

Figure 1: Workflow of Conformal Policy Learning (CPL). Given labeled data {(Xi , Ti , Yi )}ni=1 , the goal is to assign safe treatment for a new test unit Xn+1 with harm rate target α ∈ (0, 1). CPL constructs proxy label Yi† and selects a subset of data for calibration. It then computes a p-value by comparing a test conformity score to the calibration scores. The treatment is determined by whether the conformal p-value is below α, which controls the harm rate below α. !$% = *

!$ , 1 − !$ ,

1"#$ =

'$ = 1 '$ = 0

Conformity score

! % = !$ '$ + 1 − !$ 1 − '$

+

Calibration data X! , %! , &! : .! = 1

$ Selective each treatment decision as testing a counterfactual harm event Yn+1 (1) < Yn+1 (0). A p-value pn+1 satisfying calibration  P Yn+1 (1) < Yn+1 (0), pn+1 ≤ α ≤ α

yields the desired safety guarantee by assigning treatment only when pn+1 ≤ α. This connects safe policy learning to conformal inference (Vovk et al., 2005; Lei and Candès, 2021; Jin and Candès, 2023b). The key challenge here—which necessitates novel constructions of powerful conformal p-values—is that the “harm” event involves both potential outcomes and is therefore unobserved even for the labeled data. CPL addresses this difficulty through two techniques. First, it uses observable proxy labels to construct valid conformal p-values. Second, to reduce conservativeness, it selects the labeled observations used in pvalue calibration according to which treatment arm provides a sharper proxy for harm. We call this strategy “selective calibration.” The resulting p-value compares the test conformity score with the selected calibration scores computed using the proxy labels. The CPL workflow is in Figure 1. For randomized experiments, CPL provides finite-sample, distribution-free safety guarantees. Under the exchangeability conditions ensured by random assignment, we show that CPL controls the harm rate exactly below α in finite samples, without making any assumption about the outcome models used in the conformal p-values and the selective calibration rule. We first establish this result for balanced randomized experiments and then extend it to stratified experiments using weighted conformal inference (Tibshirani et al., 2019). We next study the sharpness and optimality of CPL. Because individual-level harm, Y (1) < Y (0), is not identifiable from observed data, we formulate the optimality under partial identification (Kallus, 2022; Li et al., 2023). We characterize the optimal policy that maximizes the power (probability of treatment) and average welfare subject to the safety constraint for every joint potential-outcome distribution compatible with the observed data. Then, we show that as long as the score function used in the p-values and selective calibration rule converge to oracle ones, CPL attains these global optima asymptotically. We further extend CPL to observational studies. In this setting, the treatment assignment mechanism is unknown, and the calibration weights must be estimated. We develop a learn-then-balance procedure that estimates these weights and enforces balance on functions tailored to the thresholding decisions and the partially identified harm target. Under suitable regularity conditions, the resulting policy has a doubly robust asymptotic safety guarantee: the excess harm rate vanishes when either the propensity-score model or the outcome models, but not necessarily both, are consistently estimated. Moreover, if both components converge at standard slow nonparametric rates, the excess harm rate is of a parametric order. With consistent outcome models, CPL with observational data attains the same welfare and power optima as in the randomized case. Our excess harm rate bound does not pay the price of policy class complexity common in policy learning. Finally, we evaluate CPL through simulations and an empirical study. In simulations across randomized experiments, stratified experiments, and observational studies, CPL shows robust harm-rate control and high power and welfare. In the empirical study, CPL assigns an AI-powered intervention designed to reduce conspiracy beliefs while tightly controlling the harm rate. The rest of the paper is organized as follows. Section 2 introduces the problem setup and connects 3

the safety guarantee to conformal hypothesis testing. Section 3 develops CPL for balanced randomized experiments and studies its finite-sample validity and asymptotic optimality; the framework is then extended to stratified experiments in Section 4 and observational studies in Section 5. Section 6 presents simulation studies, and Section 7 applies CPL to an empirical study of AI-powered interventions. We close the paper with a discussion on extensions and future directions in Section 8.

1.2

Related Work

This article lies in the intersection of causal inference, policy learning, and conformal inference. We summarize several important lines of related work below. Our work is connected to the established literature on policy learning (e.g., Murphy, 2003; Manski, 2004; Hirano and Porter, 2009; Zhao et al., 2012; Kitagawa and Tetenov, 2018; Athey and Wager, 2021; Jin et al., 2025b), which aims to select, among a given class of policies, the one that maximizes an objective (such as average welfare) while optionally respecting a constraint (such as harm rate). Within this literature, this work is closely related to a small but emerging literature on policy learning with safety considerations (where the exact meaning of safety varies). One line of work develops safe policy learning algorithms in settings that necessitate extrapolating beyond the observed labeled data, and the safety refers to not performing worse than a “status quo” policy (Zhang et al., 2022; Ben-Michael et al., 2025; Jia et al., 2025; Wu et al., 2025). Our safety notion of individual harm is conceptually related since it measures the harm relative to the status quo of no treatment, but distinct enough to yield completely different techniques. In addition, Ben-Michael et al. (2024) studies policy learning when the objective involves counterfactuals, providing doubly robust algorithms and regret bounds for the learned policy; the dual form of a special case in their framework (with an unknown Lagrange parameter) coincides with the welfare maximization problem subject to harm rate control; this connects with our setting and the method of Li et al. (2023). As standard in policy learning, these methods select a policy that maximizes an empirical objective within a policy class, whose performance is often measured by the regret (the gap between the true objective of the learned policy from the optimal), which is typically bounded by a statistical error term that scales with the complexity of the policy class (such as the VC-dimension). In contrast, CPL leverages conformal inference to achieve finite-sample, distribution-free safety guarantee (without a high-probability excess error bound term) when propensity scores are known; moreover, when propensity scores are unknown and estimated so inexact harm control is inevitable, our excess harm rate bound does not pay the price of the policy class complexity. Our safety notion follows from a literature in causal inference and policy learning that bounds, estimates, and controls the same harm rate notion as (2.1). A line of work studies its bounds under various assumptions such as monotonicity (Huang et al., 2012) and certain conditional independence conditions (Shen et al., 2013; Yin et al., 2018). Relatedly, due to the non-identifiability, several work focuses on establishing bounds on the conditional harm rate for binary outcomes, including Gadbury et al. (2004) without covariates, Zhang et al. (2013) with covariates, Kallus (2022) on the sharp identification bounds, and Wu et al. (2024) using a sensitivity model for the correlation between the potential outcomes. In addition, a recent independent work of Scauda et al. (2026) studies the population-level optimal welfare subject to harm rate control. Since we aim to control the harm rate without strong modeling assumptions, this work is implicitly tied to the Frechét–Hoeffding bound characterized in Zhang et al. (2013); Kallus (2022). We show CPL attains the optimal welfare subject to safety guarantee among the worst-case distributions in these work. Our method builds on the conformal p-values proposed in Jin and Candès (2023b) for i.i.d. data and Jin and Candès (2023a) for covariate shift settings, which were extended to model selection (Bai and Jin, 2024) and online settings (Xu and Ramdas, 2024). In that literature, the p-values quantify the confidence in a large, ordinary outcome (i.e., no potential outcomes) exceeding a known threshold. While we borrow the high-level intuitions, the technical route in constructing such p-values for our problem—which is the key contribution here—is sharply distinct since the two outcomes under comparison are never simultaneously observed. The resulting optimality and robustness properties are likewise quite different. The conformal inference approach also connects our work with a line of work on conformal inference for individual treatment effects (ITE) Yn+1 (1) − Yn+1 (0) (Lei and Candès, 2021; Jin et al., 2023; Yin et al., 4

2024). These work typically focuses on constructing a prediction set for the counterfactual outcome for a unit who has already received treatment or control, where inference for the ITE of a new test point with two unknown outcomes appears particularly challenging. The latter case (with binary outcomes) is the setting we address, and we study the “decision” problem rather than prediction set construction. This involves a thresholding decision rule which necessitates distinct calibration and theoretical analysis techniques.

2

Problem Setup and Conceptual Framework

2.1

Problem Setup

We assume access to a set of labeled data {(Xi , Ti , Yi )}ni=1 , where Xi ∈ X is the feature, Ti ∈ {0, 1} is the treatment assignment, and Yi ∈ {0, 1} is the binary outcome. We define the potential outcomes Yi (t) for t ∈ {0, 1} and assume the triplets {(Xi , Yi (1), Yi (0))}ni=1 are independent and identically distributed (i.i.d.) from an unknown super-population PX,Y (1),Y (0) . Assuming the Stable Unit Treatment Value Assumption (SUTVA) (Rubin, 1980), the observed outcome is given by Yi = Yi (Ti ). The joint super-population PX,Y (1),Y (0) is unidentifiable from data because researchers observe only one potential outcome for each unit (Holland, 1986). Finally, the distribution of the labeled data {(Xi , Ti , Yi )}ni=1 is induced by the unknown super-population and the treatment assignment mechanism and denoted as PX,Y,T . Throughout, we make the unconfoundedness assumption for the treatment assignment mechanism, which is standard in the causal inference literature (Imbens and Rubin, 2015). Assumption 2.1 (Unconfoundedness). The treatment assignments {Ti }ni=1 are mutually independent and obey Ti ⊥ ⊥ (Yi (0), Yi (1)) | Xi , with P(Ti = 1 | Xi = x) := e(x), and we call e(x) ∈ (0, 1) the propensity score. Our framework covers both randomized experiments (Sections 3 and 4) and observational studies (Section 5). In randomized experiments, Assumption 2.1 is satisfied by design and e(x) is known. In observational studies, Assumption 2.1 should be evaluated based on domain knowledge; even though it holds, the propensity score e(x) is unknown and needs to be estimated. Policy learning uses the labeled data to determine the treatment for a new individual (which we call the test unit/point) with observed feature Xn+1 and unobserved potential outcomes Yn+1 (1) and Yn+1 (0). Following the literature (e.g., Manski, 2004; Hirano and Porter, 2009; Zhao et al., 2012; Kitagawa and Tetenov, 2018; Athey and Wager, 2021), we assume it is from the same super-population independently of the labeled data. Assumption 2.2 (IID). The test point is drawn from (Xn+1 , Yn+1 (1), Yn+1 (0)) ∼ PX,Y (1),Y (0) independently of the labeled data. Our framework can be naturally extended to relax the i.i.d. assumption and allow for covariate shift between the labeled data i ∈ {1, . . . , n} and the test point; see Section 8 for a discussion.

2.2

Harm Rate

Recall that our goal is to develop a policy learning algorithm producing a rule π̂ : X → {0, 1} that maps features X to treatment assignment π̂(X) with the counterfactual safety guarantee:  P Yn+1 (π̂(Xn+1 )) < Yn+1 (0) ≤ α (2.1) for a pre-specified confidence level α ∈ (0, 1). Here, the probability is over the labeled data (based on which π̂ is built) and the test point. We call the left-handed side probability in (2.1) the “harm rate” following Zhang et al. (2013); the same quantity is also referred to as the “fraction of negatively affected” in the literature (Kallus, 2022; Li et al., 2023; Wu et al., 2024). Throughout, we focus on binary outcomes; this setting is most extensively studied (Zhang et al., 2013; Kallus, 2022; Wu et al., 2024).

5

The safety guarantees can be rewritten as   P Yn+1 (π̂(Xn+1 )) < Yn+1 (0) ≤ α ⇐⇒ P Yn+1 (1) < Yn+1 (0), π̂(Xn+1 ) = 1 ≤ α.

(2.2)

Thus, harm occurs when the individual treatment effect on the test unit is negative and one decides to treat the unit. Thus, a trivial way to achieve this is to treat the test unit with probability α; however, such a policy is clearly suboptimal due to limited welfare and power (whose meaning will be made precise later). Instead, we aim for a procedure that achieves this with reasonable, and sometimes optimal, welfare and power. Meanwhile, the above result shows that to control the harm rate, we should detect and avoid treating units that can be harmed by the treatment. This is the intuition for CPL.

2.3

Connecting Safety Guarantees to Hypothesis Testing

Our technical route differs from standard policy learning approaches, so it is helpful to begin with the conceptual framework: connect the safety guarantees with hypothesis testing. Equation (2.2) suggest that to make safe treatments, one needs to identify units that are likely not to be harmed by the treatment, i.e., Yn+1 (1) ≥ Yn+1 (0). We view this as testing a random null hypothesis H0 : Yn+1 (1) < Yn+1 (0). Suppose we can construct a p-value pn+1 such that  P Yn+1 (1) < Yn+1 (0), pn+1 ≤ α ≤ α. (2.3) This is a non-conventional definition of p-values because the truth of H0 is itself random, but this notion is sufficient for the desired safety guarantee: setting π̂(Xn+1 ) = 1{pn+1 ≤ α}, the validity (2.3) directly implies the safety guarantee (2.1) via the equivalence relationship in (2.2). Intuitively, a small p-value informs strong evidence against the null, i.e., the test unit is unlikely to be harmed and thus can safely receive the treatment. The remaining of the paper focuses on addressing two questions. First, how can we construct p-values that satisfy (2.3) in finite sample without making strong modeling assumptions? We address this question by expanding the recent literature of conformal p-values (Vovk et al., 2005; Bates et al., 2021; Jin and Candès, 2023b) to causal inference contexts (Imbens and Rubin, 2015; Lei and Candès, 2021), while respecting the partial-identification nature of the individual-level causal effects (e.g., Heckman et al., 1997). Second, is thresholding p-values optimal in any sense? This is important as it is not clear a priori why our approach may be desired, even if it might achieve the safety guarantee. We will show that with specifically designed p-values, our method achieves optimal power and welfare among all safe policies at the population level.

3

Conformal Policy Learning with Balanced Randomized Experiments

To fix ideas, in this section, we introduce our framework in the simple setting of balanced randomized experiments with e(x) = 1/2; this will be gradually generalized in Sections 4 and 5. In Section 3.1, we start with a single-arm construction that provides distribution-free safety guarantees. In Section 3.2, we introduce observable proxies and selective calibration that sharpens the procedure while preserving the same guarantee. In Section 3.3, we characterize the welfare-optimal policy under partial identification, followed by an explanation on the sharpness of selective calibration in Section 3.4. Throughout this section, all learned functions are fitted on a training sample independent of the labeled observations used for calibration and the test unit. In practice, this can be achieved by sample splitting (Lei et al., 2018). For notational simplicity, we use {(Xi , Ti , Yi )}ni=1 to denote the labeled observations available for calibration; selective calibration introduced in Section 3.2 may retain only a subset of these observations. Conditional on the training sample, the learned functions are treated as fixed.

3.1

Warm-up: Direct Application of Existing Conformal p-values

We begin with a simple adaptation of the existing idea to show how conformal p-values can serve as the foundation for safe policy learning. Given that we focus on binary outcomes, we can rewrite (2.3) as finding 6

a p-value pn+1 obeying  P Yn+1 (1) = 0, Yn+1 (0) = 1, pn+1 ≤ α ≤ α.

(3.1)

That is, we would like to find p-values that quantify the confidence in a small treated outcome and a large control outcome. The conformal selection (CS) framework (Jin and Candès, 2023b) provides a natural starting point. Given labeled data {(Xi , Yi )}ni=1 for an ordinary outcome Y ∈ R (i.e., no potential-outcome framing) and a known threshold c ∈ R, CS proposes conformal p-values with a similar null property P(Yn+1 ≤ c, pn+1 ≤ α) ≤ α. Setting aside the issue that we now have two potential outcomes to deal with, one simple (1) idea is to develop a conservative p-value pn+1 that satisfy  (1) P Yn+1 (1) = 0, pn+1 ≤ α ≤ α,

(3.2)

which directly implies (3.1). For treated outcome, it means we will use the treated units for calibrating p-values. The conformal p-value from CS takes the form Pn 1 + i=1 Ti · 1{V (Xi , Yi ) ≤ V (Xn+1 , 0)} (1) Pn pn+1 = , (3.3) 1 + i=1 Ti where V is any score function V : X × Y → R obeying V (x, y) ≤ V (x, y ′ ) whenever y ≤ y ′ for any x ∈ X . Example (clipped score function). One example is the clipped score function (Jin and Candès, 2023b), V (x, y) = M 1{y > 0} + (1 − µ̂1 (x)), where µ̂1 (x) is an estimator for the conditional expectation function E[Yi | Ti = 1, Xi = x] and M > 2 supx |µ̂1 (x)| is a sufficiently large constant so that (3.3) reduces to Pn 1 + i=1 Ti · 1{Yi = 0} · 1{µ̂1 (Xi ) ≥ µ̂1 (Xn+1 )} (1) Pn pn+1 = . 1 + i=1 Ti Intuitively, this p-value focuses on potentially unsafe cases (i.e., Yi (1) = 0) among treated units and examines whether the test unit is extreme with respect to µ̂1 (x). When this p-value is small, the test unit is unlikely to come from the distribution of potentially unsafe units, and thus is likely to be safe. To see why the conformal p-value (3.3) satisfies (3.2) for any score function, we outline the theoretical arguments from Jin and Candès (2023b). Let n1 = |I1 | be the number of treated units, where I1 = {i ∈ [n] : Ti = 1}. Consider the “oracle” p-value Pn 1 + i=1 Ti · 1{V (Xi , Yi (1)) ≤ V (Xn+1 , Yn+1 (1))} ∗(1) pn+1 = (3.4) 1 + n1 which is not computable since Yn+1 (1) is unobserved. The only difference between (3.3) and (3.4) is that Yn+1 (1) in the oracle p-value is replaced by 0, and we have Yi (1) = Yi for i ∈ I1 in (3.4). First, conditional on {Ti }ni=1 , in this randomized experiment, the data {(Xi , Yi (1)}i∈I1 ∪ {Xn+1 , Yn+1 (1)} are exchangeable, ∗(1) which implies P(pn+1 ≤ α) ≤ α (Vovk et al., 2005). Second, on the event that Yn+1 (1) = 0, the two p-values (1) ∗(1) ∗(1) coincide. The two facts imply P(pn+1 ≤ α, Yn+1 (1) = 0) = P(pn+1 ≤ α, Yn+1 (1) = 0) ≤ P(pn+1 ≤ α) ≤ α. (0)

Of course, one may also leverage the control samples and consider the null hypotheses H0 : Yn+1 (0) = 1. We can similarly construct the p-value (using a monotone function V as before) Pn 1 + i=1 1{Ti = 0} · 1{V (Xi , 1 − Yi ) ≤ V (Xn+1 , 0)} (0) Pn pn+1 = . 1 + i=1 1{Ti = 0} Following exactly the same rationales as above, this p-value is valid in the same sense.

7

3.2

Proposed Method: Conformal Policy Learning (1)

While both p-values in the last section are feasible and valid, they are conservative since H0 : Yn+1 (1) = 0 (0) and H0 : Yn+1 (0) = 1 are both strict implications of H0 : Yn+1 (1) = 0, Yn+1 (0) = 1. In this section, we propose the general CPL procedure that improves and subsumes the previous two options as special cases. Following the conformal inference literature, we call the subset of the labeled data used to compute the p-values the “calibration data”. Then, both p-values in Section 3.1 only use either the treatment or the control group as the calibration data, which can be substantially generalized. First, consider any inclusion function ĝ : X × {0, 1} → [0, 1] whose training process is independent of the labeled data and the test point. Conditional on data, we draw independent inclusion indicators Gi ∼ Bernoulli(ĝ(Xi , Ti )) for each i ∈ [n], which defines the calibration set Dcalib = {(Xi , Ti , Yi )}i∈Icalib ,

where

Icalib = {i ∈ [n] : Gi = 1}.

(3.5)

The inclusion into Dcalib can depend both on Ti and Xi , and is allowed to be random, although later we shall see a binary ĝ suffices for optimality. Since not all labeled data are used in calibration, we call this technique “selective calibration”. Second, to allow involving data from both arms in calibration, we combine information from the two treatment arms. Define the computable proxy outcome Yi† = Ti Yi + (1 − Ti )(1 − Yi ), so Yi† = 0 whenever Yi = 0 for a treated sample or Yi = 1 for a control sample. Generalizing the previous ideas, Yi† = 0 captures potentially unsafe units (i.e., units for which H0 might happen). With the two techniques in hand, we construct the p-value Pn 1 + i=1 Gi × 1{V (Xi , Yi† ) ≤ V (Xn+1 , 0)} Pn pn+1 = . 1 + i=1 Gi (1)

(3.6)

(0)

This covers the single-arm p-values in Section 3.1: it recovers pn+1 (resp. pn+1 ) when Gi = Ti (resp. Gi = 1 − Ti ). Another simple case is to take Gi = 1 which includes all the labeled data as the calibration set. With this p-value, our conformal policy learning simply produces a thresholding rule π̂(Xn+1 ) := 1{pn+1 ≤ α}. Algorithm 1 summarizes the procedure. Algorithm 1 Conformal Policy Learning (Randomized Experiments) Input: Labeled data {(Xi , Ti , Yi )}n i=1 , test data Xn+1 , target level α ∈ (0, 1), score V (·, ·), inclusion function ĝ(·, ·). 1: Draw independent inclusion indicators Gi ∼ Bern(ĝ(Xi , Ti )) for i = 1, . . . , n. † 2: Compute Yi = Ti Yi + (1 − Ti )(1 − Yi ) for i = 1, . . . , n. 3: Compute p-value pn+1 via (3.6). 4: Compute Tn+1 = 1{pn+1 ≤ α}. Output: Safe treatment Tn+1 .

Theorem 3.1 establishes the distribution-free safety guarantee for any score and inclusion functions whose inclusion probability in two groups sums up to a constant. The proof is in Appendix B.1. Theorem 3.1. Assume Assumption 2.1 with e(x) ≡ 0.5 and Assumption 2.2. Then, for any score function V : X × Y → R that is non-decreasing in the second argument and ĝ : X × {0, 1} → [0, 1] whose training process is independent of {(Xi , Ti , Yi )}ni=1 ∪ {Xn+1 }, and obey ĝ(x, 1) + ĝ(x, 0) = a for a constant a > 0 for any x ∈ X , it holds that P(Yn+1 (1) = 0, Yn+1 (0) = 1, pn+1 ≤ α) ≤ α for any α ∈ (0, 1). That is, the finite-sample safety guarantee (2.1) holds for π̂(Xn+1 ) = 1{pn+1 ≤ α}. Theorem 3.1 states that, in principle, users can use any score function V (·, ·) and any inclusion function ĝ(·, ·), and CPL (Algorithm 1) guarantees a controlled harm rate in finite samples. We emphasize that the two simplifying assumptions, namely e(x) = 1/2 and ĝ(x, 1) + ĝ(x, 0) = a, are used only to streamline presentation, and will be eliminated in Section 4. 8

This theorem covers the special cases we discussed in Section 3.1. The one based on the treated (resp. control) samples correspond to ĝ(x, t) = t (resp. ĝ(x, t) = 1 − t), both obeying the constant-sum condition with a = 1. As a third example, the conformal p-value using all the labeled data corresponds to ĝ(x, t) ≡ 1 and thus it also satisfies the constant-sum condition with a = 2. Finally, while (3.6) represents a restrictive class of policies, we shall see in the next subsection that specific choices of V and ĝ lead to asymptotically optimal safe policy learning algorithm among all safe policies (i.e., beyond the policy class specified by CPL). The detailed proof of Theorem 3.1 is in Appendix B.1, and we provide a sketch here to clarify the intuitions. We again consider the “oracle” p-value Pn ∗ } 1 + i=1 Gi × 1{V (Xi , Yi∗ ) ≤ V (Xn+1 , Yn+1 ∗ Pn , (3.7) pn+1 = 1 + i=1 Gi where Yi∗ = max{Yi (1), 1 − Yi (0)} so that Yi∗ = 0 if and only if Yi (1) = 0 and Yi (0) = 1, i.e., unit i is harmed. This is again not computable because Yi∗ is never observed even for the labeled data. The first key fact is that this oracle is valid as long as ĝ satisfies the constant-sum condition. Indeed, we can show ∗ that in balanced randomized experiments, the data {(Xi , Yi∗ )}Gi =1 and (Xn+1 , Yn+1 ) are exchangeable n ∗ n conditional on {Gi }i=1 , which implies P(pn+1 ≤ α | {Gi }i=1 ) ≤ α. Second, the observed p-value (3.6) ∗ replaces Yn+1 in (3.7) with the “null” value 0 and replaces the calibration label Yi∗ with Yi† , thereby upper bounding the oracle one on the “unsafe” null event. Specifically, the proxy outcome is conservative in the ∗ sense that Yi∗ ≥ Yi† ; thus, whenever Yn+1 = 0, the monotonicity of V implies pn+1 ≥ p∗n+1 . This gives ∗ n ∗ P(Yn+1 = 0, pn+1 ≤ α | {Gi }i=1 ) ≤ P(Yn+1 = 0, p∗n+1 ≤ α | {Gi }ni=1 ) ≤ P(p∗n+1 ≤ α | {Gi }ni=1 ) ≤ α. Remark 3.2 (Sharpness of conformal p-values). Our conformal p-value is subject to two sources of conservativeness compared with p∗n+1 : (i) the “oracle” labels Yi∗ are replaced by the proxy labels Yi† , and (ii) the ∗ test label Yn+1 is replaced by the threshold 0. As in conformal selection (Jin and Candès, 2023b), the issue (ii) will be addressed by a tailored “clipped” score function. On the other hand, (i) is rooted in the partial identification nature of the harm rate, and specific choices of ĝ makes our p-value sharp; we shall discuss this in more details once the optimality results in the next section are ready.

3.3

Optimality of Conformal Policy Learning

Having established the finite-sample safety guarantees for CPL, we proceed to study the optimality. We first analyze the optimal policy among all safe policies (i.e., including policies beyond our framework) under partial identification. We then show that CPL with specific choices of the score and inclusion functions asymptotically achieves global optimality. To formalize the discussion, we define some additional notations. The unknown, true joint distribution over (X, Y (1), Y (0)) is denoted as PX,Y (1),Y (0) , which induces the observable distribution (Xi , Yi , Ti ) ∼ PX,Y,T . Also, we denote the collection of distributions P over (X, Y (1), Y (0)) that are compatible with the observable distribution PX,Y,T as  P = PX,Y (1),Y (0) : PX,Y (T ),T = PX,Y (T ),T for P (T = 1 | X, Y (1), Y (0)) = e(X) , that is, the induced distribution of (X, Y, T ) under P coincides with PX,Y,T on the observables. For any distribution P and any treatment assignment rule π : X → {0, 1}, we write the harm rate as Err(π; P ) := P (Yn+1 (π(Xn+1 )) < Yn+1 (0)) = P (Yn+1 (1) = 0, Yn+1 (0) = 1, π(Xn+1 ) = 1). By the tower property, letting EP denote the expectation under P , we know   Err(π; P ) = EP π(Xn+1 ) · fP (Xn+1 ) ≤ α, where fP (x) := P (Yn+1 (1) < Yn+1 (0) | Xn+1 = x). Here we call f (x) the harm rate function. Similar to the scalar marginal harm rate Err(π, P ), the function fP (·) is not identifiable from data because Y (1) and Y (0) are never simultaneously observed. Kallus (2022) derived the sharp partial-identification upper bound for fP (x), given by γ(x) := min{1 − µ1 (x), µ0 (x)}, 9

(3.8)

where µt (x) = E[Y (t) | X = x] for t ∈ {0, 1} is the conditional mean function of each potential outcome. Namely, for any super-population PX,Y (1),Y (0) whose observed distribution PX,Y,T (induced by the same treatment assignment mechanism) is equal to PX,Y,T , its harm rate function obeys fP (x) ≤ γ(x), and there exists one such distribution whose harm rate function coincides with γ(x). The most common objective of policy learning is to maximize the welfare. Following the policy learning literature (e.g., Murphy, 2003; Manski, 2004; Hirano and Porter, 2009; Zhao et al., 2012; Kitagawa and Tetenov, 2018; Athey and Wager, 2021), we define the welfare of a policy π : X → {0, 1} as   Welfare(π; P) := E Yn+1 (π(Xn+1 )) , where the larger value of outcome Y corresponds to the larger welfare. The welfare-optimal treatment rule subject to the safety guarantee can be defined as ∗ πwelfare = argmax

Welfare(π; P)

(3.9)

π : X →{0,1}

max Err(π; P ) ≤ α.

subject to

P ∈P

Here, we optimize the welfare subject to the worst-case safety guarantee. The constraint, maxP ∈P Err(π; P ) ≤ α, ensures that, regardless of the underlying data-generating process, the harm rate is upper bounded by α. We define the oracle welfare-optimal score function swelfare (x) = −(µ1 (x) − µ0 (x))/γ(x),

(3.10)

where we use the convention 0/0 = 0, a/0 = +∞ if a > 0 and a/0 = −∞ if a < 0 throughout. The ∗ under the worst-case safety constraint. It is following theorem formally gives the optimal solution πwelfare an implication of the more general Theorem A.5 in Appendix A.2. Theorem 3.3. Assume swelfare (X) has no point mass. Then an optimal solution to (3.9) is ∗ πwelfare (x) = 1{swelfare (x) ≤ r∗ },

where

r∗ := sup {r̃ ≤ 0 : E[γ(X)1{swelfare (X) ≤ r̃}] ≤ α} .

∗ Moreover, writing τ (x) = µ1 (x)−µ0 (x), if E[γ(X)1{τ (X) > 0}] ≤ α, then r∗ = 0 and πwelfare (x) = 1{τ (x) > ∗ ∗ 0} almost surely. Otherwise, r < 0 and E[γ(X)πwelfare (X)] = α.

Theorem 3.3 shows that the optimal treatment rule for the welfare maximization is based on a cutoff on swelfare (Xn+1 ) while staying in the region τ (Xn+1 ) ≥ 0. The intuition is as follows. We can show that the ∗ problem (3.9) that defines πwelfare can be rewritten as maximize

E[π(Xn+1 )τ (Xn+1 )] + constant

subject to

E[π(Xn+1 )γ(Xn+1 )] ≤ α,

π : X →{0,1}

where τ (x) = µ1 (x) − µ0 (x) captures the conditional average treatment effect (CATE). The optimal solution thus seeks the feature regions that yield the largest welfare gain τ (X) relative to each unit of “cost” γ(X). The next question is whether, and under what conditions, the conformal policy learning method can achieve this optimal welfare. Following Theorem 3.3, we define the conformal p-value Pn 1 + i=1 Gi × 1{Yi† = 0} × 1{ŝ(Xi ) ≤ ŝ(Xn+1 )} welfare Pn , (3.11) pn+1 = 1 + i=1 Gi for the specific choices Gi = Ti 1{1 − µ̂1 (Xi ) ≤ µ̂0 (Xi )} + (1 − Ti )1{1 − µ̂1 (Xi ) > µ̂0 (Xi )}, 10

(3.12)

and

ŝ(x) = −(µ̂1 (x) − µ̂0 (x))/γ̂(x).

(3.13)

This corresponds to taking a clipped score V (x, y) = M 1{y > 0} + ŝ(x) with a sufficiently large constant M > 0 in Algorithm 1. Note that ĝ(x, 1) + ĝ(x, 0) = 1 and therefore it also satisfies the condition for the inclusion function specified in Theorem 3.1. The intuition about the expression of the inclusion function Gi is given in the next subsection, where we discuss the sharpness of our results. Consider the conformal policy learning algorithm with estimated welfare cutoff (if necessary) π̂welfare (Xn+1 ) := 1{pwelfare ≤ α} 1{µ̂1 (Xn+1 ) > µ̂0 (Xn+1 )}. n+1 Theorem 3.4 shows that when conditional expectation functions of potential outcomes are correctly specified, CPL achieves the optimal welfare asymptotically. Its proof is included in Appendix B.2. Theorem 3.4. Suppose Assumption 2.1 holds with e(x) = 0.5 and Assumption 2.2 holds. Suppose ∥µ̂t (X) − µt (X)∥L2 (PX ) = oP (1) as n → ∞ for t ∈ {0, 1}. Additionally, assume swelfare (X) has no point mass. Then, ∗ ∗ E[Y (π̂welfare (Xn+1 ))] → Welfare(πwelfare ; P) as n → ∞ where πwelfare is the optimal solution in Theorem 3.3. The consistency condition on the estimated conditional expectation functions could be achieved by flexible machine-learning-based methods. Moreover, even if the consistency condition fails, the conformal policy learning algorithm based on pwelfare still satisfies the finite-sample safety guarantee as per Algorithm 1. n+1 In Appendix A.1, we present analogous optimality results when researchers are interested in maximizing the “power”, the probability of treating the test unit. The only difference is that we should now use a power-oriented score ŝ(x) = γ̂(x).

3.4

Optimality, Sharpness, and Selective Calibration

The optimality result suggests that, in the non-trivial setting where the constraint is binding, conformal policy learning can exhaust the harm rate budget under partial identification, matching the globally optimal rule whose worst-case harm rate is exactly α. We now continue on Remark 3.2 to explain why this sharpness can be achieved. The use of clipped score essentially eliminated the conservativeness of using 0 instead of Yi∗ . We now focus on the conservativeness arising from replacing the “oracle” outcome Yi∗ = max{Yi (1), 1−Yi (0)} by the proxy label Yi† = Yi Ti +(1−Yi )(1−Ti ) ≤ Yi∗ , and discuss how this is eliminated by selective calibration. The use of Yi† is deeply rooted in the partial identification nature of the problem: even for the labeled data, the harm Yi∗ is not observed. As such, the sharpness hinges on whether, on the population level, calibration with Yi† matches the worst-case harm-rate bound that relies on γ(Xi ) in (3.8). CPL achieves so through the selective calibration mechanism. Assuming, for intuitions, that the conditional mean functions are perfect, the optimal conformal p-value uses the inclusion indicator Gi = Ti 1{1 − µ1 (Xi ) ≤ µ0 (Xi )} + (1 − Ti ) 1{1 − µ1 (Xi ) > µ0 (Xi )}. Namely, a treated unit Xi enters the calibration set if and only if 1 − µ1 (Xi ) ≤ µ0 (Xi ), in which case γ(Xi ) = 1 − µ1 (Xi ) = E[1{Yi† = 0} | Xi , Ti = 1] since Yi† = Yi (1) for Ti = 1. That is, the proxy label Yi† provides “sharp” calibration for the worst-case harm rate γ(Xi ). A similar observation applies to control units. The above discussion shows that for sharp harm rate control, one should use a labeled unit for calibration if and only if the treatment status matches whether the outcome model is “active” in the worst-case harm rate function γ(x), i.e., whether a treated unit obeys γ(Xi ) = 1 − µ1 (Xi ) or a control unit obeys γ(Xi ) = µ0 (Xi ). In the simulations, we shall evaluate the “informativeness” of labeled data and the power of CPL. Finally, we re-emphasize the model-free nature of conformal policy learning. Even though the arm informativeness—through the inclusion function ĝ—is estimated and therefore certainly imperfect, our method provides conservative harm rate control without any modeling assumptions.

11

4

Conformal Policy Learning with Stratified Experiments

In this section, we use weighted conformal inference to generalize our previous results to allow for both arbitrary, known propensity score e(x) and arbitrary inclusion function ĝ. It also serves as the foundation for CPL in observational studies, which we cover in the next section. The selective calibration process remains the same as Section 3.2. Consider any pre-trained monotone score function V : X × {0, 1} → R and inclusion function ĝ : X × {0, 1} → [0, 1]. The calibration set is the same as (3.5) where Gi ∼ Bern(ĝ(Xi , Ti )) are independently drawn inclusion indicators. When computing conformal p-values for a new test point Xn+1 ∼ PX , conditional on Gi = 1, there is a covariate shift between Dcalib and Xn+1 that arises from both the treatment assignment and the exogenous random inclusion into the calibration set. To address this shift, we define the function w : X → R+ by −1 w(x) = e(x)ĝ(x, 1) + (1 − e(x))ĝ(x, 0) . (4.1) Because e(x) := Pr(Ti = 1 | Xi = x) and ĝ(x, t) := Pr(Gi = 1 | Xi = x, Ti = t), one can show that w(x) = Pr(Gi = 1 | Xi = x)−1 , the inverse probability of being included for calibration given Xi = x, regardless of the choice of e(x) and ĝ. Given the weights, we compute Pn w(Xn+1 ) + i=1 Gi w(Xi ) 1{V (Xi , Yi† ) ≤ V (Xn+1 , 0)} Pn . (4.2) pstr = n+1 w(Xn+1 ) + i=1 Gi w(Xi ) The p-value (3.6) in Section 3 is a special case of (4.2), where the simplifying assumptions e(x) = 0.5 and ĝ(x, 1) + ĝ(x, 0) = a for a constant a > 0 implied constant weights and no reweighting is needed. Theorem 4.1 states that CPL with the above weighted conformal p-value achieves the distribution-free safety guarantee, whose proof is in Appendix B.3. Theorem 4.1. Suppose Assumption 2.1 and Assumption 2.2 hold, and e(x) is known. Then, for any score function V : X × Y → R that is non-decreasing in the second argument and for any inclusion function ĝ : X × {0, 1} → [0, 1] obeying e(X)ĝ(X, 1) + (1 − e(X))ĝ(X, 0) > 0, PX -a.s., it holds that P(Yn+1 (1) = str 0, Yn+1 (0) = 1, pstr (Xn+1 ) = 1{pstr n+1 ≤ α) ≤ α for any α ∈ (0, 1). That is, (2.1) holds for π̂ n+1 ≤ α}. As in Section 3, given consistent outcome modeling, CPL achieves optimality with Pn w(Xn+1 ) + i=1 w(Xi )Gi 1{Yi† = 0} 1{ŝ(Xi ) ≤ ŝ(Xn+1 )} str-opt Pn pn+1 = , w(Xn+1 ) + i=1 w(Xi )Gi

(4.3)

where Gi = Ti 1{1 − µ̂1 (Xi ) ≤ µ̂0 (Xi )} + (1 − Ti )1{1 − µ̂1 (Xi ) > µ̂0 (Xi )}. The score functions remains the same as the preceding case: ŝ(x) = −(µ̂1 (x) − µ̂0 (x))/γ̂(x) for welfare maximization. This is the weighted version of equation (3.11). The proof of Theorem 4.2 is in Appendix B.4. Theorem 4.2. Suppose Assumption 2.1 and Assumption 2.2 hold, and e(x) is known. Suppose ∥µ̂t (X) − P µt (X)∥L2 (PX ) → 0 as n → ∞ for t ∈ {0, 1} and swelfare (X) has no point mass. Let π̂str-opt (Xn+1 ) := 1{pstr-opt ≤ α} 1{µ̂1 (Xn+1 ) > µ̂0 (Xn+1 )} n+1 ∗ where we take the same Gi as (3.12) and ŝ(·) as (3.13). Then E[Y (π̂str-opt (Xn+1 ))] → Welfare(πwelfare ; P) as ∗ n → ∞, where πwelfare is the optimal solution in Theorem 3.3.

Analogously, CPL with the same selection indicators {Gi } as (3.12) and the estimated power-optimal score yields asymptotically optimal power; this result is deferred to Theorem A.2 in Appendix A.1.

5

Conformal Policy Learning with Observational Data

In this section, we further generalize CPL to observational studies. With observational data, the main challenge is that the propensity score, hence the weight function (4.1), is unknown and needs to be estimated. 12

While a natural idea is to estimate the weights and plug them into the conformal p-value, it remains unclear how the safety guarantee changes with the estimation quality. We present a learn-then-balance procedure to obtain the estimated weights so the resulting policy learning algorithm enjoys a double robustness property.

5.1

Conformal Policy Learning with Learn-then-Balance Weights

We construct p-values of a similar form as (4.2) and (4.3): Pn ŵn+1 + i=1 ŵi Gi 1{Yi† = 0} 1{ŝ(Xi ) ≤ ŝ(Xn+1 )} obs Pn pn+1 = , ŵn+1 + i=1 ŵi Gi

(5.1)

where the inclusion indicators are Gi ∼ Bern(ĝ(Xi , Ti )) for a function ĝ : X × {0, 1} → [0, 1] whose training process is independent of the labeled and unlabeled data. The only difference of (5.1) from (4.3) is that the weights are now estimated. Notably, here we work with a form analogous to (4.3) since our balancing weights are more easily stated in terms of the score function ŝ(·). It corresponds to (4.2) with the conformity score V (x, y) = M 1{y > 0} + ŝ(x) for a sufficiently large constant M > 2 supx |ŝ(x)|. Denoting Icalib = {i ∈ [n] : Gi = 1}, the estimated weights {ŵi }i∈Icalib ∪{n+1} in (5.1) are obtained by a learn-then-balance approach. We begin with two learned functions, one preliminary weight function w̃ : X → R+ , and one estimated harm-rate function γ̂ : X → [0, 1]. A natural choice is to define w̃(x) = 1/[ê(x)ĝ(x, 1) + (1 − ê(x)ĝ(x, 0)] where ê(x) is an estimated propensity score function, and γ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)} for estimated outcome models {µ̂t (·)}t∈{0,1} . For simplicity, we require the propensity score and outcome models to be trained independently of the calibration and test data (in practice, one can use sample splitting to fit the models (Chernozhukov et al., 2018) and the properties are analogously studied in Jin and Zubizarreta (2025)). Then, define the balancing feature vector   n  1X γ̂(Xi ) 1{ŝ(Xi ) ≤ t} ≤ α . ϕ̂(x) = γ̂(x) 1{ŝ(x) ≤ t̂}, w̃(x) ∈ R2 , where t̂ = sup t ∈ R : n i=1 Finally, we let {ŵi }i∈Icalib be the optimal solution to the following optimization program with δn = O(n−1/2 ): ( ) n X X X 1 1 1 argmin ∥w∥2 : ϕ̂k (Xi ) ≤ δn , k ∈ {1, 2}; wi ϕ̂k (Xi ) − wi = 1 . (5.2) |Icalib | n i=1 |Icalib | w⪰0 i∈Icalib

i∈Icalib

We shall see that the post-hoc processing to ensure the finite-sample approximate balancing (5.2) is essential for us to obtain the doubly robust safety guarantees even when the propensity model is misspecified. The first condition in (5.2) follows the balancing weights (Hainmueller, 2012; Zubizarreta, 2015), by enforcing a finite-sample balancing condition inspired by the desired population-level property: for the correct weights P Pn 1 1 w(·), it should hold that |Icalib i∈Icalib w(Xi )f (Xi ) ≈ n i=1 f (Xi ) for any fixed function f : X → R. Here, | we balance specifically-designed functions to obtain favorable statistical properties under general conditions on the learned functions.

5.2

Doubly Robust Safety Guarantees

We now formalize the safety guarantees of CPL (5.1) with weights {ŵi }i∈Icalib from (5.2). These results will require mild regularity conditions for the learn-then-balance optimization, which we defer to Assumption A.10. These conditions require overlap and boundedness, local stability of the population balancing solution, and regularity of the population cutoff. Under these conditions, and given that either the preliminary weight function or the outcome models are consistent, we obtain asymptotic safety guarantee for CPL. The proof of Theorem 5.1 is included in Appendix B.5, which follows from our general theory with estimated weights in Appendix C.1. Theorem 5.1 (Model double robustness). Suppose Assumption A.10 holds, and δn = O(n−1/2 ). Then, lim sup P(Yn+1 (π̂obs (Xn+1 )) < Yn+1 (0)) ≤ α n→∞

13

under either of the following conditions: (i). The preliminary weight function is consistent: ∥w̃ − w∥L2 (PX ) = oP (1) for the true weight w(·) in (4.1). (ii). The outcome models are consistent: ∥γ̂ − γ∥L2 (PX ) = oP (1). Remark 5.2. We remark that the techniques and conditions here differ from existing model double-robustness results in conformal prediction (e.g. Lei and Candès (2021)) where correct outcome model alone is enough to ensure validity; in our setting, because of the thresholding nature of the CPL policy, explicit finite-sample balance turns out to be an important element of our theoretical analysis for double robustness. Taking a step further, we show that the learn-then-balance weights lead to a product error rate structure, which further yields the OP (n−1/2 ) convergence of the harm rate of CPL given slow, nonparametric convergence rates of the estimated models. The proof of Theorem 5.3 is in Appendix B.6. Theorem 5.3 (Rate double robustness). Suppose Assumption A.10 holds and δn = O(n−1/2 ). If ∥w̃ − w∥L∞ (PX ) = OP (n−1/4 ) and ∥γ̂−γ∥L1 (PX ) = OP (n−1/4 ), then {P(Yn+1 (π̂obs (Xn+1 )) < Yn+1 (0) | An )−α}+ = OP (n−1/2 ), where An is the σ-field for the randomness in calibration data, training, and selective inclusion.

5.3

Optimal Conformal Policy Learning with Observational Studies

Since the population optimization problem only depends on the super-population of the potential outcomes and observed features, the population-level optimal solution remains the same as Theorem 3.3. The following theorem shows that under mild conditions, our method with the optimally chosen score and inclusion function achieves the optimal welfare with observational data. Its proof is in Appendix B.7. The results for power optimality are deferred to Appendix A.1. Theorem 5.4 (Welfare optimality with observational data). Let π̂obs (Xn+1 ) = 1{pobs n+1 ≤ α} 1{ŝ(Xn+1 ) < 0} where we take the same Gi as (3.12) and ŝ(·) as (3.13). Suppose Assumption A.10 holds, and δn = O(n−1/2 ). Furthermore, assume ∥ĝ − g ∗ ∥L2 (PX,T ) + ∥γ̂ − γ∥L2 (PX ) + ∥ŝ − swelfare ∥L2 (PX ) = oP (1) for g ∗ (x, t) = t 1{1 − µ1 (x) ≤ µ0 (x)} + (1 − t) 1{1 − µ1 (x) > µ0 (x)}, and swelfare (X) defined in (3.10) has no point ∗ ∗ is the optimal solution in ; P) as n → ∞, where πwelfare mass. Then E[Y {π̂obs (Xn+1 )}] → Welfare(πwelfare Theorem 3.3. As in randomized experiments, optimality requires two convergence conditions: ĝ → g ∗ , which makes the proxy-label calibration sharp, and ŝ → swelfare , which provides the optimal treatment ranking. When ĝ, ŝ, and γ̂ are constructed from plug-in outcome model estimates, these conditions follow from consistent outcome models under the usual regularity conditions. In this outcome-correct case, the balancing condition aligns the weighted calibration with the oracle harm-welfare curve, so the weight estimator need not be consistent. However, inheriting the model consistency of Theorem 5.1, one can show that the conclusion also continues to hold under consistent weights, w̄ ∝ w, without requiring γ̂ → γ, provided that ĝ → g ∗ and ŝ → sopt still hold. Yet, we do not pursue it here since the latter two convergence conditions usually require consistency of the estimated outcome models which imply the consistency of the harm rate estimation.

6

Simulation Studies

6.1

Balanced Randomized Experiments

We begin by evaluating the methods in balanced randomized experiments (Section 3). The data in the randomized experiment settings are generated using the general framework below. Data generating processes. Throughout, we generate the features X ∼ PX for some distribution PX and i.i.d. treatment indicators T ∼ Bern(1/2), independent of everything else. The potential outcomes follow P(Y (1) = 1 | X = x) = µ1 (x) and P(Y (0) = 1 | X = x) = µ0 (x) for some functions µ1 (·) and µ0 (·). The potential outcomes are coupled negatively, meaning that P(Y (1) = 0, Y (0) = 1 | X = x) = 14

min{1 − µ1 (x), µ0 (x)}. Note that any coupling would lead to the same observed data distribution PX,Y,T (and thus the same results for any method), but the negative coupling leads to the worst-case harm rate. The specific choice of PX and µ1 (·), µ0 (·) is introduced in each set of simulations below. We first design four diverse settings to compare our methods and baselines, followed by two additional settings to investigate the performance of our methods, including (i) the power of selective calibration by varying the informativeness of arms, and (ii) decisions by heterogeneous subgroups. 6.1.1

Harm rate control in diverse settings

The first set of simulations evaluates the harm rate control across diverse settings and modeling choices. We set X ∈ R20 with PX = N (0, Id ). The conditional mean functions are µt (x) = exp(ηt (x))/(1 + exp(ηt (x)), t ∈ {0, 1} for some functions ηt (x). We design four settings (see Appendix E.1 for detailed DGPs): • Setting 1: approximately linear, where ηt (x) = βt⊤ x for some βt ∈ R20 and fully-observed X ∈ R20 ; • Setting 2: approximately linear DGP but the observed covariates are X1:7 ∈ R7 ; • Setting 3: nonlinear ηt (x) involves the first 5 entries of x, with fully observed X ∈ R20 ; • Setting 4: nonlinear ηt (x) involves the first 8 entries of x, but the observed covariates are X1:5 ∈ R5 . In settings 2 and 4, missing some covariates typically reduces the accuracy of learned outcome models (finite-sample harm-rate control still holds as per Theorem 3.1). We vary the total labeled sample size ntotal ∈ {500, 1000, 2000}, fixing the number of test data m = 1000, and α = 0.1. We use Dlabel to denote the labeled data and Dtest to denote the test data. Methods and evaluation.

We compare the following three baselines and two CPL methods:

(1) Li et al., Li et al. (2023) with doubly-robust estimator and constrained optimization among depth-2 decision trees via econml python library. The policy is learned on Dlabel and evaluated on Dtest . This method requires the distributional assumption about the potential outcomes. In particular, it assumes that the potential outcomes are non-negatively correlated, which is violated in this setup. (2) Policy-Tree, which maximizes welfare E[Y (π(Xn+1 ))] among policies represented by depth-2 decision trees with econml. The policy is learned on Dlabel and evaluated on Dtest . While this method is not designed for safety control, we include it as the baseline due to its popularity as a standard practice. (3) Threshold, which uses Dlabel to obtain an estimator γ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)} and treat {j ∈ Pm [m] : γ̂(Xn+j ) ≤ γ̃} where γ̃ = max{γ : j=1 γ̂(Xn+j ) 1{γ̂(Xn+j ) ≤ γ} ≤ 0.1 · m}. This heuristic calibration is valid only when the estimator γ̂ is sufficiently accurate. (4) CPL-sel-welfare, the CPL algorithm that maximizes the welfare, i.e., Algorithm 1 with g(x, t) = t 1{1 − µ̂1 (x) ≤ µ̂0 (x)} + (1 − t) 1{1 − µ̂1 (x) > µ̂0 (x)} and a clipped score function V (x, y) = M 1{y > 0} − (µ̂1 (x) − µ̂0 (x))/γ̂(x) with a sufficiently large constant M > 0. This method always satisfies the safety guarantee and asymptotically achieves the optimal welfare if µt (x) is consistently estimated. (5) CPL-sel-power, the CPL algorithm that maximizes the power, i.e., Algorithm 1 with g(x, t) = t 1{1 − µ̂1 (x) ≤ µ̂0 (x)}+(1−t) 1{1− µ̂1 (x) > µ̂0 (x)} and a clipped score function V (x, y) = M 1{y > 0}+ γ̂(x) with a sufficiently large constant M > 0. This method always satisfies the safety guarantee and asymptotically achieves the optimal power if µt (x) is consistently estimated (Appendix A.1). For the two CPL variants, the labeled data Dlabel is randomly split into the training (75%) and calibration (25%) folds Dtrain and Dcalib . The training fold is used to fit two models µ̂1 (x) and µ̂0 (x) for µ1 (x) and µ0 (x), respectively. The function γ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)} then estimates the upper bound on the harm rate function. We adopt two model classes for training the outcome models: logistic regression, and random forest, which are used in both Threshold and CPL methods. Li et al. and Policy-Tree use random forests for nuisance component estimation to be consistent with the tree-based policy class. See Appendix E.1 for additional details on method implementation. Given the learned treatment Tn+j := π̂(Xn+j ) ∈ {0, 1} for j ∈ [m] produced by the methods, we evaluate Pm Pm 1 1 three metrics: the harm rate by m j=1 Tn+j 1{Yn+j (1) < Yn+j (0)}, the welfare by m j=1 Yn+j (Tn+j ),

15

method Setting 1 (Linear)

CPL−sel−power

Setting 2 (Linear, missing covariates)

CPL−sel−welfare

Setting 3 (Nonlinear)

Setting 4 (Nonlin. missing covariates)

0.1 0.0

RF

0.2 0.1 0.0

(b)

1000

2000

Setting 1 (Linear)

500

1000

2000

500

1000

Setting 2 (Linear, missing covariates)

2000

Setting 3 (Nonlinear)

500

1000

2000

Setting 4 (Nonlin. missing covariates)

1.00 0.75 0.50 0.25 0.00 1.00 0.75 0.50 0.25 0.00

Logistic RF

Fraction of treatment

Policy−tree

0.2

500

500

(c)

1000

2000

Setting 1 (Linear)

500

1000

2000

500

Setting 2 (Linear, missing covariates)

1000

2000

Setting 3 (Nonlinear)

500

1000

2000

Setting 4 (Nonlin. missing covariates)

0.8 0.6 0.4 0.2 0.0 0.8 0.6 0.4 0.2 0.0

Logistic RF

Average welfare

Threshold

Logistic

Realized harm rate

(a)

Li et al.

500

1000

2000

500

1000

2000

500

1000

2000

500

1000

2000

Labeled sample size

Figure 2: (a) Empirical harm rate, (b) power (probability of treatment in the test units), (c) average welfare of various methods at level α = 0.1 in the randomized experiment studies, averaged over 200 independent simulation runs. Each row represents a different prediction model (logistic regression, random forest) for (µ̂1 , µ̂0 ), and each column represents a different data-generating process. The x-axis is the total labeled sample size n. The blue dashed lines in (b) and (c) show the optimal power and welfare of CPL based on oracle correct models (µ1 , µ0 ) without any estimation error. 1 and the fraction of treatment by m

Pm

j=1 Tn+j . All metrics are averaged over 200 independent runs.

Simulation results. Figure 2 presents the results for the five methods across various settings. First, the three baselines lead to a drastic violation of the target harm rate. The Threshold method relies on accurate outcome models to ensure consistency of γ̂(·), which is difficult to satisfy with limited labeled data. Policy-tree, which focuses on welfare maximization can incur large harm rates. Finally, Li et al. achieves harm rate only when the two potential outcomes are non-negatively correlated given the features; as a result, it leads to exceedingly high harm rate due to (i) worst-case negative coupling and (ii) inconsistency due to missing covariates in settings 2 and 4. On the other hand, the two variants of CPL control the harm rate tightly at the target level, showing both the validity and sharpness of CPL. The two variants based on different score functions demonstrate negligible difference: in these settings, the rank of instances based on the two optimal scores does not drastically change the decisions by CPL. In practice, this means that researchers can approximately maximize both power and welfare together via CPL without worrying about the tradeoff between the two objectives. Finally, compared with the oracle power and welfare (blue dashed lines), the CPL methods achieve close-to-optimal performance, and the gap shows the impact of estimation error in µ̂t functions. Such gap seems to be moderate even for misspecified models (Logistic regression in settings 3 and 4).

16

method

CPL−treat

RF

(b)

0.10 0.08 0.06 0.04 0.00

0.25

0.50

0.75

1.000.00

0.25

0.50

0.75

Informativeness of the treated arm

1.00

CPL−control

CPL−sel−power

Logistic

CPL−sel−welfare

(c)

RF

Logistic

RF

0.5

Average welfare

Logistic

Fraction of treatment

Realized harm rate

(a)

0.4 0.3 0.2 0.00

0.25

0.50

0.75

1.000.00

0.25

0.50

0.75

Informativeness of the treated arm

1.00

0.6 0.5 0.4 0.3 0.00

0.25

0.50

0.75

1.000.00

0.25

0.50

0.75

1.00

Informativeness of the treated arm

Figure 3: (a) Empirical harm rate, (b) power (probability of treatment) and (c) average welfare of the seven procedures in the arm-informativeness experiments at level α = 0.1 over 200 independent runs. Within each panel, each column represents a prediction model (logistic regression, random forest) for (µ̂1 , µ̂0 ). 6.1.2

Effectiveness of selective calibration

We now use another set of experiments to dive deeper into the selective calibration mechanism. Following the discussion in Section 3.4, the harm rate control by using the conservative proxy Yi† = Ti Yi +(1−Ti )(1−Yi ) is tight if a selected sample satisfies (i) Ti = 1 and γ(Xi ) = 1 − µ1 (Xi ), or (ii) Ti = 0 and γ(Xi ) = µ0 (Xi ). Our selective calibration method in Section 3.2 aims to address this by adaptively selecting the informative arm. In this part, we vary the magnitudes of µ1 (x) and µ0 (x) to how this strategy contributes to the statistical efficiency. We vary the proportion of samples obeying γ(X) = 1 − µ1 (X) (treated arm is informative) and those obeying γ(X) = µ0 (X) (control arm is informative); see Appendix E.2 for the detailed data-generating process. In addition to the two CPL variants evaluated in Section 6.1.1, we evaluate two more procedures: (6) CPL-treat, our method in Section 3.1 with treated samples in Dcalib and V (x, y) = M y − µ̂1 (x) with a sufficiently large constant M > 0. (7) CPL-control, our method in Section 3.1 with control samples in Dcalib and V (x, y) = M y + µ̂0 (x) with a sufficiently large constant M > 0. The procedures are evaluated in terms of harm rate, power (fraction of treatment), and welfare, with the metrics averaged over 200 independent runs. The results are summarized in Figure 3, where the x-axis is the fraction of treated samples being informative (i.e., γ(X) = 1−µ1 (X)). As before, all variants of CPL control the harm rate below α = 0.1. CPL-treat and CPL-control demonstrate clear power tradeoffs: CPL-treat is more powerful when the treated arm is more informative (x-axis above 0.5), and the opposite happens otherwise. Importantly, CPL-sel-power and CPL-sel-welfare are often comparable to the more powerful single-arm variants, showing the effectiveness of selective calibration. This justifies our recommendations for the asymptotically optimal variants (Section 3.3), since in practice it is often unknown which arm might be more powerful. 6.1.3

Performance under subgroup heterogeneity

To further inspect the behavior of CPL under treatment effect heterogeneity, we design a setting with three subgroups driven by the first two features in X ∈ R20 (Figure 4 panel (a)). There is strong cross-group heterogeneity, but the units in the same group are largely similar. Studying the decisions in each group offers a zoom-in observation of “who gets treated” with different scoring functions s(·) in CPL. The first group consists of half of the population, whose worst-case harm rate γ(X) is relatively large while the conditional average treatment effect τ (X) = µ1 (X) − µ0 (X) is also large. Groups 2 and 3 consist of 1/4 of the population each, whose worst-case harm rate γ(X) and conditional average treatment effect τ (X) are both small; the main difference is that the treated samples are more “informative” in Group 2 (i.e., 1 − µ1 (X) < µ0 (X)), while the control samples are more informative in Group 3. We follow the same procedures as before to evaluate the four variants of CPL. The results averaged over 200 independent simulation runs are summarized in Figure 4. In panel (b), we

17

(a) 0.50

method Group 3: small harm rate (control uninformative) small hetero.treat. effect Group 1: large harm rate large hetero.treat. effect

0.00

−0.25

Group 2: small harm rate (treated informative) small hetero.treat. effect

−0.50 −0.50

−0.25

0.00

0.25

0.50

Fraction of treatment

Feature 2

0.25

(b) CPL−treat

Group 1

CPL−control Group 2

CPL−sel−power Group 3

CPL−sel−welfare Overall

0.20 0.15 0.10 0.05 0.00

Feature 1

Figure 4: (a) Subgroup setup, (b) Per-group fraction of Tn+1 = 1 for variants of CPL in the subgroup DGP at level α = 0.01. Panels (c-d) show random forests (RF) as the outcome model only. show the fraction of Tn+j within each group based on random forests predictions. Different score functions lead to different treatment prioritization patterns. CPL-treat concentrates the safe treatment budgets on group 2 (small γ(X) with treatment arm being informative) since it ranks instances based on µ̂1 (x), while CPL-control concentrates the budgets on group 3. In contrast, CPL-sel-power distributes the budgets relatively uniformly across groups 2 and 3 (since their harm rates are comparably small) by using the score function γ̂(x). CPL-sel-welfare, which ranks instances by balancing harm rate and treatment effect size, puts most budgets on Group 1 (large treatment effects) but assigns fewer treatments in general.

6.2

Stratified Experiments and Observational Studies

In this part, we proceed to evaluate CPL in stratified experiments and observational studies. The stratified experiments induce a known covariate shift between the calibration and test data, while in observational studies this covariate shift needs to be estimated. Simulation settings. We first sample the triplets {(Xi , Yi (1), Yi (0)}n+m i=1 using data generating processes to be specified later. In the labeled data, the treatment assignments are sampled by Ti | Xi ∼ Bernoulli(e(Xi )) independently with propensity score function e(x) to be specified later. The observed outcome is Yi = Yi (Ti ). Methods and evaluation.

Fixing the confidence level at α = 0.1, we compare the following procedures:

(1) Li et al., Li et al. (2023) with a doubly-robust estimator and constrained optimization among depth2 decision trees via econml python library, where the outcome models and propensity scores are estimated under two-fold cross-fitting. The policy is learned on Dlabel and evaluated on Dtest . This method requires that the potential outcomes are non-negatively correlated, which is violated here. (2) Policy-Tree, which maximizes welfare E[Y (π(Xn+1 ))] among depth-2 decision trees with econml. The policy is learned with Dlabel and evaluated on Dtest with similar cross-fitting estimation of outcome and propensity score models. While this method is not designed for safety control, we include it as the baseline because of its popularity as a standard practice. (3) Threshold, the same as in Section 6.1 which calibrates the threshold using cumulative estimated γ̂(X) on the test data. This method is valid only when the estimator γ̂ is sufficiently accurate. (4) CPL-sel-welfare, the method in Section 5 using the same g(x, t) and s(x) functions as in the randomized experiment case. The weights are estimated using learn-then-balance with features ϕ̂(x) = (γ̂(x) 1{s(x) ≤ t̂}, ŵ(x)), where γ̂ and ŵ are estimated using logistic regression or random forests. (5) CPL-sel-power, the method in Section 5 using the same g(x, t) and s(x) functions as in the randomized experiment case. and the weights are estimated in the same way as CPL-sel-welfare. (6) Plugin-oracle-sel-welfare, the method in Section 4 using the correct weights and the same g(x, t) and s(x) functions as CPL-sel-welfare.

18

method

Li et al.

Setting 1 (Linear)

CPL−sel−welfare

Setting 2 (Linear, missing covariates)

Plugin−oracle−sel−power

Setting 3 (Nonlinear)

Plugin−oracle−sel−welfare

Setting 4 (Nonlin. missing covariates)

0.1 0.0

RF

0.2 0.1 0.0

(b)

1000

2000

Setting 1 (Linear)

500

1000

2000

500

1000

Setting 2 (Linear, missing covariates)

2000

Setting 3 (Nonlinear)

500

1000

2000

Setting 4 (Nonlin. missing covariates)

1.00 0.75 0.50 0.25 0.00 1.00 0.75 0.50 0.25 0.00

Logistic RF

Fraction of treatment

CPL−sel−power

0.2

500

500

(c)

1000

2000

Setting 1 (Linear)

500

1000

2000

500

Setting 2 (Linear, missing covariates)

1000

2000

Setting 3 (Nonlinear)

500

1000

2000

Setting 4 (Nonlin. missing covariates)

0.8 0.6 0.4 0.2 0.0 0.8 0.6 0.4 0.2 0.0

Logistic RF

Average welfare

Policy−tree

Logistic

Realized harm rate

(a)

Threshold

500

1000

2000

500

1000

2000

500

1000

2000

500

1000

2000

Labeled sample size

Figure 5: (a) Empirical harm rate, (b) power (probability of treatment in the test units), (c) average welfare at level α = 0.1 in observational studies, averaged over N = 200 runs. Each row is a prediction model (logistic regression, random forest) for (µ̂1 , µ̂0 ), and each column is a data-generating process. The x-axis is the total labeled sample size n. The Plugin methods use the true weights. (7) Plugin-oracle-sel-power, the method in Section 4 using the correct weights and the same g(x, t) and s(x) functions as CPL-sel-power. Here, the last two plugin-oracle methods essentially evaluate CPL in the stratified experiment setting where the propensity score e(x), and hence the covariate shift weights, are known. Our theory implies their finite-sample harm rate control due to weighted exchangeability, no matter the accuracy of the scores and selective calibration functions. On the other hand, CPL-sel-power and CPL-sel-welfare evaluate the robustness of CPL in observational studies with estimated weights. 6.2.1

Harm rate control with known or estimated weights

We first evaluate the harm rate control using the same data-generating process as in Section 6.1, together with a linear or nonlinear propensity score. The empirical harm rate, fraction of treatment, and average welfare, averaged over 200 independent simulation runs, are summarized in Figure 5. First, the three baselines drastically violate the harm rate control due to similar reasons as in Section 3, now with additional estimation challenges in observational settings. The two plug-in oracles, Plugin-oracle-sel-power and Plugin-oracle-sel-welfare, confirm the finite-sample harm-rate control of Theorem 4.1. Finally, the two methods with estimated weights, CPL-sel-power and CPL-sel-welfare, also demonstrate robust harm-rate control below the target level. Although the four CPL methods rely on distinct scoring functions, they achieve similar performance in terms of fraction of treatment and average welfare, showing that the two objectives can align.

19

method

CPL−sel−power

CPL−sel−welfare

Plugin−oracle−sel−power

Plugin−oracle−sel−welfare

OR: Correct Logistic

OR: Correct Logistic

OR: Misspec. Logistic

OR: Misspec. Logistic

OR: Random Forest

PS: Correct Logistic

PS: Misspec. Logistic

PS: Correct Logistic

PS: Misspec. Logistic

PS: Random Forest

0.06 0.05 0.04 0.03

n = 1000

Realized FNA

n = 500

0.06 0.05 0.04 0.03

n = 2000

0.06 0.05 0.04 0.03 1

2

3

4

1

2

3

4

1

2

3

4

1

2

3

4

1

2

3

4

Severity of confounding

Figure 6: Empirical harm rate in the double robustness experiments; results are averaged over N = 200 independent runs. Each column corresponds to one specification of outcome and propensity score models. Each row corresponds to a total labeled sample size. The x-axis is the strength of confounding. 6.2.2

Double robustness

We now additionally examine the double robustness property established in Section 5. We design a datagenerating process where the outcome models and propensity score model are all logistic in some nonlinear transformation of the raw features X ∈ R8 ; see Appendix E.3 for details. We sample observational data {(Xi , Ti , Yi )}ni=1 as the labeled dataset, where a random subset of 75% is used as the training fold for training the models µ1 (·), µ0 (·), and e(·), while the remaining is used as the calibration data in CPL. We vary a parameter in the propensity score model to control the strength of confounding (the x-axis of Figure 6). The methods evaluated include CPL-sel-power/welfare based on estimated propensity scores, and two known-propensity baselines Plugin-oracle-sel-power/welfare used as oracle comparison only. The four procedures are implemented in the same way as in the preceding parts. Within each procedure, we vary the model classes of the outcome and propensity score models to demonstrate the double robustness property. A correct logistic model runs logistic regression over the nonlinearly-transformed features, while a misspecified logistic model runs logistic regression over the raw features. By Theorem 5.1, we expect CPL-sel-power/welfare to control the harm rate when either of them is correctly specified. Finally, we also consider that both outcome models and propensity scores are trained via random forests, which typically have slower-than-parametric convergence rates yet are less prone to model misspecification; we expect it to control the harm rate, especially when the labeled sample size is sufficiently large. The empirical harm rate averaged over 200 independent simulation runs at nominal level α = 0.05 is summarized in Figure 6. In the first three columns, when either the outcome models or the propensity score model is well-specified, we observe tight harm-rate control by the two CPL variants, which is also close to the plugin-oracle ones. This validates the double robustness theory. When the outcome models are misspecified, we observe a larger slack in the harm rate than the other two configurations when sample size is small. On the other hand, when both models are misspecified (the fourth column), the two CPL variants can violate the harm rate control, yet the violation is moderate. We note that we design the settings deliberately to fail the both-misspecified procedures. In many other settings, misspecification can create a slack in the conservative calibration through the proxy outcome Yi† , which often compensates the misspecification in the weights and keeps the harm rate below the budget even though both models are wrong. Finally, the nonparametric models in the fifth column yield tight harm rate control, showing the robustness to model misspecification and the quick convergence of the harm rate in CPL.

20

7

Empirical Application to AI-Powered Interventions

In this section, we apply CPL to a real-world dataset in the social sciences, where the AI model, ChatGPT, is used as an intervention to persuade participants out of some conspiracy beliefs, in which events are understood as being caused by secret, malevolent plots involving powerful conspirators (Costello et al., 2024). The original study found that brief conversations with AI could reduce conspiracy beliefs by 20 percent on average, and the effect was durable for at least 2 months. Given the societal concern over widespread conspiracy theories, an increasing number of researchers and policymakers are evaluating similar AI interventions as a scalable solution. If policymakers scale up such AI interventions, it is of significant importance to consider safety, as the treatment effects of AI interventions are likely to be highly heterogeneous for several reasons. First, “people believe a wide range of conspiracies, and the specific evidence brought to bear in support of even a particular conspiracy theory may differ substantially from believer to believer” (page 1, Costello et al., 2024). Second, as AI chatbots treat people with natural texts as the intervention, the content of the treatment itself is heterogeneous and unpredictable. Here, to safely scale up these types of AI-powered interventions, we use CPL to decide who should be treated by AI by controlling the harm rate with the safety guarantee. In this study, before the experiment, the participants stated a conspiracy theory they believed in and reported a numerical score quantifying their belief in it. They are then randomly assigned to treated and control groups, where treated participants engage in a live conversation with a GPT model that is instructed to talk them out of the conspiracy, while the participants in the control condition engage with a neutral conversation with the GPT model. After the experiment, they again report a numerical score quantifying their belief in the same conspiracy theory. We take all participants in the raw dataset as the analysis population. We binarize the outcome Y to indicate whether the post-experiment belief score is below 50, a cutoff the authors originally used to define their analysis population. The participants are randomly split into training (40%), calibration (40%), and test (20%) folds. There are 416 participants originally treated out of 667 participants in the test fold. The fraction of Y = 1 in the treated group is 0.274, while the fraction of Y = 1 in the control group is 0.100. We consider a stringent harm rate of α = 0.025. We build features based on participants’ demographic covariates (education, age, gender, race, religion), political and psychological covariates, AI-related covariates (familiarity and trust in generative AI), baseline belief state variables and textual embedding for the stated conspiracy. The training fold is used to fit the outcome models and conditional treatment effect models, which are used in a similar way as in the simulations to build welfare-maximizing score functions (except that we truncate on extremely small estimated harm rate for stability). See Appendix E.4 for the detailed implementation. Empirical welfare and power. We report power (the fraction of treated units in the test data) and (estimated) empirical welfare of the welfare-maximizing and power-maximizing variants of CPL. We also compare the results against three baselines as references: the first is to treat everyone in the test data (All treat), the second is to treat no one in the test data (All control), and the third is to treat 2.5% of test units, which trivially satisfies the safety constraint. Because we observe the realized outcome in the test data, we can estimate the average welfare and harm real rate of the policy as follows. Let Tn+j ∈ {0, 1} be the actual treatment assigned by the experiment, and Tn+j = π̂(Xn+j ) be the decision produced by CPL. We are interested in the average welfare Welfare(π̂) := E[Yn+j (π̂(Xn+j ))] = E[Yn+j (1) · π̂(Xn+j ) + Yn+j (0) · (1 − π̂(Xn+j ))], for which an unbiased estimate is     real real ˆ Welfare = Êtest Yn+j · π̂(Xn+j ) | Tn+j = 1 + Êtest Yn+j · (1 − π̂(Xn+j )) | Tn+j =0 , (7.1) where Êtest denotes the empirical average in the test fold. We also estimate the harm rate P(Yn+j (1) < Yn+j (0), π̂(Xn+j ) = 1) by the following conservative estimator.  ˆ =Êtest [(1 − Yn+j ) 1{1 − µ̂1 (Xn+j ) ≤ µ̂0 (Xn+j )}π̂(Xn+j ) | T real = 1 Harm n+j  real + Êtest [Yn+j 1{1 − µ̂1 (Xn+j ) > µ̂0 (Xn+j )}π̂(Xn+j ) | Tn+j =0 . (7.2) 21

Est. harm rate Num. treatment Est. welfare

All treat 0.0392 (0.0111) 667 0.274 (0.0218)

All control 0 0 0.0996 (0.0189)

Trivially-safe 0.001 (0.0002) 16 0.104 (0.0184)

CPL-welfare 0.0247 (0.0094) 605 0.254 (0.0246)

CPL-power 0.0207 (0.0086) 615 0.250 (0.0231)

Table 1: Estimated harm rate, number of treatment, and estimated welfare (standard deviation) on the test fold using different methods. “All treat” treats all test units; “All control” treats no test units; “Triviallysafe” randomly treats α-fraction of test units. The welfare/harm rate estimates are based on (7.1) and (7.2).  ˆ is unbiased and asymptotically normal for the population quantity E {(1 − µ1 (Xn+j )) 1{1 − Note that Harm  µ̂1 (Xn+j ) ≤ µ̂0 (Xn+j )} + µ0 (Xn+j ) 1{1 − µ̂1 (Xn+j ) > µ̂0 (Xn+j )}} · π̂(Xn+j ) , which upper bounds the harm rate. This is a valid conservative estimator of the harm rate even if µ̂t (x) is misspecified, and this is a consistent estimator for the sharp upper bound of the harm rate when µ̂t (x) is consistently estimated. Results. The main results are summarized in Table 1. Several points are worth noting. First, treating everyone (All treat) violates the safety constraint as its harm rate is 3.92% and exceeds 2.5%. In contrast, our proposed CPL methods achieve the tight control of harm rates at 2.5% and 2.1%, respectively. Second, while treating no one (All control) and treating only 2.5% of test units (Trivially-safe), of course, satisfy the safety constraint, they have extremely low power and welfare. In contrast, our proposed methods can treat more than 90% of test units and achieve the high average welfare. While maintaining the safety constraint, our method simultaneously achieves high power and welfare. Finally, it is important to note that the difference between our welfare-maximizing and power-maximizing variants are minimal in practice. It is true that, consistent with our theory, our power-maximizing variant has a slightly higher fraction of treated test units, and our welfare-maximizing variant has a slightly higher average welfare. However, overall, their actual policy decision on who gets treated is similar, which implies that researchers can use either variant in practice and expect to approximately optimize both power and welfare in various applications. Safe treatments by CPL. We now take a closer look at the decisions produced by the welfare-maximizing varinat of CPL. Figure 7 visualizes the treatment decisions, where panel (a) plots the test points based on the predicted harm rate γ̂(X) and predicted treatment effect τ̂ (X), while panel (b) plots the test points based on the predicted outcomes µ̂0 (X) and µ̂1 (X). CPL treats the test units with the largest values of τ̂ (X)/γ̂(X). The decision boundary is plotted in both panels, and the region not treated is indicated in light grey. While the actual harm 1{Yn+j (1) < Yn+j (0)} is not observed, we conservatively estimate it by † real real real checking the label Yn+j = Tn+j Yn+j + (1 − Tn+j )(1 − Yn+j ), where Tn+j is the actual treatment received by real real the j-th test unit, among those with ĝ(Xn+j , Tn+j ) = 1, i.e., either µ̂1 (Xn+j ) ≤ µ̂0 (Xn+j ) and Tn+j = 1 or real real µ̂1 (Xn+j ) > µ̂0 (Xn+j ) and Tn+j = 0. The colored dots are those with ĝ(Xn+j , Tn+j ) = 1, among which the † † blue dots are those with Yn+j = 1 and the red dots are those with Yn+j = 0. We observe that very few red dots in the treated region can possibly be harmed. We further examine how the treatment decision by CPL varies with participants. Figure 8 plots the moving average of predicted treated and control outcomes, as well as the fraction of safe treatments, for every value of pre-experiment belief scores (smoothed by a moving window of size 10). In general, the two outcome curves indicate that a stronger pre-treatment belief makes it less likely to be persuaded out of the conspiracy in both conditions, but the effect of the AI intervention seems strong for participants with a strong pre-treatment belief. The welfare-maximizing variant of CPL prioritizes a treatment-effect-versus-harm-rate tradeoff. It mainly treats units with firm pre-treatment belief for whom the treatment is likely to make a huge difference (the gap between the two blue curves) while the estimated harm rate is relatively low. Figure 9 similarly reports this information among participants with a specific GenAI familiarity score (panel a) and GenAI trust score (panel b). In this case, however, the outcomes and treatment decisions do not change significantly based on these features, indicating that AI intervention may be similarly safe for users with different familiarity with or trust in GenAI. 22

Harm

Maybe harmed

RF Predicted CATE

0.4

0.3

0.2

Not treated (7.8%)

0.1

0.0 0.0

0.1

0.2

Not harmed

(b)

0.3

0.4

RF Predicted treated outcome

(a)

0.6

0.4

Not treated (7.8%) 0.2

0.5

0.0

RF Predicted harm rate

0.2

0.4

0.6

RF Predicted control outcome

Figure 7: Visualization of treatment decisions by CPL using the welfare-maximizing variant, where outcome models are estimated by causal forests. Colored dots are those with ĝ(X, T ) = 1, which provide a conservative bound for harm; red dots are those who can possibly be harmed, with Y † = 0. 1.00

Fraction of safe treatment

Value

0.75

0.50

Predicted treated outcome 0.25

Predicted control outcome 0.00 0

25

50

75

100

Pre−experiment belief

Figure 8: Fraction of safe treatment (red), average predicted treated (dark blue) and control (light blue) outcomes among test participants within a moving window of self-reported pre-experiment conspiracy belief. Fraction treated

Pred. control outcome

(a) 1.00

1.00

Fraction of safe treatment

0.75

Value

Value

0.75

0.50

Fraction of safe treatment

0.50

Predicted treated outcome

Predicted treated outcome

0.25

0.00

Pred. treated outcome

(b)

0.25

Predicted control outcome 2

4

0.00 6

Predicted control outcome 2

GenAI familiarity

4

6

GenAI trust

Figure 9: Fraction of treatment by the welfare-maximizing variant (red), average predicted treated (dark blue) and control (light blue) outcomes among test units stratified by self-reported GenAI-related variables.

23

8

Discussion

In this article, we developed the CPL approach that allows for individualized treatment assignment with a safety guarantee. We prove that when learning a treatment assignment rule from randomized experiments, CPL provides the safety guarantee in a finite sample without imposing any modeling assumption about potential outcomes. Additionally, when the outcome regression model is consistently estimated as assumed in many existing methods, CPL also asymptotically achieves the optimal welfare and power with appropriate choices of score V and inclusion g functions. We then extended our results to observational studies and derived novel asymptotic doubly robust safety guarantees for CPL. We now briefly discuss several natural extensions of our proposed conformal policy learning. The first concerns external validity settings where the test units may come from a different superpopulation than the labeled data and Assumption 2.2 is violated (e.g., Egami and Hartman, 2023; Jin et al., 2025a). The most common and natural strategy is to relax the i.i.d. assumption to the covariate shift assumption, i.e., the distributions of the labeled data and the test data differ only in the covariate distribution. Under this setting, we can generalize CPL by simply multiplying the current conformal weights by additional weights wQ/P (x) := dQX / dPX (x) that account for the covariate shift density ratio where P denotes the distribution of the labeled data and Q represents the distribution of the test data. Assuming both e(x) and wQ/P (x) are known, CPL then proceeds in exactly the same way as in Section 4. When either of them is unknown, one may use similar strategies as in Section 5 to estimate and plug in these quantities. The second natural extension concerns the control of harm rate among subgroups. In many problems where fairness and equity are stressed, it is desirable to maintain the harm rate control within each subgroup, such as those defined by demographic features (Romano et al., 2020). Following the setup in the main framework, our goal is to find the treatment assignment rule π̂(Xn+1 ) ∈ {0, 1} such that, for a group indicator G(·), P(Yn+1 (π(Xn+1 )) < Yn+1 (0) | G(Xn+1 ) = 1). Given a new sample with G(Xn+1 ) = 1, this can be achieved by taking the calibration data from the same subgroup, i.e., {(Xi , Ti , Yi ) : G(Xi ) = 1}. These data are induced by the full data in the subgroup {(Xi , Yi (1), Yi (0)) : G(Xi ) = 1} which is exchangeable with the new test sample conditional on the group indicator. The same CPL procedures can then be performed within the subgroup for randomized experiments and observational studies.

Acknowledgments The authors thank Eli Ben-Michael and Molly Offer-Westort for helpful discussions at the Online Causal Inference Seminar and the Political Methodology summer meeting, respectively. We also thank seminar participants at Yale Economics, Harvard Applied Stats Workshop, University of Tokyo Applied Stats Seminar, and the Political Methodology summer meeting. Y.J. is partially supported by NSF DMS-2610282.

24

References Athey, S. and Wager, S. (2021). Policy learning with observational data. Econometrica, 89(1):133–161. Bai, H., Voelkel, J. G., Muldowney, S., Eichstaedt, J. C., and Willer, R. (2025). LLM-generated Messages can Persuade Humans on Policy Issues. Nature Communications, 16(1):6037. Bai, T. and Jin, Y. (2024). Optimized conformal selection: Powerful selective inference after conformity score optimization. arXiv preprint arXiv:2411.17983. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., and Mariman, R. (2025). Generative AI without Guardrails can Harm Learning: Evidence from high School Mathematics. Proceedings of the National Academy of Sciences, 122(26):e2422633122. Bates, S., Candès, E., Lei, L., Romano, Y., and Sesia, M. (2021). Testing for outliers with conformal p-values. arXiv preprint arXiv:2104.08279. Ben-Michael, E., Greiner, D. J., Imai, K., and Jiang, Z. (2025). Safe policy learning through extrapolation: Application to pre-trial risk assessment. Journal of the American Statistical Association, 120(551):1386– 1399. Ben-Michael, E., Imai, K., and Jiang, Z. (2024). Policy learning with asymmetric counterfactual utilities. Journal of the American Statistical Association, 119(548):3045–3058. Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters. Costello, T. H., Pennycook, G., and Rand, D. G. (2024). Durably Reducing Conspiracy Beliefs through Dialogues with AI. Science, 385(6714):eadq1814. Dell’Acqua, F., McFowland III, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. (2026). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Organization Science. Dudı́k, M., Langford, J., and Li, L. (2011). Doubly robust policy evaluation and learning. arXiv preprint arXiv:1103.4601. Egami, N. and Hartman, E. (2023). Elements of external validity: Framework, design, and analysis. American Political Science Review, 117(3):1070–1088. Gadbury, G. L., Iyer, H. K., and Albert, J. M. (2004). Individual treatment effects in randomized trials with binary outcomes. Journal of Statistical Planning and Inference, 121(2):163–174. Hainmueller, J. (2012). Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies. Political analysis, 20(1):25–46. Heckman, J. J., Smith, J., and Clements, N. (1997). Making the Most Out of Programme Evaluations and Social Experiments: Accounting for Heterogeneity in Programme Impacts. The Review of Economic Studies, 64(4):487–535. Hirano, K. and Porter, J. R. (2009). Asymptotics for Statistical Treatment Rules. Econometrica, 77(5):1683– 1701. Holland, P. W. (1986). Statistics and Causal Inference. Journal of the American Statistical Association, 81(396):945–960. Huang, Y., Gilbert, P. B., and Janes, H. (2012). Assessing treatment-selection markers using a potential outcomes framework. Biometrics, 68(3):687–696. 25

Imai, K. and Strauss, A. (2011). Estimation of heterogeneous treatment effects from randomized experiments, with application to the optimal planning of the get-out-the-vote campaign. Political Analysis, 19(1):1–19. Imbens, G. W. and Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press. Jia, Z., Ben-Michael, E., and Imai, K. (2025). Bayesian safe policy learning with chance constrained optimization: Application to military security assessment during the vietnam war. Journal of the Royal Statistical Society Series A: Statistics in Society, page qnaf122. Jin, Y. and Candès, E. J. (2023a). Model-free selective inference under covariate shift via weighted conformal p-values. arXiv preprint arXiv:2307.09291. Jin, Y. and Candès, E. J. (2023b). Selection by prediction with conformal p-values. Journal of Machine Learning Research, 24(244):1–41. Jin, Y., Egami, N., and Rothenhäusler, D. (2025a). Beyond reweighting: On the predictive role of covariate shift in effect generalization. Proceedings of the National Academy of Sciences, 122(45):e2427181122. Jin, Y., Ren, Z., and Candès, E. J. (2023). Sensitivity analysis of individual treatment effects: A robust conformal inference approach. Proceedings of the National Academy of Sciences, 120(6):e2214889120. Jin, Y., Ren, Z., Yang, Z., and Wang, Z. (2025b). Policy Learning “without” Overlap: Pessimism and Generalized Empirical Bernstein’s Inequality. The Annals of Statistics, 53(4):1483–1512. Jin, Y. and Zubizarreta, J. (2025). Cross-balancing for data-informed design and efficient analysis of observational studies. arXiv preprint arXiv:2511.15896. Kallus, N. (2022). What’s the harm? sharp bounds on the fraction negatively affected by treatment. Advances in Neural Information Processing Systems, 35:15996–16009. Kitagawa, T. and Tetenov, A. (2018). Who should be treated? empirical welfare maximization methods for treatment choice. Econometrica, 86(2):591–616. Kleinberg, J., Lakkaraju, H., Leskovec, J., Ludwig, J., and Mullainathan, S. (2018). Human decisions and machine predictions. The quarterly journal of economics, 133(1):237–293. Kosorok, M. R. and Laber, E. B. (2019). Precision medicine. Annual review of statistics and its application, 6(1):263–286. Laber, E. B., Wu, F., Munera, C., Lipkovich, I., Colucci, S., and Ripa, S. (2018). Identifying optimal dosage regimes under safety constraints: An application to long term opioid treatment of chronic pain. Statistics in medicine, 37(9):1407–1418. Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., and Wasserman, L. (2018). Distribution-free predictive inference for regression. Journal of the American Statistical Association, 113(523):1094–1111. Lei, L. and Candès, E. J. (2021). Conformal inference of counterfactuals and individual treatment effects. Journal of the Royal Statistical Society Series B: Statistical Methodology, 83(5):911–938. Li, H., Zheng, C., Cao, Y., Geng, Z., Liu, Y., and Wu, P. (2023). Trustworthy policy learning under the counterfactual no-harm criterion. In International Conference on Machine Learning, pages 20575–20598. PMLR. Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010). A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web, pages 661–670.

26

Manski, C. F. (2004). Statistical Treatment Rules for Heterogeneous Populations. Econometrica, 72(4):1221– 1246. Murphy, S. A. (2003). Optimal Dynamic Treatment Regimes. Journal of the Royal Statistical Society Series B: Statistical Methodology, 65(2):331–355. Qian, M. and Murphy, S. A. (2011). Performance guarantees for individualized treatment rules. Annals of statistics, 39(2):1180. Richens, J., Beard, R., and Thompson, D. H. (2022). Counterfactual harm. Advances in Neural Information Processing Systems, 35:36350–36365. Romano, Y., Barber, R. F., Sabatti, C., and Candès, E. (2020). With malice toward none: Assessing uncertainty via equalized coverage. Harvard Data Science Review, 2(2):4. Rubin, D. B. (1980). Randomization Analysis of Experimental Data: The Fisher Randomization Test Comment. Journal of the American statistical association, 75(371):591–593. Scauda, M., Freidling, T., and Zhao, Q. (2026). Counterfactual optimization of policy interventions: Lexical ordering and leapfrogging. arXiv preprint arXiv:2608.20505. Shen, C., Jeong, J., Li, X., Chen, P.-S., and Buxton, A. (2013). Treatment benefit and treatment harm rate to characterize heterogeneity in treatment effect. Biometrics, 69(3):724–731. Skeem, J., Scurich, N., and Monahan, J. (2020). Impact of risk assessment on judges’ fairness in sentencing relatively poor defendants. Law and human behavior, 44(1):51. Tibshirani, R. J., Barber, R. F., Candès, E. J., and Ramdas, A. (2019). Conformal Prediction Under Covariate Shift. In Advances in Neural Information Processing Systems 32, pages 2526–2536. Vovk, V., Gammerman, A., and Shafer, G. (2005). Algorithmic learning in a random world. Springer Science & Business Media. Wang, Y., Fu, H., and Zeng, D. (2018). Learning optimal personalized treatment rules in consideration of benefit and risk: with an application to treating type 2 diabetes patients with insulin therapies. Journal of the American Statistical Association, 113(521):1–13. Wu, P., Ding, P., Geng, Z., and Liu, Y. (2024). Quantifying individual risk for binary outcome. arXiv preprint arXiv:2402.10537. Wu, P., Jiang, Q., Luo, S., and Geng, Z. (2025). Safe individualized treatment rules with controllable harm rates. arXiv preprint arXiv:2505.05308. Xu, Z. and Ramdas, A. (2024). Online multiple testing with e-values. In International Conference on Artificial Intelligence and Statistics, pages 3997–4005. PMLR. Yin, M., Shi, C., Wang, Y., and Blei, D. M. (2024). Conformal sensitivity analysis for individual treatment effects. Journal of the American Statistical Association, 119(545):122–135. Yin, Y., Liu, L., and Geng, Z. (2018). Assessing the treatment effect heterogeneity with a latent variable. Statistica Sinica, pages 115–135. Zhang, Y., Ben-Michael, E., and Imai, K. (2022). Safe policy learning under regression discontinuity designs with multiple cutoffs. arXiv preprint arXiv:2208.13323. Zhang, Z., Wang, C., Nie, L., and Soon, G. (2013). Assessing the heterogeneity of treatment effects via potential outcomes of individual patients. Journal of the Royal Statistical Society Series C: Applied Statistics, 62(5):687–704.

27

Zhao, Y., Zeng, D., Rush, A. J., and Kosorok, M. R. (2012). Estimating Individualized Treatment Rules using Outcome Weighted Learning. Journal of the American Statistical Association, 107(499):1106–1118. Zubizarreta, J. R. (2015). Stable weights that balance covariates for estimation with incomplete outcome data. Journal of the American Statistical Association, 110(511):910–922.

A

Deferred discussion

A.1

Power Maximization

We first define a natural notion of power, the probability of getting treated: Power(π; P ) := P (π(Xn+1 ) = 1). The optimal treatment rule under this power notion is defined as ∗ πpower = argmax

Power(π; PX,Y (1),Y (0) )

(A.1)

π : X →{0,1}

max Err(π; P ) ≤ α.

subject to

P ∈P

We can rewrite this optimization problem as follows. ∗ πpower = argmax

E[π(Xn+1 )]

π : X →{0,1}

subject to

E[π(Xn+1 )γ(Xn+1 )] ≤ α.

where γ(x) = min{1 − µ1 (x), µ0 (x)}, which is the sharp upper bound for P(Yn+1 (1) = 0, Yn+1 (0) = 1 | Xn+1 = x). Intuitively, given the linear relaxation, the optimal treatment rule is a thresholding rule based on γ(Xn+1 ), which acts as the “cost” in this optimization problem. ∗ under the worst-case harm constraint Theorem A.1 formally establishes that the optimal solution πpower is based on a cutoff on γ(x). It is implied by a more general result in Theorem A.4 in Appendix A.2 with proof in Appendix D.1; we thus omit the proof of Theorem A.1 here. ∗ ). Assume γ(X) has no point mass. Then, the optimal solution (A.1) Theorem A.1 (Power-optimal πpower ∗ ∗ is πpower (x) = 1{γ(x) ≤ γ } for the cutoff γ ∗ = max{γ̃ ∈ [0, 1] : E[γ(X) 1{γ(X) ≤ γ̃)}] ≤ α}.

To achieve the optimal power, we define the power-optimal conformal p-value (for stratified experiments, covering the complete randomized experiments in Section 3 as a special case) Pn w(Xn+1 ) + i=1 w(Xi )Gi 1{Yi† = 0} 1{γ̂(Xi ) ≤ γ̂(Xn+1 )} Pn pstr-power = , (A.2) n+1 w(Xn+1 ) + i=1 Gi w(Xi ) where Gi = Ti 1{1 − µ̂1 (Xi ) ≤ µ̂0 (Xi )} + (1 − Ti )1{1 − µ̂1 (Xi ) > µ̂0 (Xi )}. The choice of the calibration inclusion indicators Gi is the same as the welfare-optimal procedure. This coincides with Algorithm 1 with the clipped score V (x, y) = M 1{y > 0} + γ̂(x) using a sufficiently large constant M > 0. Theorem A.2 establishes the optimality of PCL with the power-oriented conformity score. Its proof is in Appendix D.2. Theorem A.2. Suppose Assumption 2.1 and Assumption 2.2 hold, and e(x) is known. Suppose ∥µ̂t (X) − P µt (X)∥L2 (PX ) → 0 as n → ∞ for t ∈ {0, 1}, and γ(X) has no point mass. Let π̂str-power (Xn+1 ) := ∗ 1{pstr-power ≤ α} for the p-value in (A.2). Then E[Y (π̂str-power (Xn+1 ))] → Power(πpower ; P) as n → ∞, n+1 ∗ where πpower is the power-optimal solution in Theorem A.1. The theorem below establishes asymptotic power optimality under suitable conditions, in parallel to Theorem 5.4. The proof is in Appendix D.3. 28

Theorem A.3 (Power optimality with observational data). Let π̂obs (Xn+1 ) = 1{pobs n+1 ≤ α}, where we take the same Gi as (3.12) and the score ŝ(x) = γ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)}. Suppose Assumption A.10 holds, and δn = O(n−1/2 ). Furthermore, assume ∥ĝ − g ∗ ∥L2 (PX,T ) + ∥γ̂ − γ∥L2 (PX ) + ∥ŝ − γ∥L2 (PX ) = oP (1) for g ∗ (x, t) = t 1{1 − µ1 (x) ≤ µ0 (x)} + (1 − t) 1{1 − µ1 (x) > µ0 (x)}, and γ(X) has no point mass. Then ∗ ∗ E[π̂obs (Xn+1 )] → Power(πpower ; P), where πpower is the power-optimal solution in Theorem A.1.

A.2

General form of optimality

Theorem A.4 is a general form of Theorem A.1, whose proof is in Appendix D.1. Theorem A.4 (Power-optimal ϕ∗ , general form). Let γ(x) = min{P(Y (1) = 0 | X = x), P(Y (0) = 1 | X = x)}. If E[γ(X)] ≤ α, then an optimal solution to (A.1) is ϕ∗power (x) ≡ 1. Otherwise, define γ ∗ := inf {c ∈ [0, 1] : E[γ(X)1{γ(X) ≤ c}] ≥ α}, and let η ∗ ∈ [0, 1] be chosen such that E[γ(X)1{γ(X) < γ ∗ }] + η ∗ E[γ(X)1{γ(X) = γ ∗ }] = α. Then an optimal solution to (A.1) is ϕ∗power (x) = 1{γ(x) < γ ∗ } + η ∗ 1{γ(x) = γ ∗ }. In this case, the safety constraint is binding: E[γ(X)ϕ∗power (X)] = α. Theorem A.5 is a general form of Theorem 3.3, whose proof is in Appendix D.4. Theorem A.5. Consider the randomized extension of problem (3.9), where a policy is a measurable function ϕ : X → [0, 1], and ϕ(x) denotes the probability of treatment conditional on X = x. For r ≤ 0, define H(r) := E[γ(X) 1{swelfare (X) < r}]. If H(0) ≤ α, let r∗ = 0 and η ∗ = 0. If H(0) > α, let r∗ := sup {r < 0 : H(r) ≤ α}, and choose η ∗ ∈ [0, 1] such that H(r∗ ) + η ∗ E[γ(X) 1{swelfare (X) = r∗ }] = α. Then an optimal solution to the randomized extension of (3.9) is ϕ∗welfare (x) = 1{swelfare (x) < r∗ } + η ∗ 1{swelfare (x) = r∗ }. Moreover, E[γ(X)ϕ∗welfare (X)] ≤ α and r∗ · {E[γ(X)ϕ∗welfare (X)] − α} = 0. In particular, if E[γ(X) 1{τ (X) > 0}] ≤ α, then r∗ = 0 and ϕ∗welfare (x) = 1{τ (x) > 0}. Otherwise, r∗ < 0 and the safety constraint is exactly attained.

A.3

A general theory for doubly robust guarantee with estimated weights

In this section, we present the general theory for CPL with estimated weights. We consider any estimated weights {ŵi }n+1 i=1 , which may come from a pre-trained weighted function or whose estimation may depend on the data. We recall the definition of conformal p-values: Pn ŵn+1 + i=1 ŵi Gi 1{Yi† = 0} 1{ŝ(Xi ) ≤ ŝ(Xn+1 )} Pn pobs = , n+1 ŵn+1 + i=1 ŵi Gi where the inclusion indicators are Gi ∼ Bern(ĝ(Xi , Ti )) for a function ĝ : X × {0, 1} → [0, 1] whose training process is independent of the labeled and unlabeled data. We introduce several high-level conditions that the weight function and other nuisance functions need to satisfy in order to achieve the doubly robust safety guarantee. Importantly, we will later show that our learn-then-balance weights satisfy these conditions under mild model convergence conditions. First, we require the weights to satisfy the following approximate balancing condition. This will involve an estimator γ̂(·) of the worst-case harm rate function γ(x) = min{1 − µ1 (x), µ0 (x)} estimated with data independent of the calibration and test data, such as by plugging in independently estimated outcome models. Assumption A.6 (Approximate balance). Let rn > 0 be a deterministic sequence obeying rn = o(1). P 1 2 2 The weights obey ŵi ≥ 0, |Icalib i∈Icalib (ŵi − ω̂(Xi )) = OP (rn ) for some function ω̂(·) : X → R trained | ŵ

∨max

ŵ

i i∈Icalib = OP (1/n). In addition, for independently of the calibration data and test point, and n+1P i∈Icalib ŵi P n 1 t̂ = sup{t ∈ R : n i=1 γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} ≤ α}, it holds that P Pn 1 1 i∈Icalib ŵi γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} − n i=1 γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} = OP (rn ), |Icalib |

29

and

1 |Icalib |

P

i∈Icalib ŵi = 1 + OP (rn ).

In Assumption A.6, the first condition requires the estimated weights {ŵi } to converge to any independently trained weight function with parametric rates. This is easily satisfied if one set ŵi = ŵ(Xi ); we shall see that our balancing program also satisfies this general condition. The key balancing condition requires the estimated weights to approximately balance the capped score functions at the critical cutoff t̂, following ideas in the balancing weights literature (Hainmueller, 2012; Zubizarreta, 2015; Jin and Zubizarreta, 2025) but with specific choice of the balancing features to yield favorable statistical properties. The following theorem shows that, under Assumptions A.6, the conformal policy learning π̂obs (Xn+1 ) := 1{pobs n+1 ≤ α} achieves the safety guarantees asymptotically if either the outcome conditional expectation function µ̂t (x) or the weight function ŵ(x), but not necessarily both, is consistently estimated. The proof of Theorem A.7 is in Appendix C.1. Recall that Tn is the σ-algebra of the training process. Theorem A.7 (Model double robustness). Suppose Assumption A.6 holds for some rn = o(1). Define pG := P(G = 1 | Tn ) and w◦ (x) := pG w(x). Then, lim sup P(Yn+1 (π̂obs (Xn+1 )) < Yn+1 (0)) ≤ α n→∞

under either of the following conditions: (i). The weight ω̂(·) obeys ∥ω̂ − w◦ ∥L2 (PX ) = oP (1), and there exist constants c, C > 0 such that pG ≥ c and ∥w◦ ∥L2 (PX ) ≤ C with probability tending to one. (ii). The outcome models obey ∥γ̂ − γ∥L2 (PX ) = oP (1), and there exist constants c, C, η > 0 such that, with probability tending to one, q(x) ≥ c, c ≤ ω̂(x) ≤ C, x ∈ X , and the function Hn (t) := E[γ̂(X) 1{ŝ(X) ≤ t} | Tn ] satisfies Hn (t◦n − u) ≤ α − cu and Hn (t◦n + u) ≥ α + cu for 0 < u ≤ η, where t◦n := sup{t ∈ R : Hn (t) ≤ α}, and P{a < ŝ(X) ≤ b | Tn } ≤ C(b − a) for every a < b. Notably, the doubly robust safety guarantees do not require the score function ŝ(·) to converge to any true function, which inherits the model-free nature of conformal prediction. The asymptotic safety guarantee hinges on the convergence of ŵ(·) and/or γ̂(·). In particular, in the case where the weights do not converge to w(·), consistent outcome models and the balancing condition still ensure harm rate control. We further show that CPL with balanced weights achieves rate double robustness: if both the outcome models and the weight functions are consistently estimated at slow, nonparametric rates oP (n−1/4 ), the excess √ harm rate above α is of the parametric order O(1/ n). The proof of Theorem A.9 is in Appendix C.2. Recall that pG := P(G = 1 | Tn ) and w◦ (x) := pG w(x). Assumption A.8 (Slow convergence). Let ω̂(·) be the reference function in Assumption A.6, and Assume ∥ω̂ − w◦ ∥L∞ (PX ) = OP (n−1/4 ) and ∥γ̂ − γ∥L1 (PX ) = OP (n−1/4 ). Theorem A.9 (Rate double robustness). Suppose Assumption A.6 holds with rn = O(n−1/2 ) and Assumption A.8 holds. Suppose there exist deterministic constants c, C, η > 0 such that, with probability tending to one, the following regularity conditions hold: (i) q(x) ≥ c and 0 ≤ γ̂(x) ≤ 1 for every x ∈ X ; (ii) the function Hn (t) := E[γ̂(X) 1{ŝ(X) ≤ t} | Tn ] satisfies Hn (t◦n − u) ≤ α − cu and Hn (t◦n + u) ≥ α + cu for all 0 < u ≤ η, where t◦n := sup{t ∈ R : Hn (t) ≤ α}; (iii) for every a < b, P{a < ŝ(X) ≤ b | Tn } ≤ C(b − a). Then     P Yn+1 π̂obs (Xn+1 ) < Yn+1 (0) An − α + = OP (n−1/2 ).

A.4

Technical conditions for the balancing algorithm

In this section, we provide the technical conditions and proofs that our learn-then-balance weights satisfy the needed conditions in the preceding parts, which together lead to the doubly robust safety guarantees in the 30

main text. We first define some preparatory notations. Let Tn denote the σ-field generated by the training process, so that ê, ĝ, γ̂, ŝ, and w̃ are fixed conditional on Tn . Let G be a generic inclusion indicator satisfying G | X, T, Tn ∼ Bernoulli{ĝ(X, T )}. Define q(x) := e(x)ĝ(x, 1) + {1 − e(x)}ĝ(x, 0), pG := E[q(X) | Tn ], and mn := |Icalib |. For any Tn -measurable functions r : X → [0, 1] and u : X → R+ , define ψtr,u (x) := (1, hrt (x), u(x))⊤ ,

hrt (x) := r(x) 1{ŝ(x) ≤ t}, and

t◦ (r) := sup {t ∈ R : E[hrt (X) | Tn ] ≤ α} . Also, let Mr,u (t) := E[ψtr,u (X)ψtr,u (X)⊤ | G = 1, Tn ],

br,u (t) := E[ψtr,u (X) | Tn ].

Whenever Mr,u (t) is invertible, define ωtr,u (x) := ψtr,u (x)⊤ Mr,u (t)−1 br,u (t),

Wn (r, u) := ωtr,u ◦ (r) .

For the fitted functions, abbreviate t◦ := t◦ (γ̂), M (t) := Mγ̂,w̃ (t), b(t) := bγ̂,w̃ (t), and ωt := ωtγ̂,w̃ . Also, write ψt := ψtγ̂,w̃ and define n n o 1 X γ̂ t̂ := sup t ∈ R : ht (Xi ) ≤ α , n i=1

M̂n (t) :=

1 X ψt (Xi )ψt (Xi )⊤ , mn i∈Icalib

Pn −1

 3 and ψ̄n (t) := n i=1 ψt (Xi ). Let Bn (t) := b ∈ R : b1 = 1, |bk − ψ̄n,k (t)| ≤ δn , k ∈ {2, 3} . Whenever M̂n (t) is invertible, define, for any x ∈ X and t ∈ R, b̂n (t) ∈ argmin b⊤ M̂n (t)−1 b,

ân (t) := M̂n (t)−1 b̂n (t),

ω̂t (x) := ψt (x)⊤ ân (t).

b∈Bn (t)

Assumption A.10 (Regularity of the balancing program). Recall that t◦ = t◦ (γ̂), M (t) = Mγ̂,w̃ (t), and ωt = ωtγ̂,w̃ . There exist deterministic constants c0 , c1 , c2 , C, η > 0 such that, with probability tending to one over the training process, the following conditions hold. (i) For every x ∈ X , q(x) ≥ c0 , 0 ≤ γ̂(x) ≤ 1, and 0 < w̃(x) ≤ C. (ii) Let N := {t ∈ R : |t − t◦ | ≤ η}. The population balancing problem is uniformly nondegenerate and its solution is uniformly interior on N : inf t∈N λmin {M (t)} ≥ c1 , and inf t∈N ,x∈X ωt (x) ≥ c1 . (iii) The population cutoff is locally regular: for every 0 < u ≤ η, it holds that E[ht◦ −u (X) | Tn ] ≤ α − c2 u, and E[ht◦ +u (X) | Tn ] ≥ α + c2 u. (iv) The conditional distribution of ŝ(X) has a uniformly bounded density in the sense that, for every a < b, P{a < ŝ(X) ≤ b | Tn } ≤ C(b − a). Under the above regularity conditions, the following lemma shows that the learned weights ŵi are close to the population balancing rule Wn (γ̂, w̃)(Xi ). Moreover, if the preliminary weight function consistently estimates the oracle weight w, then the population balancing rule consistently estimates the normalized oracle weight w◦ ; under an OP (n−1/4 ) uniform rate, it inherits the same rate. The proof is in Appendix C.3. Lemma A.11 (Approximation properties of the balancing weights). Suppose δn = O(n−1/2 ) and Assumption A.10 holds. Then, with probability tending to one, the balancing program has the unique solution ŵi = ω̂t̂ (Xi ) for i ∈ Icalib . Define ŵn+1 := ω̂t̂ (Xn+1 ). Then |t̂ − t◦ | = OP (n−1/2 ),

1 X 2 {ŵi − Wn (γ̂, w̃)(Xi )} = OP (n−1/2 ), mn i∈Icalib

Moreover, if ∥w̃ − w∥L2 (PX ) = oP (1), then ∥Wn (γ̂, w̃) − w◦ ∥L2 (PX ) = oP (1), 31

ŵn+1 ∨ maxi∈Icalib ŵi P = OP (n−1 ). i∈Icalib ŵi

where w◦ (x) = pG w(x). Finally, if ∥w̃ − w∥L∞ (PX ) = OP (n−1/4 ), then 1 X 2 {ŵi − Wn (γ̂, w̃)(Xi )} = OP (n−1 ), ∥Wn (γ̂, w̃) − w◦ ∥L∞ (PX ) = OP (n−1/4 ). mn i∈Icalib

B

Technical proofs

B.1

Proof of Theorem 3.1

Proof of Theorem 3.1. We aim to show that ∗ P(pn+1 ≤ α, Yn+1 = 0 | {Gi }ni=1 ) ≤ α.

(B.1)

Note that by definition, Yi† = Ti Yi +(1−Ti )(1−Yi ) ≤ max{Yi (1), 1−Yi (0)} = Yi∗ . Thus, by the monotonicity ∗ of V (x, y) in y, on the event {Yn+1 = 0}, it holds deterministically that Pn ∗ 1 + i=1 Gi · 1{V (Xi , Yi∗ ) ≤ V (Xn+1 , Yn+1 )} Pn pn+1 ≥ p∗n+1 := . 1 + i=1 Gi We thus have ∗ ∗ P(pn+1 ≤ α, Yn+1 = 0 | {Gi }ni=1 ) ≤ P(p∗n+1 ≤ α, Yn+1 = 0 | {Gi }ni=1 ) ≤ P(p∗n+1 ≤ α | {Gi }ni=1 ).

(B.2)

Now, conditional on {Gi }ni=1 , we denote I := {i ∈ [n] : Gi = 1} as the index set of the selected calibration sup ∗ data. Let PX,Y ∗ be the (unknown) joint distribution of (X, Y ) induced by the unknown super-population sup ∗ P sup of (Xi , Yi (1), Yi (0)). We then have (Xn+1 , Yn+1 ) ∼ PX,Y ∗ . For any measurable subset A of X × {0, 1}, by the Bayes’ rule, we have P((Xi , Yi∗ ) ∈ A | Gi = 1) P(Gi = 1 | (Xi , Yi∗ ) ∈ A) = . ∗ P((Xn+1 , Yn+1 ) ∈ A) P(Gi = 1) Furthermore, we know P(Gi = 1 | (Xi , Yi∗ ) ∈ A) = P(Gi = 1, Ti = 1 | (Xi , Yi∗ ) ∈ A) + P(Gi = 1, Ti = 0 | (Xi , Yi∗ ) ∈ A)) = P(Gi = 1 | Ti = 1, (Xi , Yi∗ ) ∈ A) · P(Ti = 1 | (Xi , Yi∗ ) ∈ A) + P(Gi = 1 | Ti = 0, (Xi , Yi∗ ) ∈ A) · P(Ti = 0 | (Xi , Yi∗ ) ∈ A) = 1/2 · E[ĝ(Xi , 1) | (Xi , Yi∗ ) ∈ A] + 1/2 · E[ĝ(Xi , 0) | (Xi , Yi∗ ) ∈ A]. Since ĝ(x, 1) + ĝ(x, 0) ≡ a for a constant a > 0, we have P(Gi = 1 | (Xi , Yi∗ ) ∈ A) = a/2, and by the tower property,     P(Gi = 1) = E P(Gi = 1 | Xi ) = E P(Gi = 1, Ti = 1 | Xi ) + P(Gi = 1, Ti = 0 | Xi ) = a/2. Putting things together, we know that d

∗ (Xi , Yi∗ ) | Gi = 1 = (Xn+1 , Yn+1 ).

Thus, conditional on {Gi }ni=1 , the samples {(Xi , Yi )}i∈I0 ∪{n+1} are exchangeable. Noting that P ∗ 1 + i∈I0 1{V (Xi , Yi∗ ) ≤ V (Xn+1 , Yn+1 )} ∗ pn+1 = , 1 + |I0 | we know that P(p∗n+1 ≤ α | {Gi }ni=1 ) ≤ α in (B.2); see, e.g., Jin and Candès (2023b) or Vovk et al. (2005). Marginalizing over {Gi }ni=1 , we complete the proof of (B.1). Finally, taking Tn+1 = 1{pn+1 ≤ α}, we have ∗ ∗ P(Yn+1 (Tn+1 ) < Yn+1 (0)) = P(Tn+1 = 1, Yn+1 = 0) = P(pn+1 ≤ α, Yn+1 = 0) ≤ α,

thereby concluding the proof of Theorem 3.1. 32

B.2

Proof of Theorem 3.4

Proof of Theorem 3.4. Write τ (x) := µ1 (x) − µ0 (x), τ̂ (x) := µ̂1 (x) − µ̂0 (x), and recall that γ(x) = min{1 − µ1 (x), µ0 (x)} and γ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)}. We simplify the notation to s∗ (x) := −

τ (x) = swelfare (x), γ(x)

ŝ(x) := −

τ̂ (x) . γ̂(x)

Let Tn denote the sigma-field generated by the independent training process for µ̂0 and µ̂1 . Conditional on Tn , the estimated functions µ̂t , γ̂, ŝ, and the inclusion rule in (3.10) are fixed. We first show that selective calibration consistently recovers the sharp worst-case harm-rate function. Define q1 (x) := 1 − µ1 (x),

q0 (x) := µ0 (x),

and analogously q̂1 (x) := 1 − µ̂1 (x) and q̂0 (x) := µ̂0 (x). Let â(x) ∈ {0, 1} denote the arm selected by (3.10), so that â(x) = 1 when q̂1 (x) ≤ q̂0 (x) and â(x) = 0 otherwise. Define the true proxy-label risk of the selected arm by γ n (x) := q1 (x) 1{â(x) = 1} + q0 (x) 1{â(x) = 0}. Since γ(x) = min{q0 (x), q1 (x)}, we have γ n (x) ≥ γ(x). Moreover, if a∗ (x) ∈ arg mina∈{0,1} qa (x), the optimality of â(x) for the estimated risks gives 0 ≤ γ n (x) − γ(x) = qâ(x) (x) − qa∗ (x) (x) ≤ qâ(x) (x) − q̂â(x) (x) + q̂a∗ (x) (x) − qa∗ (x) (x) ≤ 2 {|µ̂1 (x) − µ1 (x)| + |µ̂0 (x) − µ0 (x)|} . The assumed L2 (PX ) convergence and the Cauchy–Schwarz inequality therefore imply ∥γ n − γ∥L1 (PX ) = oP (1). We next establish convergence of the estimated ranking. The outcome-model consistency implies ∥τ̂ − τ ∥L2 (PX ) = oP (1) and ∥γ̂ − γ∥L2 (PX ) = oP (1), where the latter follows from the Lipschitz property of the minimum function. Under the ratio conventions in (3.10), if γ(X) = 0, then s∗ (X) must take one of the three values −∞, 0, or +∞. By the no-point-mass assumption for s∗ (X), we have P{γ(X) = 0} = 0. The continuous mapping theorem then gives ŝ(X)− s∗ (X) = oP (1), where X ∼ PX is independent of the training process. The preceding convergence also holds uniformly for threshold indicators. Indeed, for every ε > 0, sup P(1{ŝ(X) ≤ t} ̸= 1{s∗ (X) ≤ t} | Tn ) t∈R

≤ P(|ŝ(X) − s∗ (X)| > ε | Tn ) + sup P{|s∗ (X) − t| ≤ ε}. t∈R

The first term is oP (1). The second term converges to zero as ε ↓ 0 because s∗ (X) has no point mass. It follows that sup P(1{ŝ(X) ≤ t} ̸= 1{s∗ (X) ≤ t} | Tn ) = oP (1). t∈R

Now define the empirical curve Fn (t) :=

1+

† i = 0} 1{ŝ(Xi ) ≤ t} i=1 Gi 1{YP , n 1 + i=1 Gi

Pn

t ∈ R,

so that pwelfare = Fn {ŝ(Xn+1 )}. Conditional on Tn , the summands are i.i.d. and bounded, thus Lemma F.3 n+1 gives   E G 1{Y † = 0} 1{ŝ(X) ≤ t} Tn sup Fn (t) − = oP (1). E[G | Tn ] t∈R

33

Under balanced randomization and the given inclusion rule, we know E[G | X, Tn ] = 1/2. Furthermore, conditional on X = x, the probability that the selected proxy label is zero is q1 (x) when â(x) = 1 and q0 (x) when â(x) = 0. Consequently,   1 E G 1{Y † = 0} 1{ŝ(X) ≤ t} Tn = E[γ n (X) 1{ŝ(X) ≤ t} | Tn ] . 2 Therefore, sup Fn (t) − H n (t) = oP (1), t∈R

H n (t) := E[γ n (X) 1{ŝ(X) ≤ t} | Tn ] .

Define the oracle harm-cost curve H(t) := E[γ(X) 1{s∗ (X) ≤ t}] . Using 0 ≤ γ ≤ 1, we have sup |H n (t) − H(t)| ≤ ∥γ n − γ∥L1 (PX ) + sup E[γ(X) |1{ŝ(X) ≤ t} − 1{s∗ (X) ≤ t}| | Tn ] = oP (1). t∈R

t∈R

Combining the preceding displays gives supt∈R |Fn (t) − H(t)| = oP (1). The function H is continuous because its jump at any t ∈ R is E[γ(X) 1{s∗ (X) = t}] = 0. Thus, for the independent test point, pwelfare − H{s∗ (Xn+1 )} ≤ sup |Fn (t) − H(t)| + |H{ŝ(Xn+1 )} − H{s∗ (Xn+1 )}| = 0P (1). n+1 t∈R

We next translate this convergence into convergence of treatment decisions. Since τ (X) = 0 implies s∗ (X) = 0, the no-point-mass condition gives P{τ (X) = 0} = 0. Hence, 1{τ̂ (Xn+1 ) > 0} − 1{τ (Xn+1 ) > 0} → 0 in probability. We also claim that P(H{s∗ (X)} = α, τ (X) > 0) = 0. To see this, let Jα := {t < 0 : H(t) = α}. Since H is nondecreasing, Jα is an interval. If a < b lie in the interior of this interval, then 0 = H(b) − H(a) = E[γ(X) 1{a < s∗ (X) ≤ b}]. Since γ(X) > 0 almost surely, it follows that P{a < s∗ (X) ≤ b} = 0. The endpoints of Jα also have probability zero because s∗ (X) has no point mass. This proves the claim. Recall that π̂welfare (Xn+1 ) = 1{pwelfare ≤ α} 1{τ̂ (Xn+1 ) > 0}. n+1 The convergence pwelfare − H{s∗ (Xn+1 )} = oP (1), the convergence of the treatment-effect indicator, and the n+1 preceding zero-probability claim imply P{π̂welfare (Xn+1 ) ̸= π0 (Xn+1 )} → 0, where π0 (x) := 1{H(s∗ (x)) ≤ α} 1{τ (x) > 0}. It remains to identify π0 with the oracle policy in Theorem 3.3. Let r∗ ≤ 0 be the cutoff in that theorem. Suppose first that the safety constraint is nonbinding. Then r∗ = 0 and, since P{τ (X) = 0} = 0, H(0) = E[γ(X) 1{τ (X) > 0}] ≤ α. For every x such that τ (x) > 0, we have s∗ (x) < 0, and therefore H{s∗ (x)} ≤ H(0) ≤ α. It follows that ∗ π0 (X) = 1{τ (X) > 0} = 1{s∗ (X) ≤ 0} = πwelfare (X)

almost surely.

Suppose instead that the safety constraint is binding. Then r∗ < 0. By the continuity of H and the definition of r∗ , we have H(r∗ ) = α. By the monotonicity of H and the preceding zero-probability result for the level set {t < 0 : H(t) = α}, we know 1{H(s∗ (X)) ≤ α} = 1{s∗ (X) ≤ r∗ }

almost surely.

Moreover, the indicator 1{τ (X) > 0} is redundant on the event {s∗ (X) ≤ r∗ } because r∗ < 0. Consequently, ∗ π0 (X) = 1{s∗ (X) ≤ r∗ } = 1{swelfare (X) ≤ r∗ } = πwelfare (X)

almost surely.

∗ We have therefore shown that P{π̂welfare (Xn+1 ) ̸= πwelfare (Xn+1 )} → 0, which implies ∗ ∗ |E[Y {π̂welfare (Xn+1 )}] − Welfare(πwelfare ; P )| ≤ P{π̂welfare (Xn+1 ) ̸= πwelfare (Xn+1 )} → 0.

This completes the proof. 34

B.3

Proof of Theorem 4.1

Proof of Theorem 4.1. The p-value (4.2) can be equivalently written as pn+1 =

w(Xn+1 ) +

P

† i∈Icalib w(Xi ) 1{V (Xi , Yi ) ≤ V (Xn+1 , 0)}

P

i∈Icalib w(Xi ) + w(Xn+1 )

.

By definition, we have Yi† = Ti Yi + (1 − Ti )(1 − Yi ) ≤ max{Yi (1), 1 − Yi (0)} = Yi∗ . Thus, by the monotonicity ∗ of V (x, y) in y, on the event {Yn+1 = 0}, it holds deterministically that P ∗ w(Xn+1 ) + i∈Icalib w(Xi ) 1{V (Xi , Yi∗ ) ≤ V (Xn+1 , Yn+1 )} ∗ P pn+1 ≥ pn+1 := . w(X ) + w(X ) i n+1 i∈Icalib We thus have ∗ ∗ P(pn+1 ≤ α, Yn+1 = 0 | {Gi }ni=1 ) ≤ P(p∗n+1 ≤ α, Yn+1 = 0 | {Gi }ni=1 ) ≤ P(p∗n+1 ≤ α | {Gi }ni=1 ).

(B.3)

The goal is then to prove the RHS of (B.3) is upper bounded by α. Conditional on {Gi }ni=1 , the data (Xi , Yi∗ ) for i ∈ Icalib are i.i.d. and follow the distribution d

P calib := P(Xi ,Yi∗ ) | Gi =1 , whereas the test point (Xi , Yi∗ ) follows the distribution P test := PXi ,Yi∗ , and we recall that P denotes the true underlying super-population distribution. Let p(x, y ∗ ) be the density function of PX,Y ∗ with respect to some base measure. The density ratio between the two distributions is p(x, y ∗ ) p(x)p(y ∗ | x) dP test (x, y) = = , dP calib p(x, y ∗ | G = 1) p(x | G = 1)p(y ∗ | x, G = 1) where we define the random variable G = ĝ(X, T ), and p(y ∗ | x, G = 1) is the conditional density of Yi∗ given Xi = x and Gi = 1, etc. By unconfoundedness, Y ∗ is independent of T conditional on X. This leads to dP test p(x)p(y ∗ | x) p(x) P(G = 1) (x, y) = = = . calib dP p(x | G = 1)p(y ∗ | x) p(x | G = 1) P(G = 1 | X = x) Now, by the definition of the Gi ’s, we know P(G = 1 | X = x) = P(G = 1, T = 1 | X = x) + P(G = 1, T = 0 | X = x) = P(G = 1 | T = 1, X = x)P(T = 1 | X = x) + P(G = 1 | T = 0, X = x)P(T = 0 | X = x) = e(x)ĝ(x, 1) + (1 − e(x))ĝ(x, 0). This implies the density ratio between P test and P calib is given by dP test (x, y) = P(G = 1) · w(x). dP calib In other words, {(Xi , Yi∗ }i∈Icalib ∪{n+1} are weighted exchangeable (Tibshirani et al., 2019), and thus following Tibshirani et al. (2019), we know that ! P ∗ w(Xn+1 ) + i∈Icalib w(Xi ) 1{V (Xi , Yi∗ ) ≤ V (Xn+1 , Yn+1 )} n P P ≤ α {Gi }i=1 ≤ α. i∈Icalib w(Xi ) + w(Xn+1 ) This proves the desired upper bound on the RHS of (B.3) hence Theorem 4.1. 35

B.4

Proof of Theorem 4.2

Proof of Theorem 4.2. The optimality follows from the asymptotic convergence of the power/welfare due to the law of large numbers, as well as the optimal results in Theorems A.1 and 3.3. Write τ (x) := µ1 (x) − µ0 (x) and τ̂ (x) := µ̂1 (x) − µ̂0 (x), and recall that γ(x) = min{1 − µ1 (x), µ0 (x)} and γ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)}. Define s∗ (x) := −

τ (x) = swelfare (x), γ(x)

ŝ(x) := −

τ̂ (x) . γ̂(x)

Let Tn denote the sigma-field generated by the independent training process. Conditional on Tn , all estimated functions are fixed. As in the proof of Theorem 3.4, let q1 (x) := 1 − µ1 (x) and q0 (x) := µ0 (x), and let â(x) ∈ {0, 1} denote the arm selected by (3.12). Thus, â(x) = 1 when 1 − µ̂1 (x) ≤ µ̂0 (x) and â(x) = 0 otherwise. Define γn (x) := q1 (x) 1{â(x) = 1} + q0 (x) 1{â(x) = 0}. The optimality of â(x) for the estimated risks gives 0 ≤ γn (x) − γ(x) ≤ 2 {|µ̂1 (x) − µ1 (x)| + |µ̂0 (x) − µ0 (x)|} , and hence ∥γn − γ∥L1 (PX ) = oP (1). The same argument as in the proof of Theorem 3.4 also gives ŝ(X) − s∗ (X) = oP (1) and sup P(1{ŝ(X) ≤ t} ̸= 1{s∗ (X) ≤ t} | Tn ) = oP (1). t∈R

We now examine the effect of weighting. Define the conditional inclusion probability ρn (x) := e(x) 1{â(x) = 1} + {1 − e(x)} 1{â(x) = 0}, so that the weight in (4.1) can be written as wn (x) = ρn (x)−1 . For covariates X, the inclusion indicator is G = 1{T = â(X)}. Conditional on (X, Tn ), note the two identities   E[Gwn (X) | X, Tn ] = 1, E Gwn (X) 1{Y † = 0} | X, Tn = γn (X). Indeed, if â(X) = 1, then G = T , wn (X) = 1/e(X), and the second conditional expectation is 1 − µ1 (X). If â(X) = 0, then G = 1 − T , wn (X) = 1/{1 − e(X)}, and the second conditional expectation is µ0 (X). Define Pn wn (Xn+1 ) + i=1 Gi wn (Xi ) 1{Yi† = 0} 1{ŝ(Xi ) ≤ t} Pn Fn (t) := , t ∈ R, wn (Xn+1 ) + i=1 Gi wn (Xi ) so that pstr−opt = Fn {ŝ(Xn+1 )}. The preceding identities imply that the training-conditional population n+1 counterpart of Fn is H n (t) := E[γn (X) 1{ŝ(X) ≤ t} | Tn ] . We now briefly verify the uniform law of large numbers. Let Z = Gwn (X). Conditional on Tn , E[Z | Tn ] = 1. Moreover, for every M > 0, we know E[(Z − M )+ | Tn ] = E[(1 − M ρn (X))+ | Tn ] ≤ E[(1 − M e(X))+ ] + E[(1 − M {1 − e(X)})+ ], which converges to zero as M → ∞ because e(X) ∈ (0, 1) almost surely. Applying Lemma F.3 to the bounded truncations and then letting M → ∞ therefore gives a uniform law of large numbers for the numerator of Fn ; the denominator follows by the same argument. In addition, wn (Xn+1 )/n = oP (1) because wn (X) ≤ max{e(X)−1 , {1 − e(X)}−1 } < ∞ almost surely. Consequently, sup |Fn (t) − H n (t)| = oP (1). t∈R

Define the oracle harm-cost curve H(t) := E[γ(X) 1{s∗ (X) ≤ t}]. Using 0 ≤ γ ≤ 1, we have sup |H n (t) − H(t)| ≤ ∥γn − γ∥L1 (PX ) + sup E[γ(X) |1{ŝ(X) ≤ t} − 1{s∗ (X) ≤ t}| | Tn ] = oP (1). t∈R

t∈R

36

It follows that sup |Fn (t) − H(t)| = oP (1). t∈R

The function H is continuous because its jump at any t ∈ R is E[γ(X) 1{s∗ (X) = t}] = 0. Hence, for − H{s∗ (Xn+1 )} = oP (1). Finally, the outcome-model consistency and the independent test point, pstr-opt n+1 P{τ (X) = 0} = 0 imply 1{τ̂ (Xn+1 ) > 0} − 1{τ (Xn+1 ) > 0} = oP (1). As shown in the proof of Theorem 3.4, the no-point-mass condition also implies P(H{s∗ (X)} = α, τ (X) > 0) = 0. Therefore, we have P{π̂str-opt (Xn+1 ) ̸= π0 (Xn+1 )} → 0, where π0 (x) := 1{H(s∗ (x)) ≤ α} 1{τ (x) > 0}. The binding and ∗ nonbinding arguments in the proof of Theorem 3.4 show that π0 (X) = πwelfare (X) almost surely. Thus, ∗ P{π̂str−opt (Xn+1 ) ̸= πwelfare (Xn+1 )} → 0.

With similar arguments as the end of the proof of Theorem 3.4, we complete the proof.

B.5

Proof of Theorem 5.1

Proof of Theorem 5.1. By Lemma A.11, the learn-then-balance weights satisfy Assumption A.6 with rn = n−1/4 and reference function ω̂(·) = Wn (γ̂, w̃)(·). Under condition (i), Lemma A.11 gives ∥ω̂ − w◦ ∥L2 (PX ) = oP (1), where w◦ = pG w ∝ w. Hence condition (i) of Theorem A.7 holds. Under condition (ii), Assumption A.10 verifies the overlap, interiority, local-slope, and bounded-density conditions in condition (ii) of Theorem A.7. The conclusion therefore follows from Theorem A.7 in either case.

B.6

Proof of Theorem 5.3

Proof of Theorem 5.3. Set ω̂ := Wn (γ̂, w̃). By Lemma A.11, the learn-then-balance weights satisfy Assumption A.6 with rn = n−1/2 and ∥ω̂ − w◦ ∥L∞ (PX ) = OP (n−1/4 ), where w◦ (x) := pG w(x). Together with the assumed rate for γ̂, this verifies Assumption A.8. Assumption A.10 verifies the remaining regularity conditions in Theorem A.9, which gives the result.

B.7

Proof of Theorem 5.4

Proof of Theorem 5.4. Let Hn (t) := E[γ̂(X) 1{ŝ(X) ≤ t} | Tn ] ,

t◦n := sup{t ∈ R : Hn (t) ≤ α},

and define H0 (t) := E[γ(X) 1{swelfare (X) ≤ t}] . By the assumed convergence of γ̂ and ŝ, the no-point-mass condition on swelfare (X), and Lemma F.4, sup |Hn (t) − H0 (t)| = oP (1). t∈R

Together with the local-crossing condition in Assumption A.10(iii), this implies t◦n − t∗raw = oP (1), where t∗raw = sup{t : H0 (t) ≤ α}. Indeed, for every fixed 0 < ε ≤ η, with probability tending to one, it holds that H0 (t◦n − ε) < α < H0 (t◦n + ε), and hence t◦n − ε ≤ t∗raw ≤ t◦n + ε. Set ω̂ := Wn (γ̂, w̃) and, writing It (x) := 1{ŝ(x) ≤ t}, define Fnγ̂ (t) :=

E[ω̂(X)ĝ(X, T )γ̂(X)It (X) | Tn ] . E[ω̂(X)ĝ(X, T ) | Tn ]

37

By the defining balance equations for ω̂, we know Fnγ̂ (t◦n ) = Hn (t◦n ) = α. Moreover, Assumption A.10(i)–(iii) implies that, for some deterministic κ > 0, with probability tending to one, Fnγ̂ (t◦n − u) ≤ α − κu and Fnγ̂ (t◦n + u) ≥ α + κu hold for every 0 < u ≤ η. Define Pn ŵn+1 + i=1 Gi ŵi 1{Yi† = 0}It (Xi ) obs Pn F̂n (t) := , ŵn+1 + i=1 Gi ŵi obs Similar to the arguments in the proof of Theorem A.7, Lemma A.11, so that pobs n+1 = F̂n {ŝ(Xn+1 )}. P Lemmas C.1 and F.4, and ŵn+1 / i∈Icalib ŵi = OP (n−1 ) give

sup F̂nobs (t) − Fn† (t) = oP (1), t∈R

Fn† (t) :=

E[ω̂(X)ĝ(X, T )m(X, T ) 1{ŝ(X) ≤ t} | Tn ] , E[ω̂(X)ĝ(X, T ) | Tn ]

where we define m(x, t) = P(Y † = 0 | X = x, T = t) = t(1 − µ1 (x)) + (1 − t)µ0 (x). Furthermore, by the definition of g ∗ , we know E[g ∗ (X, T ) 1{Y † = 0} | X] = E[g ∗ (X, T )m(X, T ) | X] = E[g ∗ (X, T )γ(X) | X]. Thus, sup |F̂nobs (t) − Fnγ̂ (t)| ≤ oP (1) + C∥ĝ − g ∗ ∥L1 (PX,T ) + C∥γ̂ − γ∥L1 (PX ) = oP (1). t∈R

Let π̂obs,raw (Xn+1 ) := 1{pobs n+1 ≤ α}. Fix 0 < ε ≤ η. The preceding results imply that, with probability tending to one, ŝ(Xn+1 ) ≤ t◦n − ε

⇒

pobs n+1 ≤ α,

ŝ(Xn+1 ) ≥ t◦n + ε

⇒

pobs n+1 > α.

Therefore, we have |π̂obs,raw (Xn+1 ) − 1{ŝ(Xn+1 ) ≤ t◦n }| ≤ 1{|ŝ(Xn+1 ) − t◦n | ≤ ε} + oP (1). Assumption A.10(iv) and the arbitrariness of ε > 0 then imply E[|π̂obs,raw (Xn+1 ) − 1{ŝ(Xn+1 ) ≤ t◦n }|] → 0. Since multiplication by an indicator cannot increase the absolute difference, the preceding display gives E[|π̂obs−welfare (Xn+1 ) − 1{ŝ(Xn+1 ) ≤ t◦n } 1{ŝ(Xn+1 ) < 0}|] → 0. We next pass to the population limit. The assumptions ∥ŝ − swelfare ∥L2 (PX ) = oP (1) and t◦n − t∗raw = oP (1), together with the no-point-mass condition on swelfare (X), imply i h E 1{ŝ(Xn+1 ) ≤ t◦n } 1{ŝ(Xn+1 ) < 0} − 1{swelfare (Xn+1 ) ≤ t∗raw } 1{swelfare (Xn+1 ) < 0} → 0. Indeed, for any ε > 0, disagreement between the first threshold indicators is contained in the union of {|ŝ − swelfare | > ε}, {|t◦n − t∗raw | > ε}, and {|swelfare − t∗raw | ≤ 2ε}. The corresponding argument at the threshold zero establishes convergence of the second indicators. It remains to identify the limiting rule with the oracle policy in Theorem 3.3. Let r∗ ≤ 0 denote the cutoff therein. We consider two cases: • If the safety constraint is binding, then H0 (0) > α, so t∗raw < 0 and t∗raw = r∗ . In this case, 1{swelfare < 0} is redundant on {swelfare ≤ t∗raw }, and hence 1{swelfare (X) ≤ t∗raw } 1{swelfare (X) < ∗ 0} = 1{swelfare (X) ≤ r∗ } = πwelfare (X) almost surely. • If the safety constraint is nonbinding, then H0 (0) ≤ α, so r∗ = 0 and t∗raw ≥ 0. Therefore, 1{swelfare (X) ≤ t∗raw } 1{swelfare (X) < 0} = 1{swelfare (X) < 0}. Since swelfare (X) has no point mass at zero, this equals ∗ 1{swelfare (X) ≤ 0} = πwelfare (X) almost surely. ∗ Combining the preceding results gives E[|π̂obs−welfare (Xn+1 ) − πwelfare (Xn+1 )|] → 0. Finally, since the outcomes are binary, we know ∗ ∗ |E[Y {π̂obs−welfare (Xn+1 )}] − Welfare(πwelfare ; P )| ≤ E[|π̂obs−welfare (Xn+1 ) − πwelfare (Xn+1 )|] → 0,

which completes the proof. 38

C

Proof of general balancing theory

C.1

Proof of Theorem A.7

Proof of Theorem A.7. Throughout this proof, we condition on the training process of the functions, so ŝ and ĝ are viewed as fixed, as well as the function ŵ(·) in the first condition of Assumption A.6. Once we prove P(Yn+1 (πobs (Xn+1 ) < Yn+1 (0) | ŝ, ĝ, ŵ) ≤ α + oP (1) conditional on these randomness, since R is uniformly bounded we also have the marginal harm rate P(Yn+1 (πobs (Xn+1 ) < Yn+1 (0)) ≤ α + oP (1). For notational simplicity, we shall omit the conditioning in the probability/expectation throughout the proof. Define τ̂ := sup{t ∈ R : F̂n (t) ≤ α},

F̂n (t) :=

ŵn+1 +

† i∈Icalib ŵi 1{Yi = 0} 1{ŝ(Xi ) ≤ t}

P

P

i∈Icalib ŵi + ŵn+1

,

so that pn+1 = F̂n (ŝ(Xn+1 )). Thus, pn+1 ≤ α is equivalent to ŝ(Xn+1 ) ≤ τ̂ . We now construct a random variable Yi∗∗ whose conditional expectation is the worst-case harm rate function γ(Xi ) and upper bounds Yi† . For t ∈ {0, 1}, define Di (t) := tYi (1) + (1 − t){1 − Yi (0)},

qt (x) := P(Di (t) = 0 | Xi = x),

so that Di (Ti ) = Yi† , q1 (x) = 1 − µ1 (x), q0 (x) = µ0 (x), and our previously defined oracle label satisfies Yi∗ = max{Di (0), Di (1)}. Let ν(x) := P(Yi∗ = 0 | Xi = x). Since {Yi∗ = 0} ⊆ {Di (t) = 0} for each t ∈ {0, 1}, we have ν(x) ≤ γ(x) = min{q0 (x), q1 (x)} ≤ qt (x). Define

  γ(x) − ν(x) , ρt (x) := qt (x) − ν(x)  0,

qt (x) > ν(x), qt (x) = ν(x).

Then ρt (x) ∈ [0, 1]; moreover, if qt (x) = ν(x), then necessarily γ(x) = ν(x). On an enlarged probability i.i.d.

space, let Ui ∼ Unif(0, 1) be independent of all existing random variables, and define Yi∗∗ := Yi∗ − 1 {Yi∗ = 1, Di (Ti ) = 0, Ui ≤ ρTi (Xi )} . By construction, Yi∗∗ ≤ Yi∗ . Furthermore, if Yi† = 1, then Di (Ti ) = 1, so the indicator in the preceding display vanishes and Yi∗∗ = Yi∗ = 1. Hence Yi† ≤ Yi∗∗ ≤ Yi∗

almost surely.

By unconfoundedness, for each t ∈ {0, 1},  P(Yi∗∗ = 0 | Xi = x, Ti = t) = ν(x) + ρt (x)P Yi∗ = 1, Di (t) = 0 | Xi = x = ν(x) + ρt (x){qt (x) − ν(x)} = γ(x). Thus, in particular, P(Yi∗∗ = 0 | Xi ) = γ(Xi ). To avoid introducing a cutoff that depends on the test point, define Sn := {Sn > 0}, define the calibration-only functions † i=1 Gi ŵi 1{Yi = 0} 1{ŝ(Xi ) ≤ t}

Pn F̂n,0 (t) :=

Sn

,

∗ F̂n,0 (t) :=

Pn

i=1 Gi ŵi , and, on the event

Pn

∗∗ = 0} 1{ŝ(Xi ) ≤ t} i=1 Gi ŵi 1{Yi

Sn

.

Under the assumptions of the theorem, Sn > 0 with probability tending to one. Moreover, assuming that the ∗ estimated weights are nonnegative, both F̂n,0 and F̂n,0 are nondecreasing functions taking values in [0, 1]. 39

The observed p-value thus obeys pobs n+1 =

 ŵn+1 + Sn F̂n,0 ŝ(Xn+1 ) ≥ F̂n,0 (ŝ(Xn+1 )) ŵn+1 + Sn

since F̂n,0 (t) ∈ [0, 1] for any t ∈ R. Furthermore, since Yi† ≤ Yi∗∗ almost surely, we have ∗ F̂n,0 (t) ≥ F̂n,0 (t)

for every t ∈ R.

It follows that   ∗ {pobs n+1 ≤ α} ⊆ F̂n,0 (ŝ(Xn+1 )) ≤ α ⊆ F̂n,0 (ŝ(Xn+1 )) ≤ α . For later use, define the calibration-only cutoffs ∗ τ̂0∗ := sup{t ∈ R : F̂n,0 (t) ≤ α}.

τ̂0 := sup{t ∈ R : F̂n,0 (t) ≤ α},

The preceding pointwise ordering implies τ̂0 ≤ τ̂0∗ . Moreover, ∗ {pobs n+1 ≤ α} ⊆ {ŝ(Xn+1 ) ≤ τ̂0 } ⊆ {ŝ(Xn+1 ) ≤ τ̂0 }.

Importantly, both τ̂0 and τ̂0∗ depend only on the calibration data, together with the auxiliary randomness used to construct Yi∗∗ , and do not depend on the test point. Let An denote the σ-field generated by these quantities. Conditional on An , and writing  Rn := P Yn+1 (1) = 0, Yn+1 (0) = 1, pobs n+1 ≤ α An , the sharp conditional upper bound P(Y (1) = 0, Y (0) = 1 | X) ≤ γ(X) gives h i  ∗ Rn ≤ E γ(Xn+1 ) 1 F̂n,0 (ŝ(Xn+1 )) ≤ α An ≤ E [γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] . ∗ Since F̂n,0 (t) is a right-continuous and non-decreasing step function, letting ŝ(X[1] ) ≤ ŝ(X[2] ) ≤ · · · ≤ ŝ(X[n] ) be the order statistics and [1], [2], . . . , [n] is a permutation of (1, . . . , n), we consider two cases: ∗ ∗ • When there exists some k ∈ [n] such that F̂n,0 (ŝ(X[k] )) = α, we know τ̂0∗ = ŝ(X[k] ), thus F̂n,0 (τ̂0∗ ) = α. ∗ ∗ • When there exists some k ∈ [n] such that F̂n,0 (ŝ(X[k−1] )) < α < F̂n,0 (ŝ(X[k] )), we know τ̂0∗ = ŝ(X[k] ), ∗ ∗ ∗ and α < F̂n,0 (τ̂ ∗ ) ≤ α + F̂n,0 (X[k] ) − F̂n,0 (X[k−1] ) = α + P

Lemma A.11.

ŵ[k−1] i∈Icalib ŵi +ŵn+1

= α + OP (1/n) by

∗ Combining the two cases yields α ≤ F̂n,0 (τ̂0∗ ) ≤ α + OP (1/n).

The above arguments yield ∗ Rn ≤ E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] + α − F̂n,0 (τ̂0∗ ) + OP (1/n) Pn Gi ŵi 1{Yi∗∗ = 0} 1{ŝ(Xi ) ≤ τ̂0∗ } Pn ≤ α + E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] − i=1 + OP (1/n), (C.1) i=1 Gi ŵi

We shall repeatedly invoke the following lemma, whose proof is in Appendix F.1. Lemma C.1. Under Assumption A.6, for any random variable {Zi } such that (Xi , Yi , Ti , Zi ) are i.i.d. across i ∈ [n] and E[Z 2 ] < ∞, and any (possibly random) function f : X → R, it holds that n n 1X 1X ŵi Zi 1{f (Xi ) ≤ t} − ω̂(Xi )Zi 1{f (Xi ) ≤ t} = OP (rn ). n i=1 t∈R n i=1

sup

40

Pn Pn Lemma C.1 with rn = o(1) implies n1 i=1 Gi ŵi = n1 i=1 Gi ω̂(Xi ) + OP (rn ), and together with Zi = Gi 1{Yi∗∗ = 0} and f = ŝ it implies Pn Pn ∗∗ = 0} 1{ŝ(Xi ) ≤ t} Gi ω̂(Xi ) 1{Yi∗∗ = 0} 1{ŝ(Xi ) ≤ t} i=1 Gi ŵi 1{Y Pin Pn sup = OP (rn ). − i=1 t∈R i=1 Gi ŵi i=1 Gi ω̂(Xi ) Taking t = τ̂0∗ in the above display, and following (C.1), we have Pn Gi ω̂(Xi ) 1{Yi∗∗ = 0, ŝ(Xi ) ≤ τ̂0∗ } Pn Rn ≤ α + E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] − i=1 + OP (rn + 1/n). i=1 Gi ω̂(Xi ) (C.2) Consider Zi := Gi ω̂(Xi ){1(Yi∗∗ = 0) − γ(Xi )}. Since Gi is generated independently of (Yi (0), Yi (1), Ui ) conditional on (Xi , Ti ), and P(Yi∗∗ = 0 | Xi , Ti ) = γ(Xi ), we have E[Zi | Xi ] = ω̂(Xi )E [E [Gi {1(Yi∗∗ = 0) − γ(Xi )} | Xi , Ti ]|Xi ] = ω̂(Xi )E [ĝ(Xi , Ti ){P(Yi∗∗ = 0 | Xi , Ti ) − γ(Xi )}|Xi ] = 0. Thus, invoking Lemma F.3 with this Zi and s(·) = ŝ(·) yields Pn ∗∗ √ = 0} − γ(Xi )] 1{ŝ(Xi ) ≤ t} i=1 Gi ω̂(Xi )[1{Y Pi n = OP (1/ n). sup t∈R i=1 Gi ω̂(Xi ) Continuing with (C.2), this implies Rn ≤ α + E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] −

Pn

√ )γ(Xi ) 1{ŝ(Xi ) ≤ τ̂0∗ } i=1 Gi ω̂(X Pin + OP (1/ n + rn ). i=1 Gi ω̂(Xi ) (C.3)

Case 1: ∥ω̂ − w◦ ∥L2 (PX ) = oP (1). Let It (x) := 1{ŝ(x) ≤ t}. Conditional on Tn , uniform empirical convergence gives Pn ω̂(X )γ(Xi )It (Xi ) E[q(X)ω̂(X)γ(X)It (X) | Tn ] i=1 G Pi n i sup − = oP (1). E[q(X)ω̂(X) | Tn ] t∈R i=1 Gi ω̂(Xi ) Since 0 ≤ q, γ, It ≤ 1, we have sup |E[q(X){ω̂(X) − w◦ (X)}γ(X)It (X) | Tn ]| = oP (1), t∈R

and E[q(X){ω̂(X) − w◦ (X)} | Tn ] = oP (1). Moreover, q(x)w◦ (x) = pG , and therefore E[q(X)w◦ (X)γ(X)It (X) | Tn ] = E[γ(X)It (X) | Tn ]. E[q(X)w◦ (X) | Tn ] Since pG is bounded away from zero, it follows that Pn ω̂(X )γ(Xi )It (Xi ) i=1 G Pi n i sup − E[γ(X)It (X) | Tn ] = oP (1). t∈R i=1 Gi ω̂(Xi ) Evaluating this display at t = τ̂0∗ in (C.3) gives Rn ≤ α + oP (1). Case 2: ∥γ̂ − γ∥L2 (PX ) = oP (1).

For ξ ∈ {γ̂, γ}, define the training-conditional population curves

Hnξ (t) := E[ξ(X) 1{ŝ(X) ≤ t} | Tn ],

Fnξ (t) :=

41

E[ω̂(X)ĝ(X, T )ξ(X) 1{ŝ(X) ≤ t} | Tn ] . E[ω̂(X)ĝ(X, T ) | Tn ]

Let t◦n := sup{t ∈ R : Hnγ̂ (t) ≤ α} and recall Pn Gi ŵi γ̂(Xi ) 1{ŝ(Xi ) ≤ t} Pn . Q̂n (t) := i=1 i=1 Gi ŵi Lemmas C.1 and F.4, applied to the numerators and denominators, give sup |Q̂n (t) − Fnγ̂ (t)| = oP (1), t∈R

∗ sup |F̂n,0 (t) − Fnγ (t)| = oP (1). t∈R

Here the second display also uses P(Yi∗∗ = 0 | Xi , Ti ) = γ(Xi ). Moreover, since the denominator of Fnξ is bounded away from zero, sup |Fnγ (t) − Fnγ̂ (t)| + sup |Hnγ (t) − Hnγ̂ (t)| ≤ C∥γ̂ − γ∥L1 (PX ) = oP (1). t∈R

t∈R

By Lemma F.4 and the assumed local slope, t̂ − t◦n = OP (n−1/2 ). The balancing condition in Assumption A.6, together with the definition of t̂ and the fact that the maximal jump of its defining empirical curve is at most n−1 , gives Q̂n (t̂) = α + oP (1). Hence Fnγ̂ (t̂) = α + oP (1) and, by the bounded-density condition, Fnγ̂ (t◦n ) = α + oP (1). The same curve has a positive local slope at t◦n . Indeed, for t1 < t2 in a sufficiently small neighborhood of t◦n , the lower bounds on q and ω̂, together with their boundedness, imply Fnγ̂ (t2 ) − Fnγ̂ (t1 ) ≥ c{Hnγ̂ (t2 ) − Hnγ̂ (t1 )} for some deterministic c > 0. Therefore, for every fixed ε > 0, with probability tending to one, Fnγ (t◦n − ε) < α < Fnγ (t◦n + ε). ∗ ∗ yields Combining this strict separation with supt |F̂n,0 (t) − Fnγ (t)| = oP (1) and the monotonicity of F̂n,0

τ̂0∗ − t◦n = oP (1). Finally, conditional on An , the test point is independent of the training and calibration data. Thus, by (C.3) and the preceding uniform convergence, Rn ≤ α + Hnγ (τ̂0∗ ) − Fnγ (τ̂0∗ ) + oP (1). The bounded-density condition, τ̂0∗ − t◦n = oP (1), and the uniform differences between the γ- and γ̂-curves imply Hnγ (τ̂0∗ ) = α + oP (1) and Fnγ (τ̂0∗ ) = α + oP (1). Consequently, Rn ≤ α + oP (1). Marginalizing over An completes the proof in this case.

C.2

Proof of Theorem A.9

Proof of Theorem A.9. We reuse the proof of Theorem 5.1 up to (C.3) to prove the rate double robustness result. Note that the results up to (C.3) did not use any assumptions on the fitted models. Define ω̂ † (x) :=

ω̂(x) . pG

Since q(x) ≥ c, we have pG = E[q(X) | Tn ] ≥ c. Therefore, Assumption A.8 implies ∥ω̂ † − w∥L∞ (PX ) = ∥ω̂ − w◦ ∥L∞ (PX ) /pG = OP (n−1/4 ). Moreover, replacing ω̂ by ω̂ † does not change any weighted empirical or population ratio. Recall that An is the σ-field that includes the randomness in the labeled data, the training process, and generating {Gi }, 42

and we denote the An -conditional harm rate Rn := P(Yn+1 (πobs (Xn+1 )) < Yn+1 (0) | An ). The balancing condition in Assumption A.6 with rn = O(n−1/2 ), together with (C.3), implies Pn Gi ω̂ † (Xi )γ(Xi ) 1{ŝ(Xi ) ≤ τ̂0∗ } Pn Rn ≤ α + E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] − i=1 † i=1 Gi ω̂ (Xi ) Pn n √ Gi ω̂ † (Xi )γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} 1X Pn − + i=1 γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} + OP (1/ n), † n i=1 i=1 Gi ω̂ (Xi ) where the second and the third terms invoked Lemma C.1 twice. By Assumption A.6, we have n n √ √ √ √ 1X mn 1X Gi ω̂(Xi ) = Gi ŵi + OP (1/ n) = {1 + OP (1/ n)} + OP (1/ n) = pG + OP (1/ n). n i=1 n i=1 n

√

Pn

† i=1 Gi ω̂ (Xi ) = 1 + OP (1/

n), which further yields

Rn ≤ α + E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] −

1X γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} n i=1

Consequently, n1

n

+

n n √ 1X 1X Gi ω̂ † (Xi )γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} − Gi ω̂ † (Xi )γ(Xi ) 1{ŝ(Xi ) ≤ τ̂0∗ } + OP (1/ n) n i=1 n i=1

≤ α + E[γ(X) 1{ŝ(X) ≤ τ̂0∗ } | An ] − E[γ̂(X) 1{ŝ(X) ≤ t̂} | An ] n n √ 1X 1X + ĝ(Xi , Ti )ω̂ † (Xi )γ̂(Xi ) 1{ŝ(Xi ) ≤ t̂} − ĝ(Xi , Ti )ω̂ † (Xi )γ(Xi ) 1{ŝ(Xi ) ≤ τ̂0∗ } + OP (1/ n), n i=1 n i=1 where X ∼ PX is an independent copy. Here, the second inequality uses Lemma F.3 for Zi = (Ĝi − ĝ(Xi , Ti ))ω̂ † (Xi )γ(Xi ) or Zi = (Ĝi − ĝ(Xi , Ti ))ω̂ † (Xi )γ̂(Xi ), as well as Lemma F.3 for Zi = γ̂(Xi ) and Hn (t) = E[γ̂(X) 1{ŝ(X) ≤ t} | Tn ] evaluated at t = t̂ which is adapted to the σ-field An . For the oracle weight w(X), we know E[Gi w(Xi )] = E[ĝ(X, T )w(X)] = 1, and due to the covariate shift between calibration data conditional on Gi = 1 and the test data, E[γ(Xn+1 ) 1{ŝ(Xn+1 ) ≤ τ̂0∗ } | An ] = E[ĝ(X, T )w(X)γ(X) 1{ŝ(X) ≤ τ̂0∗ } | An ], E[γ̂(Xn+1 ) 1{ŝ(Xn+1 ) ≤ t̂} | An ] = E[ĝ(X, T )w(X)γ̂(X) 1{ŝ(X) ≤ t̂} | An ], where the expectations are over an independent copy (X, T ) ∼ PX,T . Thus, denoting  r̂(x, t) = ĝ(x, t) · γ̂(x) 1{ŝ(x) ≤ t̂} − γ(x) 1{ŝ(x) ≤ τ̂0∗ } which is a function adapted to An and obeys r̂(x, t) ∈ [−1, 1] for any value of (x, t), since ĝ(x, t) ∈ [0, 1] and γ̂(x), γ(x) ∈ [0, 1]. Then, for an independent copy (X, T ) ∼ PX,T , Rn ≤ α + ≤α+

n √ 1X † ω̂ (Xi )r̂(Xi , Ti ) − E[w(X)r̂(X, T ) | An ] + OP (1/ n) n i=1 n n o √ 1 X † 1 Xn ω̂ (Xi ) − w(Xi ) r̂(Xi , Ti ) + w(Xi )r̂(Xi , Ti ) − E[w(X)r̂(X, T ) | An ] +OP (1/ n). n i=1 n i=1 | {z } | {z } (a)

(b)

We now proceed to control the two terms separately via uniform convergence. Let Tn denote the σ-field generated by the training processes, conditional on which ĝ, γ̂, and ŝ are fixed. Writing Pn for the empirical measure of (Xi , Ti )ni=1 and P for the expectation over an independent copy, define f1,u (x, t) := w(x)ĝ(x, t)γ̂(x) 1{ŝ(x) ≤ u},

f2,v (x, t) := w(x)ĝ(x, t)γ(x) 1{ŝ(x) ≤ v}. 43

By definition, as the cutoffs τ̂0∗ and t̂ in the definition of r̂(·, ·) is fixed given An , we know √ |(b)| = (Pn − P )f1,t̂ − (Pn − P )f2,τ̂0∗ ≤ sup |(Pn − P )f1,u | + sup |(Pn − P )f2,v | = OP (1/ n), u∈R

v∈R

where the last equality follows from applying Lemma F.3 twice, provided that ∥wĝγ̂∥L2 (PX,T ) +∥wĝγ∥L2 (PX,T ) = OP (1) based on the boundedness of w(·), ĝ, γ, and γ̂. For term (a), we define for any u, v ∈ R the function ru,v (x, t) := ĝ(x, t) {γ̂(x) 1{ŝ(x) ≤ u} − γ(x) 1{ŝ(x) ≤ v}} , so that r̂ = rt̂,τ̂0∗ . We then have |(a)| = Pn {(ω̂ † − w)rt̂,τ̂0∗ } ≤ P {(ω̂ † − w)rt̂,τ̂0∗ } + sup (Pn − P ){(ω̂ † − w)ru,v } . u,v∈R

By the Hölder’s inequality, P {(ω̂ † − w)rt̂,τ̂0∗ } ≤ ∥ω̂ † − w∥L∞ (PX ) ∥r̂∥L1 (PX,T ) . Moreover, sup (Pn − P ){(ω̂ † − w)ru,v } u,v∈R

    ≤ sup (Pn − P ) (ω̂ † − w)ĝγ̂ 1{ŝ ≤ u} + sup (Pn − P ) (ω̂ † − w)ĝγ 1{ŝ ≤ v} . v∈R

u∈R

Conditional on Tn , both terms are indexed by nested threshold classes. Since ĝ, γ̂, γ ∈ [0, 1] and ∥ω̂ † − w∥L2 (PX ) = OP (1), applying Lemma F.3 twice gives supu,v∈R (Pn − P ){(ω̂ † − w)ru,v } = OP (n−1/2 ). Thus, |(a)| ≤ ∥ω̂ † − w∥L∞ (PX ) ∥r̂∥L1 (PX,T ) + OP (n−1/2 ). Combining the above two bounds gives Rn ≤ α + ∥ω̂ † − w∥L∞ (PX ) ∥r̂∥L1 (PX,T ) + OP (n−1/2 ).

(C.4)

We now proceed to show ∥r̂∥L1 (PX,T ) = OP (n−1/4 ). By definition, as ĝ ∈ [0, 1] and γ ∈ [0, 1], ∥r̂∥L1 (PX,T ) ≤ γ̂(·) 1{ŝ(·) ≤ t̂} − γ(·) 1{ŝ(·) ≤ τ̂0∗ } L1 (P ≤ γ̂(·) 1{ŝ(·) ≤ t̂} − γ(·) 1{ŝ(·) ≤ t̂} L1 (P ≤ ∥γ̂ − γ∥L1 (PX,T ) +

X,T )

X,T )

+ γ(·) 1{ŝ(·) ≤ t̂} − γ(·) 1{ŝ(·) ≤ τ̂0∗ } L1 (P

X,T )

1{ŝ(·) ≤ t̂} − 1{ŝ(·) ≤ τ̂0∗ } L1 (P ) ≤ OP (n−1/4 ) + O(|τ̂0∗ − t̂|), X,T

where the last inequality uses the bounded-density condition on ŝ(·). In the following, we show that |τ̂0∗ − t̂| = OP (n−1/4 ) under the given conditions. Recall that Hn (t) := E[γ̂(X) 1{ŝ(X) ≤ t} | Tn ] and t◦n := sup{t ∈ R : Hn (t) ≤ α}. By Lemma F.4, n

1X γ̂(Xi ) 1{ŝ(Xi ) ≤ t} − Hn (t) = OP (n−1/2 ). t∈R n i=1

sup

Together with condition (ii) of Theorem A.9, the cutoff conclusion of Lemma F.4 gives |t̂ − t◦n | = OP (n−1/2 ). Since Assumption A.6 holds with rn = O(n−1/2 ), Lemmas C.1 and F.4, together with P(Yi∗∗ = 0 | Xi , Ti ) = γ(Xi ), give   E ω̂ † (X)ĝ(X, T )γ(X) 1{ŝ(X) ≤ t} Tn ∗ ◦ −1/2 ◦ . sup |F̂n,0 (t) − Fn (t)| = OP (n ), where Fn (t) := E[ω̂ † (X)ĝ(X, T ) | Tn ] t∈R 44

  Write Dn := E ω̂ † (X)ĝ(X, T ) | Tn . For every integrable function h, we have E[w(X)ĝ(X, T )h(X) | Tn ] = E[h(X) | Tn ]. It follows that |Dn − 1| ≤ ∥ω̂ † − w∥L∞ (PX ) = OP (n−1/4 ), and   sup E ω̂ † (X)ĝ(X, T )γ(X) 1{ŝ(X) ≤ t} Tn − Hn (t) ≤ ∥ω̂ † − w∥L∞ (PX ) + ∥γ̂ − γ∥L1 (PX ) = OP (n−1/4 ). t∈R

Since Dn = 1 + oP (1) and 0 ≤ Hn (t) ≤ 1, on the event Dn ≥ 1/2, it holds that supt∈R |Fn◦ (t) − Hn (t)| = OP (n−1/4 ). Therefore, by the triangle inequality, ∗ ∗ εn := sup |F̂n,0 (t) − Hn (t)| ≤ sup |F̂n,0 (t) − Fn◦ (t)| + sup |Fn◦ (t) − Hn (t)| = OP (n−1/4 ). t∈R

t∈R

t∈R

Let En denote the event on which Hn (t◦n − u) ≤ α − cu and Hn (t◦n + u) ≥ α + cu for every 0 < u ≤ η. By condition (ii) of Theorem A.9, we know P(En ) → 1. Set ε̄n := εn +n−1 and un := 2ε̄cn . Since εn = OP (n−1/4 ), we have un = OP (n−1/4 ) and P{En ∩ {un ≤ η}} → 1. On this event, ∗ F̂n,0 (t◦n − un ) ≤ Hn (t◦n − un ) + εn ≤ α − cun + εn < α, ∗ F̂n,0 (t◦n + un ) ≥ Hn (t◦n + un ) − εn ≥ α + cun − εn > α. ∗ Because F̂n,0 is non-decreasing, it follows that t◦n − un ≤ τ̂0∗ ≤ t◦n + un . Therefore, we have |τ̂0∗ − t◦n | = −1/4 OP (n ), and hence |τ̂0∗ − t̂| ≤ |τ̂0∗ − t◦n | + |t̂ − t◦n | = OP (n−1/4 ).

Putting it back to (C.4), we conclude the proof of Theorem A.9.

C.3

Proof of Lemma A.11

P Proof of Lemma A.11. Let N := {t ∈ R : |t − t◦ | ≤ η} and write Pn,G f := m−1 n i∈Icalib f (Xi ) and PG f := E[f (X) | G = 1, Tn ]. We work conditional on Tn and on the event in Assumption A.10, whose probability tends to one. Conditional on Tn , the pairs (Xi , Gi ) are i.i.d., with P(Gi = 1 | Xi = x, Tn ) = q(x) and P(Gi = 1 | Tn ) = pG . All the conditional stochastic bounds below therefore also hold marginally. For brevity, throughout the first part of the proof write ψt := ψtγ̂,w̃ , M (t) := Mγ̂,w̃ (t), b(t) := bγ̂,w̃ (t), and ωt := ωtγ̂,w̃ . The coordinates of ψt and their pairwise products are uniformly bounded and indexed by the nested sets {ŝ(X) ≤ t}. Applying Lemma F.4 coordinatewise gives sup ∥ψ̄n (t) − b(t)∥ = OP (n−1/2 ), t∈N

mn = pG + OP (n−1/2 ), n

(C.5)

sup ∥M̂n (t) − M (t)∥op = OP (n−1/2 ),

sup |Pn,G 1{ŝ(X) ≤ u} − PG 1{ŝ(X) ≤ u}| = OP (n−1/2 ).

t∈N

u∈R

Since pG and inf t∈N λmin {M (t)} are bounded away from zero, M̂n (t) is invertible uniformly over t ∈ N with probability tending to one. Pn Define Hn (t) := n−1 i=1 γ̂(Xi ) 1{ŝ(Xi ) ≤ t} and H(t) := E[γ̂(X) 1{ŝ(X) ≤ t} | Tn ]. Lemma F.4 gives supt |Hn (t) − H(t)| = OP (n−1/2 ). The local-slope condition in Assumption A.10 therefore yields |t̂ − t◦ | = OP (n−1/2 ).

(C.6)

In particular, t̂ ∈ N with probability tending to one. We next characterize the optimizer. Fix t ∈ N and let Ψt be the mn × 3 matrix whose row indexed by i ∈ Icalib is ψt (Xi )⊤ . For any b ∈ Bn (t), the unique vector u ∈ Rmn minimizing ∥u∥2 subject to ⊤ −1 m−1 b, with objective value mn b⊤ M̂n (t)−1 b. Hence the unconstrained solution to n Ψt u = b is u = Ψt M̂n (t) the balancing program is ŵi (t) = ω̂t (Xi ). By the definition of Bn (t), equation (C.5), and δn = O(n−1/2 ), we know supt∈N ∥b̂n (t) − b(t)∥ = OP (n−1/2 ). The identity M̂n−1 − M −1 = M̂n−1 (M − M̂n )M −1 then gives sup ∥ân (t) − M (t)−1 b(t)∥ = OP (n−1/2 ),

sup

t∈N

t∈N ,x∈X

45

|ω̂t (x) − ωt (x)| = OP (n−1/2 ).

(C.7)

By Assumption A.10(ii), we have inf t∈N ,x∈X ωt (x) ≥ c1 . Thus, with probability tending to one, ω̂t (x) ≥ 0 uniformly over t ∈ N and x ∈ X . The unconstrained solution is therefore feasible for the nonnegative balancing program and, by strict convexity, is its unique solution. Taking t = t̂ proves ŵi = ω̂t̂ (Xi ). We now compare the empirical weights with Wn (γ̂, w̃) = ωt◦ . Since q(x) ≤ 1 and pG is bounded away from zero, Assumption A.10(iv) implies PG {a < ŝ(X) ≤ b} ≤ C(b − a) for all a < b. Consequently, uniformly over t, t′ ∈ N , ∥b(t) − b(t′ )∥ + ∥M (t) − M (t′ )∥op ≤ C|t − t′ |,

∥M (t)−1 b(t) − M (t′ )−1 b(t′ )∥ ≤ C|t − t′ |.

Write M (t)−1 b(t) = (a0 (t), ah (t), aw (t))⊤ . Then ωt (x) − ωt′ (x) = ψt (x)⊤ {M (t)−1 b(t) − M (t′ )−1 b(t′ )} + ah (t′ ){hγ̂t (x) − hγ̂t′ (x)}, and hence |ωt (x) − ωt′ (x)|2 ≤ C|t − t′ |2 + C 1{min(t, t′ ) < ŝ(x) ≤ max(t, t′ )}.

(C.8)

Equations (C.5) and (C.6) imply Pn,G 1{min(t̂, t◦ ) < ŝ(X) ≤ max(t̂, t◦ )} = OP (n−1/2 ). It follows from (C.8) that Pn,G {ωt̂ (X) − ωt◦ (X)}2 = OP (n−1/2 ).

(C.9)

On the other hand, equation (C.7) gives Pn,G {ω̂t̂ (X) − ωt̂ (X)}2 = {ân (t̂) − M (t̂)−1 b(t̂)}⊤ M̂n (t̂){ân (t̂) − M (t̂)−1 b(t̂)} = OP (n−1 ).

(C.10)

Combining (C.9)–(C.10) proves 1 X 2 {ŵi − Wn (γ̂, w̃)(Xi )} = OP (n−1/2 ). mn i∈Icalib

The boundedness of ψt , together with Assumption A.10(ii), implies |ωt (x)| ≤

sup t∈N ,x∈X

sup

∥ψt (x)∥ sup ∥M (t)−1 ∥op sup ∥b(t)∥ ≤ C.

t∈N ,x∈X

t∈N

t∈N

Together with (C.7) and P(t̂ ∈ N ) → 1, this gives supx∈X |ω̂t̂ (x)| = OP (1). Since ŵi = ω̂t̂ (Xi ) for i ∈ Icalib and ŵn+1 = ω̂t̂ (Xn+1 ), it follows that ŵn+1 ∨ maxi∈Icalib ŵi = OP (1). The normalization constraint gives P −1/2 ) and pG is bounded away from zero. Therefore, i∈Icalib ŵi = mn , while mn /n = pG + OP (n ŵn+1 ∨ maxi∈Icalib ŵi P = OP (n−1 ). ŵ i i∈Icalib Suppose now that ∥w̃ − w∥L2 (PX ) = oP (1). Let a◦ := (0, 0, pG )⊤ . Since q(x)w(x) = 1, we have, uniformly over t ∈ N , that b(t) = E[q(X)ψt (X)w(X) | Tn ]. On the other hand, since ψt (x)⊤ a◦ = pG w̃(x), we have M (t)a◦ = E[q(X)ψt (X)w̃(X) | Tn ]. Consequently, we have b(t) − M (t)a◦ = E[q(X)ψt (X){w(X) − w̃(X)} | Tn ]. By Assumption A.10, ψt is uniformly bounded and M (t)−1 is uniformly bounded over t ∈ N . Hence, by the Cauchy–Schwarz inequality, sup ∥M (t)−1 b(t) − a◦ ∥ ≤ C∥w̃ − w∥L2 (PX ) = oP (1).

t∈N

46

Moreover, since w◦ (x) = pG w(x), we have ωt (x) − w◦ (x) = ψt (x)⊤ {M (t)−1 b(t) − a◦ } + pG {w̃(x) − w(x)}. It follows that supt∈N ∥ωt − w◦ ∥L2 (PX ) = oP (1). Taking t = t◦ and recalling that Wn (γ̂, w̃) = ωt◦ gives ∥Wn (γ̂, w̃) − w◦ ∥L2 (PX ) = oP (1). Finally, suppose εn := ∥w̃ − w∥L∞ (PX ) = OP (n−1/4 ). The preceding argument now yields sup ∥M (t)−1 b(t) − a◦ ∥ ≤ Cεn ,

sup |ah (t)| ≤ Cεn ,

t∈N

t∈N

where ah (t) denotes the coefficient on hγ̂t in M (t)−1 b(t). Therefore, supt∈N ∥ωt − w◦ ∥L∞ (PX ) = OP (n−1/4 ), and in particular, ∥Wn (γ̂, w̃) − w◦ ∥L∞ (PX ) = OP (n−1/4 ). To sharpen the empirical approximation rate, recall that, uniformly over t, t′ ∈ N ,   |ωt (x) − ωt′ (x)|2 ≤ C|t − t′ |2 + C sup |ah (u)|2 1{min(t, t′ ) < ŝ(x) ≤ max(t, t′ )}. u∈N

The bounds already established above imply |t̂ − t◦ | = OP (n−1/2 ) and Pn,G 1{min(t̂, t◦ ) < ŝ(X) ≤ max(t̂, t◦ )} = OP (n−1/2 ). Since supu∈N |ah (u)|2 = OP (n−1/2 ), it follows that Pn,G {ωt̂ (X) − ωt◦ (X)}2 = OP (n−1 ). Combining this with Pn,G {ω̂t̂ (X) − ωt̂ (X)}2 = OP (n−1 ) gives 1 X 2 {ŵi − Wn (γ̂, w̃)(Xi )} = OP (n−1 ), mn i∈Icalib

which completes the proof.

D

Technical proof for results in appendix

D.1

Proof of Theorem A.4

Proof of Theorem A.4. Following Kallus (2022) and Li et al. (2023), for any decision rule ϕ : X → {0, 1}, max Err(ϕ; P ) = E[ϕ(X)γ(X)], P ∈P

where γ(x) = min{P(Y (1) = 0 | X = x), P(Y (0) = 1 | X = x)} solely relies on the observed distribution, and E[·] is with respect to the observed distribution. On the other hand, the power is the expectation of ϕ(X) under the observed covariate distribution. The optimization problem (with randomized policy) is maximize E[ϕ(X)] ϕ : X →[0,1]

subject to E[ϕ(X)γ(X)] ≤ α. If E[γ(X)] ≤ α, then ϕ∗power = 1 maximizes the power since its power is one. Otherwise, suppose E[γ(X)] > α. Consider any decision rule ψ(·) obeying E[ψ(X)γ(X)] ≤ α, and write ϕ∗ (·) = ϕ∗power (·) for notational convenience. Note that  γ ∗ · E[ψ(X)] − E[ϕ∗ (X)]  ∗  = E γ · (ψ(X) − ϕ∗ (X)) 1{γ(X) < γ ∗ } 47

    + E γ ∗ · (ψ(X) − ϕ∗ (X)) 1{γ(X) > γ ∗ } + E γ ∗ · (ψ(X) − ϕ∗ (X)) 1{γ(X) = γ ∗ } . The first term, as ϕ∗ (X) = 1 when γ(X) < γ ∗ and hence ψ(X) − ϕ∗ (X) ≤ 0, obeys     E γ ∗ · (ψ(X) − ϕ∗ (X)) 1{γ(X) < γ ∗ } ≤ E γ(X)(ψ(X) − ϕ∗ (X)) 1{γ(X) < γ ∗ } . Similar arguments yield     E γ ∗ · (ψ(X) − ϕ∗ (X)) 1{γ(X) > γ ∗ } ≤ E γ(X)(ψ(X) − ϕ∗ (X)) 1{γ(X) > γ ∗ } ,     E γ ∗ · (ψ(X) − ϕ∗ (X)) 1{γ(X) = γ ∗ } = E γ(X)(ψ(X) − ϕ∗ (X)) 1{γ(X) = γ ∗ } . Adding the three terms up, we know  γ ∗ · E[ψ(X)] − E[ϕ∗ (X)] ≤ E[γ(X)ψ(X)] − E[γ(X)ϕ∗ (X)] ≤ α − α = 0 since ψ(X) is a feasible solution and the satefy constraint is binding for ϕ∗ . Finally, as γ ∗ ≥ 0, this implies E[ψ(X)] ≤ E[ϕ∗ (X)]. The arbitrariness of ψ implies the optimality of ϕ∗ .

D.2

Proof of Theorem A.2

Proof of Theorem A.2. Let â(x) ∈ {0, 1} denote the arm selected by the inclusion rule, and define γ n (x) := {1 − µ1 (x)}1{â(x) = 1} + µ0 (x)1{â(x) = 0}. As in the proof of Theorem 4.2, the optimality of â(x) for the estimated proxy-label risks gives 0 ≤ γ n (x) − γ(x) ≤ 2 {|µ̂1 (x) − µ1 (x)| + |µ̂0 (x) − µ0 (x)|} . Hence ∥γ n − γ∥L1 (PX ) = oP (1). The assumed outcome-model consistency also implies ∥γ̂ − γ∥L1 (PX ) = oP (1). Conditional on the training process, let wn (x) := ρn (x)−1 .   The weighted identities in the proof of Theorem 4.2 give E[Gwn (X) | X] = 1 and E Gwn (X)1{Y † = 0} X = γ n (X). It follows by the same uniform-law-of-large-numbers argument that the weighted empirical conformal curve converges uniformly to H(t) := E[γ(X)1{γ(X) ≤ t}]. Consequently, ρn (x) := e(x)1{â(x) = 1} + {1 − e(x)}1{â(x) = 0},

pstr−power − H{γ(Xn+1 )} = oP (1). n+1 Since γ(X) has no point mass, the same threshold argument as in the proof of Theorem 4.2 gives  ∗ P π̂str−power (Xn+1 ) ̸= πpower (Xn+1 ) → 0, where the binding and nonbinding cases are characterized by Theorem A.4. Therefore, |E[π̂str−power (Xn+1 )]− ∗ Power(πpower ; P )| → 0, which completes the proof.

D.3

Proof of Theorem A.3

Proof of Theorem A.3. Define Hn (t) := E[γ̂(X) 1{γ̂(X) ≤ t} | Tn ] ,

t◦n := sup{t ∈ R : Hn (t) ≤ α},

and H0 (t) := E[γ(X) 1{γ(X) ≤ t}]. By ∥γ̂ − γ∥L2 (PX ) = oP (1), the no-point-mass condition on γ(X), and Lemma F.4, we have sup |Hn (t) − H0 (t)| = oP (1). t∈R

48

Together with the local-crossing condition in Assumption A.10(iii), this implies t◦n − t∗power = oP (1), where t∗power := sup{t ∈ R : H0 (t) ≤ α}. Set ω̂ := Wn (γ̂, w̃) and define Fnγ̂ (t) :=

E[ω̂(X)ĝ(X, T )γ̂(X) 1{γ̂(X) ≤ t} | Tn ] . E[ω̂(X)ĝ(X, T ) | Tn ]

By the defining population balance equations, we know Fnγ̂ (t◦n ) = Hn (t◦n ) = α. Moreover, Assumption A.10(i)–(iii) implies that Fnγ̂ crosses α at t◦n with a slope bounded away from zero, with probability tending to one. Define F̂nobs (t) :=

ŵn+1 +

† = 0} 1{γ̂(Xi ) ≤ t} i=1 Gi ŵi 1{Y Pni , ŵn+1 + i=1 Gi ŵi

Pn

−1 obs P ŵn+1 ), so that pobs n+1 = F̂n {γ̂(Xn+1 )}. By Lemma A.11, Lemmas C.1 and F.4, and ŵi = OP (n i∈Icalib

we obtain the uniform convergence of the empirical curve F̂nobs (t) to its training-conditional population counterpart. Furthermore, by the definition of g ∗ , we know E[g ∗ (X, T ) 1{Y † = 0} | X]   = E E[1{Y † = 0} | X]T 1{1 − µ1 (X) ≤ µ0 (X)} X] + E E[1{Y † = 0} | X](1 − T ) 1{1 − µ1 (X) > µ0 (X)} X]   = E (1 − µ1 (X))T 1{1 − µ1 (X) ≤ µ0 (X)} X] + E µ0 (X)(1 − T ) 1{1 − µ1 (X) > µ0 (X)} X] = E[g ∗ (X, T )γ(X) | X] Therefore, sup |F̂nobs (t) − Fnγ̂ (t)| ≤ oP (1) + C∥ĝ − g ∗ ∥L1 (PX,T ) + C∥γ̂ − γ∥L1 (PX ) = oP (1). t∈R

Fix 0 < ε ≤ η. The local-crossing and uniform convergence results imply that, with probability tending to one, γ̂(Xn+1 ) ≤ t◦n − ε

⇒

pobs n+1 ≤ α,

whereas

γ̂(Xn+1 ) ≥ t◦n + ε

⇒

pobs n+1 > α.

Consequently, |π̂obs (Xn+1 ) − 1{γ̂(Xn+1 ) ≤ t◦n }| ≤ 1{|γ̂(Xn+1 ) − t◦n | ≤ ε} + oP (1). Assumption A.10(iv) and the arbitrariness of ε > 0 yield E [|π̂obs (Xn+1 ) − 1{γ̂(Xn+1 ) ≤ t◦n }|] → 0. Finally, since γ̂ → γ in L2 (PX ), t◦n → t∗power , and γ(X) has no point mass,   E π̂obs (Xn+1 ) − 1{γ(Xn+1 ) ≤ t∗power } → 0. ∗ By Theorem A.1, we know πpower (x) = 1{γ(x) ≤ t∗power }. Hence   ∗ ∗ E[π̂obs (Xn+1 )] − E[πpower (Xn+1 )] ≤ E π̂obs (Xn+1 ) − πpower (Xn+1 ) → 0, ∗ which proves E[π̂obs (Xn+1 )] → Power(πpower ; P).

D.4

Proof of Theorem 3.3

Proof of Theorem 3.3. The result follows directly from Theorem A.5. In particular, since swelfare (X) has no point mass, P{swelfare (X) = r∗ } = 0, so the randomization at the cutoff in Theorem A.5 is immaterial. ∗ Therefore, the optimal randomized rule in Theorem A.5 reduces almost surely to πwelfare (x) = 1{swelfare (x) ≤ r∗ }. The binding/nonbinding characterizations follow from the two cases in Theorem A.5.

49

Proof of Theorem A.5. We first characterize the worst-case safety constraint. For any P ′ ∈ P, define fP ′ (x) := P ′ {Y (1) = 0, Y (0) = 1 | X = x}. For a randomized policy ϕ, whose randomization is independent of the potential outcomes conditional on X, its harm rate under P ′ is Err(ϕ; P ′ ) = E[ϕ(X)fP ′ (X)]. Since fP ′ (x) ≤ P ′ {Y (1) = 0 | X = x} = 1 − µ1 (x) and fP ′ (x) ≤ P ′ {Y (0) = 1 | X = x} = µ0 (x), we have fP ′ (x) ≤ γ(x) for every P ′ ∈ P. This upper bound is sharp. Indeed, for every x, consider the following conditional distribution of the potential outcomes: P ′ {Y (1) = 0, Y (0) = 1 | X = x} = γ(x), P ′ {Y (1) = 0, Y (0) = 0 | X = x} = 1 − µ1 (x) − γ(x), P ′ {Y (1) = 1, Y (0) = 1 | X = x} = µ0 (x) − γ(x), P ′ {Y (1) = 1, Y (0) = 0 | X = x} = µ1 (x) − µ0 (x) + γ(x). All four probabilities are nonnegative. The first three nonnegativity claims follow directly from γ(x) = min{1 − µ1 (x), µ0 (x)}, while the last follows from γ(x) ≥ µ0 (x) − µ1 (x). They sum to one and yield the conditional marginals P ′ {Y (1) = 1 | X = x} = µ1 (x) and P ′ {Y (0) = 1 | X = x} = µ0 (x). Combining this conditional coupling with the original distribution of X and the same treatment-assignment mechanism therefore defines a distribution in P. Consequently, for every randomized policy ϕ, max Err(ϕ; P ′ ) = E[γ(X)ϕ(X)].

P ′ ∈P

Write τ (x) := µ1 (x) − µ0 (x). The welfare of a randomized policy is Welfare(ϕ; P ) = E[µ1 (X)ϕ(X) + µ0 (X){1 − ϕ(X)}] = E[µ0 (X)] + E[τ (X)ϕ(X)]. Thus, up to the constant E[µ0 (X)], the randomized extension of (3.9) is equivalent to maximize ϕ:X →[0,1]

E[τ (X)ϕ(X)]

subject to

E[γ(X)ϕ(X)] ≤ α.

We next verify that the cutoff and randomization probability in the theorem are well defined. Recall that τ (x) under the stated ratio conventions, and define, for r ≤ 0, swelfare (x) = − γ(x) H(r) := E[γ(X) 1{swelfare (X) < r}] . The function H is nondecreasing and left-continuous on (−∞, 0]. Indeed, if rk ↑ r, then 1{swelfare (X) < rk } increases pointwise to 1{swelfare (X) < r}, so the conclusion follows from the dominated convergence theorem. Moreover, H(r) → 0 as r → −∞. To see this, if swelfare (X) is finite, then 1{swelfare (X) < r} → 0. If swelfare (X) = −∞, then necessarily γ(X) = 0, so γ(X) 1{swelfare (X) < r} = 0 for every r. The dominated convergence theorem therefore gives the claim, and hence the set defining r∗ is nonempty. Suppose first that H(0) > α. Since H(r) ↑ H(0) as r ↑ 0, there exists some r0 < 0 such that H(r0 ) > α. It follows that r∗ < 0. By the definition of r∗ , there exists a sequence rk ↑ r∗ such that H(rk ) ≤ α. The left-continuity of H therefore implies H(r∗ ) ≤ α. On the other hand, the right limit of H at r∗ is lim H(r) = E[γ(X) 1{swelfare (X) ≤ r∗ }] = H(r∗ ) + E[γ(X) 1{swelfare (X) = r∗ }] .

r↓r ∗

This limit must be at least α. Otherwise, there would exist some r > r∗ sufficiently close to r∗ such that H(r) < α, contradicting the definition of r∗ as the supremum. Therefore, H(r∗ ) ≤ α ≤ H(r∗ ) + E[γ(X) 1{swelfare (X) = r∗ }] . This establishes the existence of η ∗ ∈ [0, 1] satisfying the equality in the theorem. If the expectation multiplying η ∗ is zero, the preceding inequalities imply H(r∗ ) = α, and we may take η ∗ = 0. If instead 50

H(0) ≤ α, the theorem sets r∗ = 0 and η ∗ = 0. Under the stated ratio conventions, swelfare (x) < 0 if and only if τ (x) > 0. Hence, in this case, ϕ∗welfare (x) = 1{τ (x) > 0},

E[γ(X)ϕ∗welfare (X)] = H(0) ≤ α.

It follows in both cases that ϕ∗welfare is feasible and satisfies (−r∗ ) {E[γ(X)ϕ∗welfare (X)] − α} = 0 by complementary slackness. It remains to establish optimality. Let ψ : X → [0, 1] be any feasible randomized policy. By the definition of ϕ∗welfare , it holds pointwise that {τ (x) + r∗ γ(x)}{ψ(x) − ϕ∗welfare (x)} ≤ 0. Indeed, if swelfare (x) < r∗ , then τ (x) + r∗ γ(x) ≥ 0 and ϕ∗welfare (x) = 1, so the second factor is nonpositive. If swelfare (x) > r∗ , then τ (x) + r∗ γ(x) ≤ 0 and ϕ∗welfare (x) = 0, so the second factor is nonnegative. If swelfare (x) = r∗ , the first factor is zero. The same conclusions continue to hold when γ(x) = 0 under the stated ratio conventions. Consequently, Welfare(ψ; P ) − Welfare(ϕ∗welfare ; P ) = E[τ (X){ψ(X) − ϕ∗welfare (X)}] = E[{τ (X) + r∗ γ(X)}{ψ(X) − ϕ∗welfare (X)}] − r∗ E[γ(X){ψ(X) − ϕ∗welfare (X)}] ≤ −r∗ {E[γ(X)ψ(X)] − E[γ(X)ϕ∗welfare (X)]} ≤ −r∗ {α − E[γ(X)ϕ∗welfare (X)]} = 0. The first inequality follows from the preceding pointwise inequality. The second follows from the feasibility of ψ and the fact that −r∗ ≥ 0. The final equality follows from complementary slackness. Since ψ was arbitrary, this proves the optimality of ϕ∗welfare .

E

Simulation details

E.1

Details for experiments in Figure 2

This section includes the omitted details that produce Figure 2. Data generating process. Across all the four settings, we set the feature dimension as X ∈ R20 . In Settings 1-2, the linear coefficients are obtained by a random i.i.d. draw from Unif[(1, 2)] and fixed before running all the experiments. i.i.d.

In Setting 1, we draw X ∼ N (0, I) and set µ1 (x) = logit−1 (0.0835x1 + 0.097x2 + 0.0885x3 + 0.0575x4 + 0.0675x5 + 1) and µ0 (x) = 1 − 0.8µ1 (x) − 0.2 logit−1 (−0.87x4 − 0.58x5 − 0.675x6 − 0.54x7 − 0.57x8 ). Both functions are then truncated to [0.05, 0.95]. i.i.d.

In Setting 2, we draw X ∼ N (0, Σ) with Σ = 0.5 1 1⊤ +0.5I and set µ1 (x) = logit−1 (x⊤ θ1 + 0.8) where [θ1 ]1:10 = (1.66, 1.71, 1.93, 1.72, 1.99, 1.98, 1.52, 1.57, 1.37, 1.53) × 0.02, and µ0 (x) = 1 − 0.9µ1 (x) − 0.4 logit−1 (x⊤ θ0 ) where [θ0 ]4:13 = −(1.73, 1.92, 1.97, 1.70, 1.78, 1.37, 1.78, 1.76, 1.59, 1.32)×0.2; both functions are then truncated to [0.05, 0.95]. Finally, we take X1:7 as the observed features. In the nonlinear Setting 3, we draw entries of X i.i.d. from Unif(0, 1) and set µ1 (x) = logit−1 (f1 (x)), where f1 (x) = signal · {sin(3πx1 ) + cos(2πx2 ) + x23 − 1(x4 > 0.5)x24 + 0.5 1(x5 > 0.3) + e−x4 } + 1; we then truncate µ1 (x) to [0.05, 0.95]. Next, we define µ0 (x) = 1 − 0.2µ1 (x) − 0.8µ1 (x)2 . In Setting 4, we draw X from the same correlated Gaussian design as in Setting 2, and set µ1 (x) = logit−1 (f1 (x)) with f1 (x) = 0.2 · {sin(3πx1 ) + cos(2πx8 ) + 0.5x23 − 1(x4 > 0.5)x25 + 0.5 1(x6 > 0.3) + e−x7 } + 0.8 + 0.1x8 , truncated to [0.05, 0.95]. We then define µ0 (x) = 1 − µ1 (x) + 0.1 1(x6 > 0.5) − 0.2x28 , truncated to [0.08, 0.95]. Finally, after generating the outcomes and treatments, we take X1:5 as the observed features.

51

Additional implementation details. For Li et al., we implement the doubly-robust estimator ûFNA (π) in Li et al. (2023, Lemma 6.1) with uniform cost c(X) = 1, where the µt (·) are fitted using the same random forest classifier in scikit-learn Python library with 2-fold cross-fitting. The policy class π ∈ Π is depth-2 policy trees from the econml Python library. Given these estimators we fit the optimization problem (2) in Li et al. (2023) via the dual form and a bi-search to find the dual variable. The Policy-Tree baseline follows the standard use of econml library, where we use Dlabel to learn the policy and apply to Dtest .

E.2

Details for additional experiments in Section 6.1 i.i.d.

For all the additional experiments in Section 6.1, we sample X ∈ [0, 1]20 ∼ Unif([0, 1]20 ) and assign treatment independently as T ∼ Bernoulli(1/2). We state the three sets of DGPs below. Arm informativeness and selective calibration. In this set of experiments, we vary the scale of µ1 (x) and µ0 (x) to vary whether the harm rate relies on the treated or control outcome, i.e., whether ρ(x) = 1 − µ1 (x) or ρ(x) = µ0 (x). Let q ∈ {0, 0.25, 0.5, 0.75, 1} be a parameter that controls the comparison of the two outcome models. We define the latent score rarm (x) = sin(πx1 x2 ) + x3 − 0.5x4 + 0.3 1{x5 > 0.4}, and qr ∈ R be the q-th population quantile of rarm (X). For rarm (x) ≤ qr , we set µ1 (x) = 1 − ρ(x),

µ0 (x) = ρ(x) + ∆(x),

whereas for rarm (x) > qr , we set µ0 (x) = ρ(x),

µ1 (x) = 1 − ρ(x) − ∆(x).

Thus, the parameter q controls the fraction of units for which the treated arm is the informative arm. Here we use a nonlinear harm rate function ρ(x) = 0.08 + 0.22 expit{sin(2πx1 ) + 0.7 cos(2πx2 ) + 0.5(x3 − 0.5)2 − 0.6 1{x4 > 0.6} + 0.5e−x5 }, and the treatment effect function ∆(x) = {0.35 + 0.35 + 0.25 expit(rδ (x))} · max{1 − 2ρ(x) − 0.02, 0.02}, where rδ (x) = cos(2πx3 ) − 0.6 sin(2πx5 ) + 0.8x6 − 0.4 1{x7 > 0.5}, and expit(u) = 1/(1 + e−u ). Subgroup heterogeneity. x̃2 = x2 − 0.5, and define

For this experiment we use a three-group construction. Let x̃1 = x1 − 0.5 and v(x) = σ(sin(2πx3 ) − 0.8 cos(2πx4 ) + 0.6x5 ) .

We partition the covariate space into three groups: the first group G1 consists of samples with x̃1 > 0; the second group G2 consists of samples with x̃1 ≤ 0 and x̃2 ≤ 0; the third group G3 consists of samples with x̃1 ≤ 0 and x̃2 > 0. For x ∈ G1 , we set µ0 (x) = 0.35,

µ1 (x) = 0.65.

For x ∈ G2 ∪ G3 , we define a small heterogeneous treatment effect τ (x) = max{0.004, min{0.03, 0.008 + 0.01125 0.3 + 0.7v(x) }. Then for x ∈ G2 , we set µ1 (x) = max{0.94, min{0.995, 0.95 + 0.02v(x)}},

µ0 (x) = max{0.9, min{0.99, µ1 (x) − τ (x)}},

so both outcomes are near one and the treatment effect is small and positive. Note that in G2 , the harm rate is small, and one recognizes this only when using the treated outcomes for calibration. For x ∈ G3 , we set µ1 (x) = max{0.001, min{0.06, µ0 (x) − τ (x)}},

µ0 (x) = max{0.005, min{0.08, 0.02 + 0.03v(x)}}

so both arms are near zero and the treatment effect is small and negative. In G3 , the harm rate is close to zero, and the method recognizes this only when using the control outcomes for calibration.

52

E.3

Details for experiments in Section 6.2

Data generating process for Figure 5. In the four settings, the process that generates (X, Y (1), Y (0)) is the same as those in Figure 2 for RCT experiments. Given the features {Xi }, we sample the treatments {Ti } independently from Bernoulli(e(Xi )) according to a propensity score e(x) = P(T = 1 | X = x) = logit−1 (ηe (x)). In the linear settings 1-2, we set ηe (x) = 0.25 · (1.1x1 + 0.9x2 + 0.6x3 − 0.4x4 + 0.5x5 + 0.3x6 ). In the nonlinears 3-4, we set ηe (x) = 0.25 · (sin(3πx1 ) + cos(2πx8 ) + 0.1x23 − 0.3x4 ). Data generating process for double robustness results. We generate covariates X ∈ R20 with entries i.i.d. from Unif(0, 1). We define three latent threshold indicators by q1 (x) = 1{x1 > 0.5}, q2 (x) = 1{x3 > 0.5}, and q3 (x) = 1{x5 > 0.5}. These induce three hidden subgroups: G1 (x) = 1{q1 (x) = 1, q2 (x) = 1}, G2 (x) = 1{q1 (x) = 1, q2 (x) = 0}, and G3 (x) = 1{q1 (x) = 0, q3 (x) = 1}. The potential outcomes satisfy Y (t) | X = x ∼ Bernoulli(µt (x)) for t ∈ {0, 1}. We set µ0 (x) = 0.07 + 0.02x7 + 0.01(x8 − 0.5) + 0.38 G1 (x) + 0.03 G3 (x) and τ (x) = µ1 (x) − µ0 (x) = 0.03 + 0.01(x8 − x7 ) − 0.48 G1 (x) + 0.42 G2 (x) + 0.14 G3 (x), and then truncate µ0 (x) to [0.03, 0.82] and set µ1 (x) = µ0 (x) + τ (x) truncated to [µ0 (x) + 0.01, 0.95]. Treatment is assigned observationally according to the propensity score e(x) = logit−1 (−1.2 + β G1 (x) + 1.2 G2 (x) + 0.6 G3 (x)), truncated to [0.1, 0.9], and we sample T | X = x ∼ Bernoulli(e(x)). Here the parameter β ∈ {2.8, 3.4, 4.0, 4.4} controls the strength of confounding in four levels shown in the x-axis of Figure 6. The observed outcome is Y = T Y (1) + (1 − T )Y (0). We generate (Y (1), Y (0)) under the negative coupling scheme. Methods for double robustness results. To study double robustness, we compare five nuisance-model regimes. In both correct, both the propensity and outcome models are fit by logistic regression on the transformed features (q1 , q2 , q3 , G1 , G2 , G3 , x7 , x8 ). In both wrong, both are fit by logistic regression on the raw covariates X. In ps correct only, the propensity model uses the transformed features while the outcome models use the raw covariates, and in outcome correct only the reverse is used. In the rf regime, both the propensity and outcome models are fit by random forests on the raw covariates. The other data splitting and model fitting details are the same as other experiments.

E.4

Real data analysis details

For the real-data analysis, we define the binary outcome as Yi = I(Post Belief Specifici < 50) and the treatment indicator as Ti = I(messageTypei = Active). Because the study is a randomized experiment, treatment assignment is known by design, so we do not estimate a propensity score model. We split the sample in two stages. First, we reserve 20% of observations as the test set. We then split the remaining 80% evenly into a 40% training set and a 40% calibration set. All prediction models are fit on the training set only, and predictions are generated for the calibration and test sets. The feature set includes demographic covariates (Education Cat, AgeYears, Race *, Gender *, religion), political and psychological covariates (Extremism, AOT, IH, Party *), AI-related covariates (genai fam 1, genai use 1, genai trust, Sureness 1), and baseline state variables (Pre Belief Specific, LowConfidence). Missing values in the structured covariates are imputed using mean imputation. These variables are then standardized using scaling parameters estimated on the training set. To incorporate the text information in the stated conspiracy, we use a pre-computed embedding for the (pre-treatment) conspiracy for each observation. We apply principal components analysis (PCA) to the training-sample embeddings and retain the first 20 principal components. The final feature vector is the concatenation of the standardized structured covariates, the standardized dialogue-length variables, and 20 PCA components. We estimate the conditional mean outcomes under treatment and control separately using the regression forest in econml Python package. Specifically, one regression forest is fit on the treated training subsample to estimate µ1 (x) = E[Y | X = x, T = 1], and a second regression forest is fit on the control training subsample to estimate µ0 (x) = E[Y | X = x, T = 0]. These two forests produce predictions µ̂1 (x) and µ̂0 (x) on the calibration and test sets. To obtain a direct estimate of treatment heterogeneity, we additionally fit a causal

53

forest via the econml package on the full training sample using the same features. This model produces a direct estimate of the conditional average treatment effect, τ̂CF (x), where τ (x) = µ1 (x) − µ0 (x). Our welfare-based selection rule combines the direct causal-forest estimate of treatment benefit with a conservative proxy for treatment risk. Define ρ̂(x) = min{1 − µ̂1 (x), µ̂0 (x)}. We then define the welfare score as Swelfare (x) = −τ̂CF (x)/ max{ρ̂(x), c}, where c > 0 is a small numerical floor to avoid instability when ρ̂(x) is close to zero. In our implementation, we set c = 0.025. Our power-based selection rule directly uses the ρ̂(x) above in the score.

F

Auxiliary lemmas

Lemma F.1. Suppose the distribution of a random variable X ∈ R has no point mass. Then supt P(t − δ < X ≤ t) as a function of δ converges to zero as δ → 0. Proof of Lemma F.1. Since X has no point mass, the c.d.f. F (t) := P(X ≤ t) is continuous and nondecreasing. Consider any constant ϵ > 0. There exists constants a, b ∈ R such that F (a) ≤ ϵ and 1−F (b) ≤ ϵ. Second, since F is continuous on the compact set [a, b], the Heine–Cantor theorem implies F is uniformly continuous on [a, b], which means there exists some δ > 0 such that supx,y∈[a,b],|x−y|≤δ |F (x) − F (y)| ≤ ϵ. Now we take any sequence a = t1 < t2 < · · · < tM = b, such that ti+1 − ti ≤ δ. By the arguments above, we know the following holds: P(ti < X ≤ ti+1 ) ≤ ϵ,

P(X ≤ t1 ) ≤ ϵ,

P(X > tM ) ≤ ϵ.

Therefore, for any t ∈ [t1 + δ, tM + δ], there exists some i such that δi ≤ t − δ < t ≤ δi+2 and thus P(t − δ < X ≤ t) ≤ 2ϵ. For any t ≤ t1 + δ, we know P(t − δ < X ≤ t) ≤ P(X ≤ t1 + δ) Therefore, for any ϵ > 0, we know supt P(X ≤ t < ŝ(X)) ≤ oP (1) + ϵ. The arbitrariness of ϵ > 0 the implies the desired result. Lemma F.2. Suppose two fixed functions s1 , s2 : X → R obeys ∥s1 (X) − s2 (X)∥L2 = o(1) and one of them has no point mass. Then for any fixed t ∈ R, we have P(s1 (X) ≤ t) − P(s2 (X) ≤ t) = o(1). Proof of Lemma F.2. Without loss of generality we assume s1 (X) has no point mass, so the mapping t 7→ P(s1 (X) ≤ t) is continuous on t ∈ R. The L2 convergence implies the convergence in probability. That is, P(|s1 (X)−s2 (X)| > ϵ) → 0 for any fixed ϵ > 0. Due to the continuity of P(s1 (X) ≤ t) in t ∈ R and the above convergence, for any δ > 0, we can find a sufficiently small ϵ > 0 such that P(s1 (X) ≤ t−ϵ) ≥ P(s1 (X) ≤ t)−δ, P(s1 (X) ≤ t + ϵ) ≤ P(s1 (X) ≤ t) + δ, and P(|s1 (X) − s2 (X)| > ϵ) ≤ δ. Therefore, P(s2 (X) ≤ t) ≤ P(s2 (X) ≤ t, |s1 (X) − s2 (X)| > ϵ) + P(s2 (X) ≤ t, |s1 (X) − s2 (X)| ≤ ϵ) ≤ P(|s1 (X) − s2 (X)| > ϵ) + P(s1 (X) ≤ t + ϵ) ≤ P(s1 (X) ≤ t) + 2δ. By the same arguments, P(s1 (X) ≤ t − ϵ) ≤ P(|s1 (X) − s2 (X)| > ϵ) + P(s2 (X) ≤ t), which further implies P(s2 (X) ≤ t) ≥ P(s1 (X) ≤ t) − 2δ. The arbitrariness of δ > 0 thus implies the desired result.

54

Lemma F.3. Let Tn be a σ-field representing a possibly random training process. Conditional on Tn , let {(Xi , Zi )}ni=1 be i.i.d. copies of (X, Z), where Z ∈ R, and let s : X → R be fixed. Suppose E[Z 2 | Tn ] = OP (1). Define n

Hn (t) :=

1X Zi 1{s(Xi ) ≤ t}, n i=1

H(t) := E[Z 1{s(X) ≤ t} | Tn ],

∆n := sup |Hn (t) − H(t)|. t∈R

Then ∆n = OP (n−1/2 ), where the stochastic order is with respect to both the training process and the evaluation sample. In addition, suppose that Z ≥ 0 almost surely conditional on Tn , and define   n 1X τ̂ := sup t ∈ R : Zi 1{s(Xi ) ≤ t} ≤ α , τ ∗ := sup {t ∈ R : E[Z 1{s(X) ≤ t} | Tn ] ≤ α} . n i=1 Assume that there exist constants c > 0 and δ > 0 such that, with probability tending to one over the training process, H(τ ∗ − u) ≤ α − cu, H(τ ∗ + u) ≥ α + cu for every 0 < u ≤ δ. Then τ̂ − τ ∗ = OP (n−1/2 ). Proof of Lemma F.3. For t ∈ R, set gt (x, z) := z 1{s(x) ≤ t}. Conditional on Tn , the functions gt are fixed and (Xi , Zi )ni=1 are i.i.d. By the conditional symmetrization inequality, " # n X 2 εi Zi 1{s(Xi ) ≤ t} Tn , E[∆n | Tn ] ≤ E sup n t∈R i=1 where ε1 , . . . , εn are i.i.d. Rademacher random variables independent of everything else. Condition further on (Xi , Zi )ni=1 , reorder the observations so that s(X(1) ) ≤ · · · ≤ s(X(n) ), and let aj := Z(j) . Since the sets {i : s(Xi ) ≤ t} are nested, sup

n X

εi Zi 1{s(Xi ) ≤ t} ≤ max

0≤k≤n

t∈R i=1

k X

εej aj ,

j=1

Pk where (e εj )nj=1 is again a Rademacher sequence. Setting Mk := j=1 εej aj , Doob’s L2 maximal inequality gives   n X 2 n Eε max |Mk | (Xi , Zi )i=1 , Tn ≤ 4 a2j . 0≤k≤n

j=1

Consequently,  E[∆n | Tn ] ≤

4  E n

n X

!1/2 Zi2

i=1

 4  1/2 Tn  ≤ √ E[Z 2 | Tn ] . n

Let Vn := {E[Z 2 | Tn ]}1/2 . For any K, M > 0, conditional Markov’s inequality yields √ 4K P( n ∆n > M ) ≤ P(Vn > K) + . M Since Vn = OP (1), this proves ∆n = OP (n−1/2 ). For the second claim, let En denote the event on which the local-slope condition holds, and define   cδ . An := En ∩ ∆n ≤ 4 Then P(An ) → 1. On An , let un := 4∆n /c, so that un ≤ δ. If ∆n = 0, then Hn ≡ H and hence τ̂ = τ ∗ . Otherwise, Hn (τ ∗ + un ) ≥ H(τ ∗ + un ) − ∆n ≥ α + cun − ∆n = α + 3∆n > α, 55

whereas Hn (τ ∗ − un ) ≤ H(τ ∗ − un ) + ∆n ≤ α − cun + ∆n = α − 3∆n < α. Because Zi ≥ 0, the function Hn is nondecreasing. It follows that τ ∗ − un ≤ τ̂ ≤ τ ∗ + un , and therefore |τ̂ − τ ∗ | ≤ 4∆n /c on An . Since ∆n = OP (n−1/2 ) and P(An ) → 1, we conclude that τ̂ − τ ∗ = OP (n−1/2 ). Lemma F.4. Consider a sequence of random functions fˆn : X × Z → R+ and ŝn : X → R obeying ∥fˆn − f ∥L2 (PX,Z ) = oP (1),

∥ŝn − s∥L2 (PX,Z ) = oP (1),

where f ∈ L2 (PX,Z ) is nonnegative and the distribution of s(X) has no point masses. Let {(Xi , Zi )}ni=1 be i.i.d. samples from PX,Z and independent of the training processes of fˆn and ŝn . Define n

Ĥn (t) =

1Xˆ fn (Xi , Zi ) 1{ŝn (Xi ) ≤ t}, n i=1

e n (t) = E [f (X, Z) 1{ŝn (X) ≤ t}] , H

h i Hn (t) = E fˆn (X, Z) 1{ŝn (X) ≤ t} ,

H(t) = E [f (X, Z) 1{s(X) ≤ t}] ,

and ˆ n = sup |Ĥn (t) − H(t)|, ∆

∆n = sup |Hn (t) − H(t)|.

t∈R

t∈R

e n , and H are taken over an independent copy (X, Z) ∼ PX,Z . Then Here, the expectations defining Hn , H e n (t)| = oP (1), sup |Hn (t) − H

e n (t)| = oP (1), sup |Ĥn (t) − H

t∈R

t∈R

ˆ n = oP (1). In addition, suppose that H(t∗ − ε) < α < H(t∗ + ε) holds for every ε > 0, and ∆n = oP (1)m ∆ where t̂ = sup{t ∈ R : Ĥn (t) ≤ α} and t∗ = sup{t ∈ R : H(t) ≤ α}. Then t̂ − t∗ = oP (1). Pn Proof of Lemma F.4. Write Pn g := n−1 i=1 g(Xi , Zi ) and P g := E[g(X, Z)], and let Tn be the σ-field generated by the training processes of fˆn and ŝn . Define  e n (t)|, Dn := sup |H e n (t) − H(t)|. An := sup (Pn − P ) fˆn 1{ŝn ≤ t} , Cn := sup |Hn (t) − H t∈R

t∈R

t∈R

ˆ n ≤ An + Cn + Dn , and Then ∆n ≤ Cn + Dn , ∆ e n (t)| ≤ An + Cn . sup |Ĥn (t) − H t∈R

We first control An . Conditional on Tn , the functions fˆn and ŝn are fixed and the evaluation observations are i.i.d. By conditional symmetrization, " # n X 2 E[An | Tn ] ≤ E sup εi fˆn (Xi , Zi ) 1{ŝn (Xi ) ≤ t} Tn , n t∈R i=1 where ε1 , . . . , εn are i.i.d. Rademacher random variables independent of everything else. Conditional further on the evaluation sample, reorder the observations so that ŝn (X(1) ) ≤ · · · ≤ ŝn (X(n) ) and set aj := fˆn (X(j) , Z(j) ). Since the sets {i : ŝn (Xi ) ≤ t} are nested, they are prefixes of this ordering. Thus, by Doob’s L2 maximal inequality and Jensen’s inequality, " Eε sup

n X

# εi fˆn (Xi , Zi ) 1{ŝn (Xi ) ≤ t}

t∈R i=1

(Xi , Zi )ni=1 , Tn

 1/2 n X ≤ 2 a2j  . j=1

56

Consequently,   !1/2 n 4 4  Xˆ Tn  ≤ √ ∥fˆn ∥L2 (PX,Z ) . E[An | Tn ] ≤ E fn (Xi , Zi )2 n n i=1 Since ∥fˆn ∥L2 (PX,Z ) ≤ ∥f ∥L2 (PX,Z ) + ∥fˆn − f ∥L2 (PX,Z ) = OP (1), conditional Markov’s inequality gives An = OP (n−1/2 ) = oP (1). Next,  Cn = sup P (fˆn − f ) 1{ŝn ≤ t} ≤ P |fˆn − f | ≤ ∥fˆn − f ∥L2 (PX,Z ) = oP (1). t∈R

It remains to control Dn . For each t ∈ R, let En,t := {1{ŝn (X) ≤ t} ̸= 1{s(X) ≤ t}} . By the Cauchy–Schwarz inequality,  1/2 . Dn ≤ ∥f ∥L2 (PX,Z ) sup P (En,t ) t∈R

For any η > 0, En,t ⊆ {|s(X) − t| ≤ η} ∪ {|ŝn (X) − s(X)| > η}, and therefore sup P (En,t ) ≤ ω(η) + P (|ŝn (X) − s(X)| > η) ≤ ω(η) + η −2 ∥ŝn − s∥2L2 (PX,Z ) , t∈R

where ω(η) := supt∈R P (|s(X) − t| ≤ η). Since the distribution of s(X) has no point masses, its distribution function is continuous and hence uniformly continuous, which implies ω(η) → 0 as η ↓ 0. For every fixed η > 0, the second term is oP (1). Letting first n → ∞ and then η ↓ 0 yields supt P (En,t ) = oP (1), and hence Dn = oP (1). We have thus shown e n (t)| = Cn = oP (1) sup |Hn (t) − H t∈R

and e n (t)| ≤ An + Cn = oP (1). sup |Ĥn (t) − H t∈R

Moreover, ∆n ≤ Cn + Dn = oP (1),

ˆ n ≤ An + Cn + Dn = oP (1). ∆

Finally, fix ε > 0 and define γε :=

n   o 1 ε ε min α − H t∗ − , H t∗ + − α > 0. 2 2 2

ˆ n < γε }, On the event {∆

  ε ε ˆ Ĥn t∗ − ≤ H t∗ − + ∆n < α 2 2 and   ε ε ˆ Ĥn t∗ + ≥ H t∗ + − ∆n > α. 2 2 Since fˆn ≥ 0, the function Ĥn is nondecreasing. Hence ε ε t∗ − ≤ t̂ ≤ t∗ + , 2 2 so ˆ n ≥ γε ) → 0. P(|t̂ − t∗ | ≥ ε) ≤ P(∆ Therefore, t̂ − t∗ = oP (1). 57

F.1

Proof of Lemma C.1

Proof of Lemma C.1. By the Cauchy-Schwarz inequality, it holds deterministically for any t ∈ R that n n 1X 1X ŵi Zi 1{f (Xi ) ≤ t} − ŵ(Xi )Zi 1{f (Xi ) ≤ t} n i=1 n i=1 n

1X Zi (ŵi − ŵ(Xi )) 1{f (Xi ) ≤ t} n i=1 v v v v u n u n u n u n u1 X u1 X u1 X u1 X 2 2 2 t t t ≤ Z 1{f (Xi ) ≤ t} (ŵi − ŵ(Xi )) ≤ Z t (ŵi − ŵ(Xi ))2 . n i=1 i n i=1 n i=1 i n i=1 =

Since E[Zi2 ] < ∞, we know n1

Pn

2 i=1 Zi = OP (1) by the Markov’s inequality, and therefore

v v u n u n u1 X u1 X ŵi Zi 1{f (Xi ) ≤ t} − ŵ(Xi )Zi 1{f (Xi ) ≤ t} t Zi2 t (ŵi − ŵ(Xi ))2 = OP (rn ) sup n i=1 n i=1 n i=1 t∈R n i=1 n 1X

n 1X

by Assumption A.6.

58

Record · ID 919390 · SHA-256 cafb836bfc392497
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.