Conceptio › Archive › arXiv CS
arXiv CSopen access

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias arXiv:2609.16971v1 [stat.ML] 15 Sep 2026

Nicolas Alexander Ihlo∗, Merle Behr† Faculty of Informatics and Data Science University of Regensburg, Germany September 15, 2026

In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individuals than for others. To address this two-fold challenge of prediction and interpretation, we introduce an algorithm based on decision trees and random forests for estimating individual treatment effects. Our algorithm is simple: it operates exactly like a standard random forest, but with a different splitting criterion, and requires no additional workarounds such as double machine learning or orthogonalization as used in Generalized random forests. It handles observational studies with varying treatment propensities without requiring separate estimation of the full propensity function. This is achieved by combining two splitting criteria—one targeting heterogeneity in the treatment effect, the other targeting bias correction for the average treatment effect—which together improve split point selection and automatically distinguish confounders from features responsible for heterogeneity. As a result, interpretation follows directly from the fitted tree structure itself, that is, from which features the trees split on and with which split statistics, without requiring separate post-hoc analysis. For the theoretical analysis of this algorithm, we consider a change point model with step functions for potential outcomes and treatment propensity and provide insights into the theoretical underpinnings of our approach. Simulation studies show that our ∗ †

[email protected] [email protected]

1

simple algorithm achieves comparable, and often better, prediction accuracy than existing methods, while substantially improving interpretability. Keywords: causal inference, statistical machine learning, interpretable machine learning, heterogeneous treatment effect estimation

1. Introduction Estimating treatment effects, i.e., changes in an outcome due to some intervention, is of interest across many fields [Manski, 1993, Imbens and Rubin, 2015, Angrist and Pischke, 2015], including medicine, marketing, public policy, and economics. Beyond average effects, many applications benefit from individual-level estimates: for example, estimating the effectiveness of a drug for a specific patient enables individualized medicine. For such estimates to be useful in practice, however, accurate predictions are often not enough by themselves; understanding why the treatment effect differs across individuals, i.e., interpreting the estimation method, is equally important. In this paper, we focus on the estimation of individual treatment effects with a particular emphasis on interpretability. In many settings, treatment effects cannot be estimated from randomized experiments and one has to rely on observational data instead. This introduces the problem of confounding: features that influence both the outcome and the treatment propensity, and which can bias treatment effect estimates even when they do not affect the treatment effect itself [Hernan and Robins, 2020]. Crucially, this means a feature can appear relevant to treatment effect estimation for two entirely different reasons: either because it is a confounder that must be corrected for, or because it genuinely drives treatment effect heterogeneity. Distinguishing between these two cases is precisely the interpretation we are interested in. Throughout this paper, we assume that all confounders are observed, i.e., we rule out hidden confounding, and focus on the estimation of individual treatment effects under observed confounding. Machine learning (ML) methods are well suited to this task, as they can learn flexible structures directly from data with little manual adjustment. General supervised learning methods, originally developed for classification or regression, can be adapted for treatment effect estimation, for example via metalearners [Künzel et al., 2019] and double/debiased machine learning (DML) [Chernozhukov et al., 2018]; other methods are purpose-built for causal inference, building on neural networks, for example TarNet [Shalit et al., 2017], DragonNet [Shi et al., 2019], and RA-Net [Curth and van der Schaar, 2021]. However, many of the most powerful ML methods act as black boxes and offer little insight into their prediction mechanism, even though such insight is often required for interpretation and downstream decision-making. Random forest (RF) [Breiman, 2001], originally developed for supervised learning but since extended to causal inference as well, offers a compromise between these two goals: it achieves state-of-the-art prediction accuracy, particularly on tabular data [Shmuel et al., 2025], while remaining, to a certain extent, interpretable through its underlying tree structure. This structure has, for instance, been used to gain insight into the mechanism of the fitted forest, e.g., by feature importance statistics such as mean decrease in impurity

2

(MDI) [Breiman, 2001], as well as several extensions and other approaches, e.g., MDI+ [Agarwal et al., 2025], iRF [Basu et al., 2018], LSSFind [Behr et al., 2022], and TreeSHAP [Lundberg et al., 2020]. Several RF-type algorithms have been proposed for estimating treatment effects, and it is this line of work that we build on and extend in this paper; we review these approaches in the following subsection. Our main motivation for building on RF, rather than a black-box method such as a neural network, is its interpretability, and accordingly, our goal is to preserve as much of this interpretability as possible in the RF-type algorithm we propose.

1.1. Previous Tree-based Methods for Causal Inference A popular adaptation of RF is causal forest [Wager and Athey, 2018], which itself uses a simplified version of the causal-inference-specific adaptation of decision trees, causal tree, introduced by Athey and Imbens [2016]. While causal forest proposes two possible procedures, we focus here on procedure 1 (double sample trees), as only it incorporates outcomes in the construction of the trees and is therefore able to model treatment effect heterogeneity via the tree structure. Both causal tree and causal forest build trees by selecting splits that maximize the variance of treatment effect predictions, a criterion derived from an analogy to mean squared error (MSE) minimization in regression (see Section 3.1 for details); however, this analogy relies on properties of the regression setting that do not hold in general for causal inference, and the resulting susceptibility to confounding has been noted previously, e.g., by Athey et al. [2019]. Notably, among the tree-based methods discussed below, causal forest is the only one that preserves the simple structure of the original RF: the individual treatment effect is estimated directly as the average prediction of a single ensemble of decision trees, each of which remains interpretable on its own. Other adaptations of tree-based methods for causal inference typically sacrifice this simplicity. Metalearners [Künzel et al., 2019], for instance, do not estimate the treatment effect directly with a single forest, but instead combine separate forests fit to the outcomes under treatment and under control, so that the treatment effect estimate is no longer the direct output of one interpretable ensemble. Generalized random forest (GRF) [Athey et al., 2019], used as an alternative implementation of causal forest, instead modifies the splitting criterion via gradient-based approximations to an estimating equation and a correspondingly changed prediction mechanism, while additionally incorporating an orthogonalization step to reduce the influence of confounders. This orthogonalization step is similar in spirit to DML, reflected in the name “CausalForestDML” used for its implementation in the popular package EconML [Battocchi et al., 2019]. While effective at reducing confounding bias, this two-step procedure again departs from the direct, single-ensemble structure that makes RF interpretable in the first place. A method to improve feature importance for treatment effect heterogeneity in GRF was recently developed by Bénard and Josse [2025], but requires additional post-hoc steps which are computationally costly, and, more importantly, does not address our goal of interpreting the fitted tree structure itself. Further adaptations include orthogonal random forest

3

Sample Treatment T Treated outcome Y T =1 Control outcome Y T =0 Feature X1 Feature X2

1

2

3

4

5

6

0

1 0.5

0

1 1.5

0

0.5

1.5

1 1.5

0.5

0.5 0 0

0.5

0 0

0.5 0 0

1.5

1 0

1.5 1 1

1.5

1 1

Table 1: Example for data in a potential outcomes model, with binary treatment T , potential outcomes Y T =1 and Y T =0 , and two binary features X1 , X2 . Note that for each sample only the outcomes corresponding to the treatment assignment can be observed (shown in bold). The outcomes in italics are unobserved and cannot be used for estimating the treatment effect. [Oprescu et al., 2019]; see Jiang et al. [2021] for an overview and comparison of these methods. Causal forest is therefore the only method that retains the interpretable, single-ensemble structure of RF for individual treatment effect estimation. However, as noted above, it does not explicitly account for confounders during split selection, so splits driven purely by confounding cannot be distinguished from splits that reflect genuine treatment effect heterogeneity, undermining the very interpretability its tree structure would otherwise offer. In this paper, we propose a modification of RF for causal inference, denoted as IntCF (interpretable causal forest), that follows the general idea of causal forest but uses a different splitting criterion, one that explicitly accounts for confounders at the split selection step. This direct inclusion of both heterogeneity and confounder detection in the tree structure opens the possibility for improved interpretation, especially for distinguishing features based on how they affect the treatment effect prediction, while preserving the simple, single-ensemble structure that makes RF interpretable.

1.2. Confounding versus Heterogeneity: an Illustrative Example In the following, we provide an illustrative example, which demonstrates why the original splitting criterion of causal forest [Wager and Athey, 2018] fails to result in interpretable tree structures under confounding, and how our new splitting criterion in IntCF improves on this. Consider the data in Table 1, with binary treatment T , potential outcomes Y T =1 and Y T =0 , and two binary features X1 , X2 . For each sample, only the upright, bold outcome can be observed; the outcome in italics is unobservable. The treatment effect Y T =1 − Y T =0 is zero for every sample. However, directly estimating the average treatment effect from the observed outcomes gives a biased estimate: the difference between the mean observed treated outcome 0.5+1.5+1.5 = 76 and the mean observed 3 0.5+0.5+1.5 5 1 control outcome = 6 is 3 ̸= 0, because of the higher prevalence of treatment 3 among samples with higher potential outcomes. Splitting the data into two subsets based on feature X1 would eliminate this bias, leading to conditional average treatment effect estimates of 0 in both subsets. However,

4

IntCF

Causal forest

X1 ≤ 0.5 d Bias) d Split criterion: max(Het,

X2 ≤ 0.5 d = 1 Split criterion: Het 18 no separate criteria for heterogeneity and bias samples: 1, 2, 3, 4, 5, 6 estimate τ̂ = 13

d=0 Heterogeneity criterion: Het d = 1 Bias criterion1 : Bias 9 samples: 1, 2, 3, 4, 5, 6 estimate τ̂ = 13

samples: 1, 2, 3 estimate τ̂ = 0

samples: 1, 2, 3, 4 estimate τ̂ = 12

samples: 4, 5, 6 estimate τ̂ = 0

samples: 5, 6 estimate τ̂ = 0

Figure 1: Decision stumps after the first split for the data in Table 1, using the combined splitting criterion of our proposed method, IntCF (left), and the criterion of the original causal forest algorithm [Wager and Athey, 2018] (right). τ̂ denotes the estimated treatment effect in each node, computed as the difference between the mean observed outcome under treatment and under control. IntCF selects the split on X1 , correctly identifying it as bias correction rather than heterogeneity, and yields unbiased estimates in both child nodes; causal forest instead selects the split on X2 , which appears to indicate heterogeneity but remains biased due to the unaddressed confounding from X1 . the splitting criterion used in causal forest [Wager and Athey, 2018], which is based on heterogeneity in outcomes, will not select this split, as the estimated outcome is the same on both sides. Even worse, the split on feature X2 , which results in an estimate of 12 for the treatment effect of samples 1–4, an apparent effect that is itself an artifact of confounding rather than true heterogeneity, would seem preferable. This is in contrast to the modification we propose in IntCF below. There, we explicitly combine two different splitting criteria, one for heterogeneity in the treatment effect and one for bias correction. With this approach, the change in predicted average treatment effect from splitting on X1 is recognized as bias correction rather than heterogeneity, so the split is still used when building the tree. Interpretability is preserved, since the split is explicitly labeled as reducing bias rather than as improving the estimated treatment effect heterogeneity. Applying IntCF to this example yields the decision stump shown in Figure 1 (left) after the first split. While IntCF estimates the correct treatment effect in this example, causal forest produces a biased estimate, as the confounding introduced by feature X1 is not removed (Figure 1, right). Our combined splitting criterion in IntCF correctly indicates that the selected split corrects bias rather than reflecting treatment effect heterogeneity. 1

Note that our bias criterion is an estimate for the squared bias.

5

1.3. Outline of Paper The remainder of this paper is organized as follows. In Section 2, we introduce the considered setting, notation, and data model. In Section 3, we introduce the IntCF algorithm, with a particular focus on our new splitting criteria, as well as an additional step for honest predictions and validation of splits. In Section 4, we derive theoretical properties of our new splitting criteria under a change point model. In Section 5, an extensive simulation study, including a semi-synthetic benchmark as well as a real data example, demonstrates that IntCF achieves competitive prediction accuracy while substantially improving interpretability. Finally, in Section 6, we discuss our results and outline open questions.

2. Preliminaries 2.1. Data and Potential Outcomes Model We consider a potential outcomes framework [Rubin, 1974], in which the data generating process is described by the joint distribution of (T, X, Y T =0 , Y T =1 ), with binary treatment T ∈ {0, 1}, a feature vector X ∈ Rd , and potential outcomes Y T =0 and Y T =1 . For each unit, only the potential outcome corresponding to the realized treatment is observed, i.e., we observe Y ∼ Y T , while the other potential outcome remains unobserved. We aim to predict the conditional average treatment effect (CATE)   τ (x) := E Y T =1 − Y T =0 | X = x , (1) for a given feature vector x ∈ Rd ; throughout this paper, we refer to τ (x) as the individual treatment effect, since it is the treatment effect estimate available given the observed covariates x. Beyond predicting τ (x), we are interested in interpreting treatment effect heterogeneity, i.e., in identifying which of the d features actually influence τ (x). We observe N independent and identically distributed copies of (T, X, Y ), (ti , xi = (xi,j )dj=1 , yi ),

i = 1, . . . , N.

The main difficulty in estimating τ (x) is that, for each observation i, only one of yiT =1 and yiT =0 is observed; the individual treatment effect τi := yiT =1 − yiT =0 can therefore never be observed, so standard machine learning methods for regression, which require direct observations of the target variable, cannot be applied to estimate τ directly. We further define p(x) := P (T = 1 | X = x) (2)

6

for the treatment propensity, and µ(x) :=

 1  T =0 E Y + Y T =1 | X = x 2

(3)

for the main outcome.

2.2. Decision Trees and RFs Decision trees for classification and regression tasks were proposed by Breiman et al. [1984] and later extended to RFs [Breiman, 2001]. RFs use a collection of decision trees, whose predictions are combined (for regression, usually by taking the average) to obtain an ensemble prediction. Each tree divides the feature space into subsets by performing binary splits based on a threshold applied to one of the features at each inner node. For deciding which split to use, a splitting criterion is applied, most commonly the impurity decrease according to an impurity measure, for example mean squared error (MSE) for regression. Specifically, when selecting a split for a node eparent , which contains samples with indices in Jparent , a collection of potential splits is considered; this might be all possible splits or a subset of these. Each potential split creates two child nodes, eL and eR , containing samples with indices JL and JR , respectively, such that Jparent = JL ⊔ JR is a disjoint union. We denote the number of samples in the left and right child nodes by nL = |JL | and nR = |JR |, respectively. In a regression setting, each training sample has an attached response value yi . For each of these three nodes, a prediction and an impurity value are computed; comparing these then yields the impurity decrease for the candidate split. Finally, the potential split with the highest impurity decrease is selected as the actual split, and the procedure is repeated recursively at the two child nodes until a stopping criterion is reached.

3. Interpretable Causal Tree and Forest Algorithm 3.1. Correcting Heterogeneity Splitting Rule for Confounding As outlined in Section 2.2, a central element for the construction of decision trees from training data is the splitting rule. In this section, we will derive a new splitting rule, which can be applied to data from a potential outcomes model in a causal inference setting. Recall that the main difficulty compared to a regression setting is that for each observation i, only one of yiT =1 and yiT =0 is observed; the individual treatment effect τi = yiT =1 − yiT =0 can never be observed directly, hence, standard splitting rules as used for decision trees in regression, cannot be applied directly. Previous work has therefore considered different splitting criteria which overcome this problem [see, e.g., Athey and Imbens, 2016, Wager and Athey, 2018, Athey et al., 2019]. To motivate our new splitting criteria, as well as previously considered splitting criteria for causal forest [Wager and Athey, 2018], we start with a recap of the impurity decrease via MSE in the regression setting, the standard split statistic of RF implementations. There,

7

the splitting criterion is obtained by maximizing among possible splits Jparent = JL ⊔ JR the decrease in MSE, that is,   X X X 1 1  ∆RF (L, R) := (ŷparent − yi )2 − (ŷL − yi )2 + (ŷR − yi )2  (4) N N i∈Jparent

i∈JL

i∈JR

1 with node predictions obtained as averages, that is, ŷparent := nL +n R similar 1 X 1 X ŷL := yi and ŷR := yi . nL nR i∈JL

P

i∈Jparent yi and

(5)

i∈JR

Note that for such averages we have ŷparent =

nL · ŷL + nR · ŷR . nL + n R

(6)

Using (5), it is easy to see that (4) is equal to  1 nL (ŷparent − ŷL )2 + nR (ŷparent − ŷR )2 . N

(7)

Moreover, using (6), it is easy to see that (7) is equal to nL · n R (ŷL − ŷR )2 . N (nL + nR )

(8)

+nR Note that (8) equals, up to a factor of nLN (which is constant within each node), the variance among the predictions at the two child nodes. For the causal inference setting with outcome of interest being the (unobserved) treatment effect τi , the direct analog to the splitting criterion from (4) is to consider   X X 1 X 1  (τ̂parent − τi )2 − (τ̂L − τi )2 + (τ̂R − τi )2  . (9) N N i∈Jparent

i∈JL

i∈JR

For the node predictions τ̂parent , τ̂L and τ̂R one can obtain natural estimates with the difference-in-means estimator at a specific node e, that is, τ̂e :=

X X 1 1 yi − yi , |{i ∈ Je : ti = 1}| |{i ∈ Je : ti = 0}| i∈Je ti =1

(10)

i∈Je ti =0

e.g., for the parent node e = parent, for the left child node e = L, or for right child node e = R. The problem with the decrease in MSE as in (9) is, however, that the individual treatment effects τi are not observed and hence, one cannot use (9) for a splitting criterion. Therefore, the proposal of Wager and Athey [2018] for causal forest was to consider the analog of (8) instead. That means, they propose to select splits Jparent = JL ⊔ JR such

8

that the variance among the predicted treatment effects at the child nodes is maximized, i.e., they maximize at each node d := ∆CF (L, R) := Het

nL · n R (τ̂L − τ̂R )2 . N (nL + nR )

(11)

We will call this the heterogeneity criterion. The major problem with using the heterogeneity criterion (11) as a replacement for the decrease in MSE in (9), is that the corresponding analog equations to (5) and (6), which were needed to derive the equivalence between (4) and (8) in the regression setting, do not hold in the causal inference setting. More precisely, the analog to (5) is given by the condition τ̂L =

1 X τi nL

and τ̂R =

i∈JL

1 X τi nR

(12)

i∈JR

and the analog to (6) is given by the condition τ̂parent =

nL τ̂L + nR τ̂R . nL + n R

(13)

In contrast to the regression setting, (12) and (13) are not guaranteed to hold in the causal inference setting. However, if these conditions do hold, one can derive equivalence between (11) and (9), as the following two lemmas show. Lemma 1. If (12) holds, then the decrease in MSE (9) equals the mean squared change in prediction  1 nL (τ̂parent − τ̂L )2 + nR (τ̂parent − τ̂R )2 . (14) N The proof is given in the appendix. Note that (14) is the analog to (7) from the regression setting. Lemma 2. If (13) holds, then the mean squared change in prediction (14) equals the d in (11). heterogeneity criterion Het The proof is given in the appendix. Both (12) and (13) may fail to hold under confounding, and either violation breaks d in (11) and the decrease in MSE the equivalence between the heterogeneity criterion Het in (9). However, the two assumptions differ fundamentally in nature. A violation of (13) is a property of the candidate split itself: for any given split, τ̂parent , τ̂L , and τ̂R can be computed directly from the observed data, and one can check whether the parent estimate coincides with the sample-size-weighted average of the child estimates. Note that even in a randomized controlled trial, i.e., in the absence of any confounding, (12) and (13) will typically not hold exactly, simply due to finite-sample noise. However, without confounding both conditions hold in expectation. Under confounding, in contrast, they need not hold even in expectation. A strong violation of (13), beyond what would be expected from sampling noise alone, therefore indicates a systematic effect rather than a chance fluctuation and, as illustrated in our example above, arises precisely when

9

treatment propensity differs between the two child nodes, i.e., it is a direct symptom of confounding along the considered split. This makes it possible to explicitly measure and correct for such violations as part of the splitting criterion. In contrast, (12) is not a property of any particular split, but a standing identification assumption: it requires that, conditional on already being in a given node, the differencein-means estimator (10) is unbiased for the true average treatment effect within that node, i.e., that no unobserved confounding remains once conditioned on the node. This within-node unconfoundedness assumption underlies the use of the difference-in-means estimator at any node regardless of the tree structure, and is required by essentially every tree-based method for treatment effect estimation, including our own leaf-level estimates. Since it is not tied to any single candidate split, it cannot be assessed or corrected by comparing τ̂parent , τ̂L , and τ̂R at a given step, but only by conditioning on the relevant confounders while growing the tree as a whole. We therefore focus our new splitting criterion on correcting for violations of (13), while retaining (12) as a standing assumption, consistent with its role throughout the tree-based causal inference literature. As Lemma 2 shows, whenever the mean squared change in prediction (14) and the d in (11) are not equal, this can be directly attributed to heterogeneity criterion Het a violation of (13) and hence, a bias correction. We therefore treat the difference between (14) and (11) as one component of our splitting criterion, quantifying the extent to which a split corrects for confounding, as opposed to detecting heterogeneity in the treatment effect. The following Lemma gives an explicit expression how this difference corresponds to the violation of (13). Lemma 3. The difference of (14) and (11) is given by d := nL + nR Bias N

  nL τ̂L + nR τ̂R 2 τ̂parent − . nL + n R

(15)

We call (15) the bias criterion. The proof is given in the appendix. In total, we have shown the following decomposition MSE decrease ≈

 1 d + Bias. d nL (τ̂parent − τ̂L )2 + nR (τ̂parent − τ̂R )2 = Het N

(16)

Note that the decomposition in (16) is analog to a classical bias-variance decomposition. d and Bias, d capture different reasons for improved predictions: The two criteria, Het identified heterogeneity (which we measure by the variance of predictions) and bias correction (which we estimate by the square of change of average prediction). To improve interpretability of the final trees, we want to separate these two reasons. Therefore, at each node, we select the split that has the strongest effect from either of d in (11) and the bias criterion these two sources. As both the heterogeneity criterion Het d in (15) are directly derived from the mean squared change (14), their values are Bias comparable. Hence, with IntCF we propose to use the maximum of both criteria to select the split. That is, our splitting criterion for IntCF is obtained by maximizing among possible splits Jparent = JL ⊔ JR the maximum of heterogeneity and bias improvement,

10

that is,   d Bias d . ∆IntCF (L, R) := max Het,

(17)

For our theoretical analysis, we also consider signed versions of these criteria, which we refer to as signed splitting criteria, defined by √ ± nL · n R d (τ̂L − τ̂R ), (18) Het := nL + n R the signed heterogeneity, and d ± := τ̂parent − nL τ̂L + nR τ̂R , Bias nL + n R

(19)

+nR the signed bias. By squaring and then multiplying by nLN (a constant factor at each d d ±, node), one recovers the heterogeneity criterion Het from the signed heterogeneity Het d from the signed bias Bias d ±. and the bias criterion Bias

3.2. Honest Validation of Splits In this section, we introduce a second modification to standard RF that we propose for IntCF, complementing the new splitting criterion (17) introduced in Section 3.1. Importantly, this modification does not affect how the trees themselves are grown: trees are constructed exactly as before, simply using the combined splitting criterion (17) in place of the standard impurity decrease. Instead, this modification concerns how we later summarize, at each node, the heterogeneity and bias contribution identified during tree construction, for example when computing a mean-decrease-in-impurity-type measure of feature importance. While building the tree, especially for nodes close to the root, the criteria used to select splits might not yet be informative: as discussed in Section 3.1, the within-node unconfoundedness assumption (12) need not hold until the tree has conditioned on the relevant confounders, so for nodes near the root the estimates entering the heterogeneity d (11) and the bias criterion Bias d (15) may still be affected by confounding criterion Het that is only corrected at later splits. To address this problem, we use out-of-bag (or hold-out) samples to summarize the heterogeneity and bias contribution of each split. This overall idea is not new and has been used in various forms before, both for prediction, where it is known as honest prediction, and for feature importance, e.g., in debiased or out-of-bag variants of MDI [Li et al., 2019, Zhou and Hooker, 2021]. We build on both of these ideas. First, we also use out-of-bag or hold-out samples for our final predictions; this is simply the standard honest-prediction procedure of double-sample trees, as used in causal forest [Wager and Athey, 2018], and not itself a new contribution. Second, and in addition, we use these samples as validation set to re-evaluate the split statistics at every node, similar to out-of-bag MDI approaches d (11) and the bias for RF. However, simply recomputing the heterogeneity criterion Het d criterion Bias (15) on out-of-bag samples, using the same difference-in-means estimator

11

at each node, is not sufficient in the causal inference setting: while using held-out data addresses the classical overfitting bias that motivates honest estimation and out-of-bag MDI, it does not address confounding. In particular, the within-node unconfoundedness assumption (12) need not hold until the tree has conditioned on the relevant confounders, regardless of whether the underlying estimate is computed on training or on out-of-bag data. We therefore develop a modified validation procedure in the remainder of this section, which draws on the fully grown tree (or forest) to correct for this remaining confounding. To use the validation set to also validate split attribution at all nodes, we extend the double-sample-trees procedure to compute predictions from this data set not only at leaf nodes, but also at inner nodes. To obtain both an estimate of corrected bias and of identified heterogeneity, we compute two separate estimates for each node, using the following procedure. For the first estimate, we feed the validation samples through the considered tree and, at each node, compute an estimate as before using the difference-in-means estimator (10), but now based on validation samples instead of training samples. Denote the indices of validation samples in node e by Ie ; throughout this section, ne , nL , nR , and N refer to the corresponding counts of validation samples, analogous to their earlier use for training samples in Section 3.1. At node e, we call this estimate τ̂ef =

X X 1 1 yi − yi |{i ∈ Ie : ti = 1}| |{i ∈ Ie : ti = 0}| i∈Ie ti =1

(20)

i∈Ie ti =0

the forward prediction. It might still be influenced by bias that is only corrected at later splits. As with double-sample trees, at leaf nodes these estimates are used as predictions for samples falling in the respective leaf. For the second estimate, we use the final predictions τ̂ival for the validation samples (these can be either the predictions from the single tree or from an entire RF). At each node e, a new estimate τ̂eb for the CATE can then be computed by averaging the final predictions of validation samples contained in this node, i.e., τ̂eb =

1 X val τ̂i . |Ie |

(21)

i∈Ie

n τ̂ b +n τ̂ b

b For inner nodes, this estimate can also be computed as τ̂parent = LnLL +nRR R . Since these estimates can therefore be computed iteratively from the leaves up to the root, we call them backward predictions. If the final predictions are unbiased, so are the backward predictions. One task in validating split attribution is to estimate the heterogeneity between treatment effects in the child nodes at each split. For this, we use a formula analogous to the heterogeneity splitting criterion (11), but now using the backward predictions, i.e.,

d v := Het

2 nL · n R  b τ̂L − τ̂Rb . N (nL + nR )

12

(22)

We refer to (22) as the heterogeneity validation criterion. For the bias, such a direct transfer is not possible: unlike the heterogeneity criterion, the backward predictions already incorporate bias corrections made throughout the tree, not only the correction attributable to the split under consideration. Instead, we can estimate the squared bias corrected by the remainder of the tree by comparing forward 2 +nR f b and backward estimates, using nLN τ̂parent − τ̂parent . But this also includes bias that was only corrected in descendant nodes. For that reason, we instead use the decrease in squared bias,  2 n  2 n  2 L R b f b f b d v := nL + nR τ̂ f Bias − τ̂ − τ̂ − τ̂ − τ̂ − τ̂ , (23) parent parent L L R R N N N which we refer to as the bias validation criterion. In each node, one can show that sum of the heterogeneity validation criterion and the bias validation criterion equals the decrease in mean squared error obtained by replacing the forward estimate of the average treatment effect with the final individual predictions, as the following lemma shows. Lemma 4. For the validation criteria (22) and (23) we have         X X X 2 2 2 v v f d + Bias d = 1  Het τ̂parent − τ̂ival − τ̂Lf − τ̂ival − τ̂Rf − τ̂ival  . N i∈Iparent

i∈IL

i∈IR

The proof is given in the appendix. Again, note that the decomposition in Lemma 4 is analog to a classical bias-variance decomposition, similar to (16) for the splitting criteria.

3.3. Summary of Algorithm We now combine the two modifications introduced above—the amended splitting criterion ∆IntCF (17) from Section 3.1 and the honest validation procedure from Section 3.2—into a complete algorithm for growing decision trees and RFs. Analogous to how a RF is built from individual decision trees (Section 2.2), we first describe how to grow a single tree, which we call the interpretable causal tree (IntCT); the corresponding forest, IntCF, is then obtained by combining many such trees, as described below. 3.3.1. Interpretable Causal Tree (IntCT) 1. Split the samples into a training set J and a validation set I. 2. Build the tree based on the training samples J: a) At the current node, consider all possible splits, determine the resulting sample sets for the two child nodes, and compute the treatment-effect estimates τ̂L , τ̂R for these child nodes using the difference-in-means estimator (10). d (11) and the b) For each possible split, compute the heterogeneity criterion Het d bias criterion Bias (15).

13

  d Bias d (17), i.e., the split c) Select the split that maximizes ∆IntCF = max Het, with the largest value of either criterion. d) Create the new child nodes and assign them their respective training samples. e) If no stopping criterion (e.g., maximum depth or minimum leaf size) has been reached, repeat the above steps at each child node. 3. Validate the splits and compute predictions using the validation samples (see Section 3.2); this step retraces the tree built in Step 2, but now operates on I instead of J: a) Compute the forward predictions τ̂ef (20) for all nodes of the tree. b) At the leaves, use the forward prediction as the tree’s prediction for samples that fall into the respective leaf; this is the standard “honest prediction” procedure, as proposed, e.g., in [Wager and Athey, 2018]. c) Use the tree (or the forest) to compute estimated treatment effects τ̂ival for all validation samples. d) Use these estimated treatment effects to compute the backward predictions τ̂eb according to (21) for all nodes of the tree. d v (22) and the bias valie) Compute the heterogeneity validation criterion Het d v (23) from the forward and backward predictions, and dation criterion Bias store their values at each node. 4. The final result is the tree together with its leaf-level predictions and the validation d v , Bias d v stored at each split. criterion values Het 3.3.2. Interpretable Causal Forest (IntCF) In a forest, the same algorithm is used to grow each tree, with the training/validation split of Step 1 performed independently for each tree, so that every sample is used, either in the training or in the validation role, in every tree of the forest. This plays a role analogous to the bootstrap sampling used in standard RFs, where the samples not drawn in the bootstrap, the out-of-bag samples, serve as the validation set. As with standard RFs, only a random subset of features is considered at each split when growing each tree. 3.3.3. Feature Importance Following the original mean-decrease-in-impurity approach to feature importance [Breiman et al., 1984], we sum, over all nodes that split on a given feature, a measure of that split’s contribution to obtain the feature’s importance. Instead of the impurity decrease, we use a validation criterion for this purpose: by default, only the heterogeneity validation d v (22) is used, but the bias validation criterion Bias d v (23), or the sum of criterion Het both, can also be used if desired. In many implementations, for example scikit-learn [Pedregosa et al., 2011], it is customary to scale the importance measure so that it sums to 1. In contrast to this

14

standard MDI procedure, the bias validation criterion can take negative values, so feature importance scores that include it may also be negative. We therefore scale our feature importance scores so that the sum of the non-negative scores is 1.

4. Population-level Analysis of the Splitting Criteria under a Change Point Model In this section, we provide theoretical support for the central claim of this paper: that the combined splitting criterion introduced in Section 3.1 correctly distinguishes splits driven by confounding-induced bias from splits that reflect genuine heterogeneity in the treatment effect. Rather than finite-sample guarantees, we study this question at the population level, i.e., in the limit as the sample size N → ∞, under a stylized change point model in which the treatment effect, the propensity, and the main outcome are step functions of a single covariate. The remainder of this section is organized as follows. Section 4.1 derives the population d ± and the signed bias criterion Bias d ±. limits of the signed heterogeneity criterion Het Section 4.2 explains why, due to the marginalization inherent to decision-tree splits, it suffices to study a univariate change point model, which we introduce formally. Section 4.3 then shows that any such change point model decomposes into a component with a constant treatment effect and a component for which the difference-in-means estimator is unbiased. Finally, Section 4.4 shows that our two splitting criteria correctly recover this decomposition: the population bias criterion attains an extremum at the true change point exactly when bias is present, and the population heterogeneity criterion attains an extremum there exactly when the treatment effect is genuinely heterogeneous. This provides theoretical justification for the algorithm’s ability to separate these two sources of improvement, which underlies its interpretability.

4.1. Population-level Analysis Rather than deriving finite-sample results, we focus on the corresponding population analogs of (18) and (19) in order to elucidate general structural properties of these splitting criteria. To this end, consider a decision tree with an inner node e and the associated hyper-rectangle R(e). At node e, we examine a candidate split along a given variable, say x1 , at threshold k. Throughout the analysis, we impose the following simplifying assumptions: A1 The feature vector X is uniformly distributed on the unit hypercube, that is, X ∼ U ([0, 1]d ). A2 The outcomes Y T =1 , Y T =0 are bounded. A3 The positivity assumption holds, meaning that there exists some ϵ > 0 such that, for all x ∈ [0, 1]d , ϵ ≤ P (T = 1 | X = x) ≤ 1 − ϵ.

15

A4 The hyper-rectangle R(e), with positive volume µ(R(e)) > 0, the choice of the splitting variable (without loss of generality assumed to be x1 ), as well as the considered threshold k are independent of the training data D = {(y1 , x1 , t1 ), . . . , (yN , xN , tN )}. Note that, since decision trees are invariant under monotone transformations of the covariates, the first assumption is essentially equivalent to assuming independence among the different features. Assumptions A1 and A2 are standard in the analysis of decisiontree-based methods [see, for example, Behr et al., 2022, Wager and Athey, 2018]. Assumption A3 is a requirement for causal inference in general [compare, for example, Rosenbaum and Rubin, 1983]. The final assumption is clearly violated for our intCT algorithm, since all split points in the tree are selected in a data-dependent manner, implying that the resulting hyper-rectangle R(e) likewise depends on the training data D. Nevertheless, extending the subsequent results to this setting appears to be relatively straightforward and mainly of a technical nature; see Remark 1 for details. Under Assumptions A1–A4, it follows directly from the law of large numbers and the continuous mapping theorem that, when the training data D are generated i.i.d. according to the data-generating process P (Y T =0 , Y T =1 , X, T ), for any fixed threshold k p d± → Het Het± (k),

p d± → Bias Bias± (k),

as N → ∞.

The corresponding population quantities are given by p    Het± (k) := k(1 − k) · ȲLT =1 − ȲLT =0 − ȲRT =1 − ȲRT =0 ,

(24)

(25)

and    Bias± (k) := Ȳ T =1 − Ȳ T =0 − k · ȲLT =1 − ȲLT =0 − (1 − k) · ȲRT =1 − ȲRT =0

(26)

where, for a ∈ {0, 1},   Ȳ T =a := E Y T =a | X ∈ R(e), T = a ,   ȲLT =a := E Y T =a | X1 ≤ k, X ∈ R(e), T = a ,   ȲRT =a := E Y T =a | X1 > k, X ∈ R(e), T = a are expectations of observed outcomes in the whole node as well as to left and right of the considered split, respectively. In the following, we study the properties of the population counterparts Het± (k) and Bias± (k) associated with our proposed splitting criteria. Remark 1. To dispense with Assumption 4, it would suffice to establish the convergence result in (24) uniformly over all hyper-rectangles R(e) with a fixed minimal volume. More precisely, one would require a result of the form that, for any ϵ > 0, " # ± d − Het± (k) > ϵ → 0 as N → ∞, P sup Het R,k,j∈{1,...,d}, µ(R)>δ

and an analogous statement for the bias term. The assumption of a minimal volume can be enforced in practice by imposing a maximal tree depth and by requiring splitting thresholds

16

k to be bounded away from zero and one; both restrictions are standard in the analysis of RFs [see, e.g., Scornet and Hooker, 2026]. Establishing such uniform convergence results should be feasible using standard tools from empirical process theory, including uniform convergence arguments and concentration inequalities [see, e.g., Behr et al., 2022]. Since this extension is largely technical and does not affect the qualitative insights of our analysis, we do not pursue it further here and instead focus on the population quantities Het± (k) and Bias± (k) in (25) and (26) directly.

4.2. Marginalization Effects By construction, when a decision tree considers a split along a given variable Xj , say X1 , the influence of the remaining variables X2 , . . . , Xd is marginalized out through averaging, as reflected in the conditional expectations in (25) and (26). Consequently, once a specific variable X1 is selected, decision trees can capture only the marginal effect of this variable. This is an intrinsic property of decision trees and gives rise to well-known limitations of such methods. This marginalization is visible directly in the population criteria themselves: Het± (k) and Bias± (k) in (25) and (26) depend on the joint distribution of (X, T, Y T =0 , Y T =1 ) only through X1 and the conditioning event X ∈ R(t), which is fixed across both candidate children. It therefore suffices, for the purpose of studying the behavior of a single candidate split, to consider a data-generating process in which only one covariate, say x1 , carries any signal. Here, we study a univariate change point model for CATE, propensity, and main outcome, in analogy to the general definitions of τ, p, µ in (1), (2) and (3): τ (x1 ) := β0 + β1 1x1 ≤γ , p(x1 ) := q0 + q1 1x1 ≤γ ,

(27)

µ(x1 ) := α0 + α1 1x1 ≤γ . Each coefficient governs a distinct aspect of the model. The slope β1 governs treatment effect heterogeneity: β1 = 0 corresponds to a constant treatment effect τ (x1 ) = β0 , i.e., a model without treatment effect heterogeneity. Analogously, q1 governs the degree of confounding through x1 : q1 = 0 corresponds to a constant propensity p(x1 ) = q0 , i.e., a setting in which treatment assignment does not depend on x1 , as in a randomized trial. Finally, α1 governs whether x1 has a direct, prognostic effect on the main outcome: α1 = 0 corresponds to a constant main outcome µ(x1 ) = α0 . Intuitively, confounding bias for the difference-in-means estimator arises when x1 affects both treatment assignment (q1 ̸= 0) and the potential outcomes; Theorem 5 below makes this precise. By construction, this model has a unique natural split point at γ: splitting exactly there separates τ , p, and µ into constant pieces on either side, which is what allows the resulting heterogeneity and bias to be exactly quantified. With µ and τ specified by the change point model above—now understood as depending on x only through its first coordinate x1 —the potential outcomes are defined as 1 1 Y T =0 (x) = µ(x) − τ (x) + ε0 and Y T =1 (x) = µ(x) + τ (x) + ε1 2 2

17

(28)

with mean-zero noise terms ε0 , ε1 that are assumed to be independent of one another and of X.

4.3. Bias–Heterogeneity Decomposition of the Change Point Model In the following, we provide theoretical insight into why the population splitting criteria Het± (k) and Bias± (k) in (25) and (26) are able to distinguish between heterogeneity in the treatment effect and bias induced by confounding. To this end, we first observe that any change point model of the form (27) can be decomposed into the sum of two components: one corresponding to a model with a constant treatment effect, and another for which the difference-in-means estimator is unbiased. This decomposition is formalized in the following lemma. More precisely, consider a potential outcome model (Y T =0 , Y T =1 , X, T ) with treatment effect τ (x), main outcome µ(x), and propensity p(x) as defined in Section 2.1. We say that the difference-in-means estimator is unbiased if     E [τ (X)] = E Y T =1 | T = 1 − E Y T =0 | T = 0 , (29) or, equivalently (recall Equation 28), if     E [τ (X)] = E µ(X) + 12 τ (X) | T = 1 − E µ(X) − 21 τ (X) | T = 0 . Theorem 5. Consider a potential outcome model (Y T =0 , Y T =1 , X, T ) with a single uniformly distributed covariate X ∼ U ([a, b]), a < b ∈ R, potential outcomes as in (28), and main outcome, treatment effect, and propensity following a change point model as in (27), i.e., µ(x) = α0 + α1 · 1x≤γ ,

τ (x) = β0 + β1 · 1x≤γ ,

p(x) = q0 + q1 · 1x≤γ .

Assume that Assumption A3 holds. Then µ(x) and τ (x) can be decomposed as µ(x) = µB (x) + µH (x)

and

τ (x) = τ B (x) + τ H (x),

such that, for suitable constants α0′ , α1′ , α0′′ , α1′′ , β0′ , β0′′ , β1′′ , the following holds: 1. (Change point model with constant treatment effect) µB (x) = α0′ + α1′ · 1x≤γ and τ B (x) = β0′ . 2. (Change point model with unbiased difference-in-means) µH (x) = α0′′ + α1′′ · 1x≤γ and τ H (x) = β0′′ + β1′′ · 1x≤γ , where       E τ H (X) = E µH (X) + 21 τ H (X) | T = 1 −E µH (X) − 21 τ H (X) | T = 0 . (30) The proof is given in the appendix. Remark 2. Note that this decomposition is not unique. The condition (30) does not restrict α0′ and α0′′ as well as β0′ and β0′′ , so that α0 and β0 can be freely distributed onto the respective pair of parameters. If the treatment assignment is randomized, i.e., p(x) constant in [0, 1], any parameters will fulfill (30), so the same freedom of choice also extends to α1 .

18

Figure 2: Illustration of the decomposition of a change point model into a component with constant treatment effect and a component with unbiased difference-in-means, as in Theorem 5. The original model was defined by µ(x) = 0+2·1x≤0.6 , τ (x) = 1 + 3 · 1x≤0.6 and p(x) = 0.25 + 0.5 · 1x≤0.6 . Intuitively, this decomposition isolates exactly the two reasons a split can improve the estimated treatment effect (cf. Section 4.2): the bias component (µB , τ B ) has, by construction, a constant treatment effect, so any change in the estimated treatment effect attributable to this component must be due to bias correction rather than genuine heterogeneity. Conversely, the heterogeneity component (µH , τ H ) satisfies the unbiasedness condition (30) by construction, so any change in its estimated treatment effect must reflect genuine heterogeneity rather than bias correction. We make this precise in Section 4.4 below. Figure 2 illustrates such a decomposition as described in Theorem 5.

4.4. Identifying Heterogeneity and Bias with Signed Splitting Criteria By Theorem 5 and the linearity of expectation, it follows directly that the population splitting criteria Het± (k) and Bias± (k) in (25) and (26) can be decomposed accordingly.

19

To make this explicit, we first rewrite Het± (k) and Bias± (k) in terms of the treatment effect function τ (x) and the main outcome function µ(x) as follows: h  p  k(1 − k) · E µ(X) + 21 τ (X) | X1 ≤ k, X ∈ R(e), T = 1   − E µ(X) − 12 τ (X) | X1 ≤ k, X ∈ R(e), T = 0   − E µ(X) + 21 τ (X) | X1 > k, X ∈ R(e), T = 1  i − E µ(X) − 12 τ (X) | X1 > k, X ∈ R(e), T = 0 .

(31)

  Bias± (k) = E µ(X) + 21 τ (X) | X ∈ R(e), T = 1   − E µ(X) − 12 τ (X) | X ∈ R(e), T = 0 h   − k · E µ(X) + 12 τ (X) | X1 ≤ k, X ∈ R(e), T = 1  i − E µ(X) − 12 τ (X) | X1 ≤ k, X ∈ R(e), T = 0 h   − (1 − k) · E µ(X) + 21 τ (X) | X1 > k, X ∈ R(e), T = 1  i − E µ(X) − 12 τ (X) | X1 > k, X ∈ R(e), T = 0 .

(32)

Het± (k) =

and

Consequently, under the assumptions of Theorem 5, we obtain the decomposition ± Het± (k) = Het± B (k) + HetH (k)

and

± Bias± (k) = Bias± B (k) + BiasH (k),

(33)

B B where Het± B (k) is defined analogously to (31), with µ and τ replaced by µ and τ ± ± ± as in Theorem 5, and where HetH (k), BiasB (k), and BiasH (k) are defined in complete analogy. Note that, while there is some freedom in the choice of decomposition in ± Theorem 5, any such decomposition results in the same values for Het± B (k), HetH (k), ± ± BiasB (k), and BiasH (k). An illustrative example of this decomposition for both the bias and heterogeneity splitting criteria is shown in Figure 3.

Theorem 6. Consider a potential outcome model (Y T =0 , Y T =1 , X, T ) with a single uniformly distributed covariate X ∼ U ([a, b]), a < b ∈ R, potential outcomes as in (28), and main outcome, treatment effect, and propensity following a change point model as in (27), i.e., µ(x) = α0 + α1 · 1x≤γ ,

τ (x) = β0 + β1 · 1x≤γ ,

p(x) = q0 + q1 · 1x≤γ .

± ± Fix a decomposition as in Theorem 5, and define Het± B (k), HetH (k), BiasB (k), and ± BiasH (k) as in (33). Then, under Assumption A3, the following statements hold:

1. If the difference-in-means estimate is biased as in (29), i.e.,     E [τ (X)] ̸= E Y T =1 | T = 1 − E Y T =0 | T = 0 , ± then Bias± B (k) attains a global extremum at k = γ and HetB (γ) = 0.

20

Figure 3: The signed heterogeneity and bias criteria, Het± (k) and Bias± (k), together ± ± ± with their decomposition into Het± B (k) + HetH (k) and BiasB (k) + BiasH (k) as in (33), for the model shown in Figure 2. 2. If the treatment effect τ (x) is non-constant, then Het± H (k) attains a local extremum at k = γ and Bias± (γ) = 0. H The proof is given in the appendix. Remark 3. By the relationship between signed and unsigned splitting criteria, as well as the details of the proof of Theorem 6, the extrema in Theorem 6 of the signed splitting criteria correspond to maxima of their unsigned counterparts while zeros are preserved. As the splitting criteria are non-negative, the zeros correspond to minima of the splitting criteria. The interpretation of Theorem 6 is as follows. Theorem 5 shows that, locally at any fixed node, the potential outcome model marginalized with respect to a given splitting variable can be decomposed into two distinct components: one component that captures all heterogeneity in the treatment effect, and another component that captures

21

all bias of the local difference-in-means estimator. This decomposition is intrinsic to the marginalization mechanism of decision trees and holds at the population level for each candidate split. As a consequence, a split at a given node may improve the local estimation of heterogeneous treatment effects for two fundamentally different reasons: either because it captures genuine heterogeneity in the treatment effect function, or because it reduces bias in the local average treatment effect induced by confounding. Standard decision tree–based approaches do not distinguish between these two sources of improvement, as they rely only on the heterogeneity criterion when selecting splits. Theorem 6 shows that this separation can be achieved at the population level by the proposed splitting criteria. More precisely, the heterogeneity component of the signed heterogeneity criterion, Het± H (k), attains an extremum at the true jump location γ whenever the treatment effect is non-constant, while the bias component of the signed bias criterion, Bias± B (k) is maximized at the same location whenever bias is present. The behavior predicted by Theorem 6 can also be observed empirically in the illustrative example shown in Figure 3. Thus, the population versions of our splitting criteria Het± (k) and Bias± (k) in (25) and (26) recover the split positions associated with treatment effect heterogeneity and bias, respectively, within the appropriate model components, with the former guaranteed up to a local optimum. Overall, this result provides a theoretical explanation for the empirical improvements observed for the proposed splitting criteria in our simulation studies, presented in the next section.

5. Simulations and Application In this section, we apply IntCF on semi-synthetic and synthetic data sets, which allow for evaluation of prediction accuracy by having known ground-truth causal effects, as well as on a real data example. For synthetic data sets we are further able to compare importance measures to the relation of features to outcomes defined by the data model. Throughout, we pay particular attention to the trade-off between prediction accuracy, model simplicity, and interpretability, since this trade-off is the central motivation for our approach.

5.1. Prediction Accuracy on Real-data Benchmarks First, we compare prediction accuracy of our method with other established tree-based treatment effect estimation methods on four established data sets using the framework of Balazadeh et al. [2025]. As tree-based competitors we consider generalized random forests (GRF) [Athey et al., 2019], CausalForestDML as implemented in the EconML package [Battocchi et al., 2019], and causal forest [Wager and Athey, 2018]. These models were employed to gain predictions for average treatment effect (ATE) and CATE on four data sets with ground-truth causal effects: IHDP [Ramey et al., 1992, Hill, 2011], ACIC 2016 [Dorie et al., 2019], Lalonde CPS and PSID cohorts [LaLonde, 1986] with causal effects from RealCause [Neal et al., 2021]. The predictions were evaluated against ground-truth ATE and CATE using relative ATE error and precision in estimation of heterogeneous treatment effects (PEHE) [see Balazadeh et al., 2025].

22

Method

Mean ATE Relative Error ± Standard Error (↓ better)

Mean PEHE ± Standard Error (↓ better) IHDP

ACIC 2016 Lalonde CPS (×103)

Lalonde PSID (×103)

IHDP

ACIC 2016 Lalonde CPS

Lalonde PSID

IntCF

2.99±0.49

1.57±0.10

9.99±0.06 15.93±0.27

0.21±0.04

0.15±0.04

0.39±0.02

0.18±0.02

causal forest

3.72±0.61

2.19±0.12

9.58±0.03 16.02±0.17

0.22±0.04

0.28±0.06

0.24±0.01

0.09±0.01

0.18±0.03 0.07±0.02 0.08±0.01 0.05±0.01

0.82±0.02 1.03±0.01

0.85±0.02 1.05±0.01

GRF 3.67±0.61 1.32±0.30 12.33±0.06 Forest DML 4.53±0.73 1.48±0.31 12.95±0.04

22.91±0.17 22.99±0.15

Table 2: CATE & ATE results. Columns correspond to benchmark suites: IHDP, ACIC 2016, Lalonde CPS/PSID. (left half ) mean PEHE, (right half ) mean ATE relative error. Lalonde PEHE is in thousands. The best and second best entries in each column are highlighted. This table is modified from Balazadeh et al. [2025], with results for IntCF in the first line, results for causal forest [Wager and Athey, 2018] in the second line and the results for GRF and Forest DML of Balazadeh et al. [2025] as comparison in the remaining two lines. Results are shown in Table 2. Overall, IntCF is mostly comparable with causal forest, with some improvement on data sets IHDP and ACIC 2016, while causal forest produces better ATE estimates on the Lalonde data sets. The results for the other two methods follow a different pattern: they perform better in estimating ATE for data sets IHDP and ACIC 2016, but their results for PEHE and ATE on the Lalonde data sets are significantly worse than those of IntCF. Overall, IntCF performs best in two settings, is second best in three settings, and third out of four in the remaining three settings, so that we conclude that IntCF is very competitive in terms of prediction accuracy compared to other established tree-based competitors. Beyond the tree-based competitors considered here, [Balazadeh et al., 2025] also benchmark several additional, non-tree-based methods on the same four data sets, using the identical evaluation setup we adopt here; we refer to Balazadeh et al. [2025] for the full comparison. Some of these methods, most notably CausalPFN, achieve higher prediction accuracy than any of the tree-based methods considered in Table 2, including IntCF, in most settings. However, CausalPFN and related methods are complete black boxes: unlike tree-based approaches, they offer no direct way to inspect which features drive a given prediction, or to what extent, and hence cannot answer the type of interpretability questions that motivate this paper. Tree-based methods therefore occupy a useful middle ground between prediction accuracy and interpretability. Within this class of methods, IntCF is particularly attractive: as shown above, it is competitive in prediction accuracy with the best tree-based alternatives, while, as we show in the remainder of this section, achieving substantially better interpretability than all of them.

5.2. Interpretation Accuracy on Synthetic Data While the data sets in Section 5.1 allow for evaluation of prediction accuracy by providing ground-truth causal effects, they do not lend themselves to evaluation of interpretability of models, as the mechanisms behind these effects are not known. For this reason, we generated synthetic data to compare importance measures to the structure of treatment

23

effects as defined by the data model. Four data models were used to generate the synthetic data sets for this section, each with a complexity parameter s. This parameter scaled the dimension of the data, so that features for all models are generated uniformly from [0, 1]d with d = 5s, and the parameter also affected the functions µ, τ, p defining the models. These models were defined as follows: 1. Interaction model: An adaption of a, so-called, locally-spiky-sparse (LSS) model, based on [Basu et al., 2018, Behr et al., 2022], with interactions for the treatment effect, but constant baseline and randomized treatment assignment: p(X) = 0.5, µ(X) = 0, s X τ (X) = 1X2i−1 ≤γ · 1X2i ≤γ , i=1

where γ = 0.7. The features responsible for treatment effect heterogeneity are therefore X1 , . . . , X2s , appearing in interacting pairs (X2i−1 , X2i ); since p and µ do not depend on X at all, this model contains no confounders. Note that the marginal effect (see Section 4.2) of any single feature in this LSS model again corresponds to a change point model, directly connecting this simulation setting to our theoretical analysis in Section 4. 2. Change point model: An additive change point model with both confounding features and features responsible for treatment effect heterogeneity: 2s

p(X) = 0.2 +

0.6 X 1Xi ≤γ , s i=s+1

µ(X) =

τ (X) =

2s X

1Xi ≤γ ,

i=s+1 s X

1Xi ≤γ ,

i=1

this time with γ = 0.5. The features responsible for treatment effect heterogeneity are therefore X1 , . . . , Xs , while Xs+1 , . . . , X2s act as pure confounders, affecting both the propensity and the main outcome but not the treatment effect itself. 3. Linear model: A model with linear functions for propensity, main effect and

24

treatment effect: 2s

p(X) = 0.25 +

0.5 X Xi , s i=s+1

µ(X) =

τ (X) =

2s X

Xi ,

i=s+1 s X

Xi .

i=1

As in the change point model, the features responsible for treatment effect heterogeneity are X1 , . . . , Xs , while Xs+1 , . . . , X2s act as pure confounders. 4. Mixed model: Combination of linear and LSS model: 3s 0.6 X p(X) = 0.2 + 1Xi ≤γ , s i=2s+1

µ(X) =

3s X

Xi ,

i=2s+1

τ (X) =

s X

2s

1X2i−1 ≤γ · 1X2i ≤γ +

1 X Xi , s i=s+1

i=1

with γ = 0.7. The features responsible for treatment effect heterogeneity are therefore X1 , . . . , X2s , with Xs+1 , . . . , X2s entering both the interaction and the linear part of τ ; X2s+1 , . . . , X3s act as pure confounders. In all four models, any remaining features beyond those listed above, up to the total dimension d = 5s, are pure noise, unrelated to p, µ, or τ . The ROC-AUC statistic which we introduced below treats a feature as a positive only if it appears in the formula for τ ; confounders and noise features are therefore both treated as negatives, even though confounders do affect p and µ. Data was generated with normal distributed noise terms where the variance was equal to the variance of treatment effects. Again, we compare our method IntCF to the three other tree-based causal machine learning methods, causal forest [Wager and Athey, 2018], GRF [Athey et al., 2019] and CausalForestDML from EconML [Battocchi et al., 2019]. For the evaluation, both predictions and feature importance generated by these models were considered. Both GRF and CausalForestDML use a maximum depth and decay exponent for their feature importance calculation. We set maximum depth to infinity (or a very high value if not possible) and decay exponent to zero to ensure that the importance measures are comparable with those of the other estimation methods. For predictions, the precision in estimation of heterogeneous treatment effects (PEHE) was calculated and normalized by dividing it by the PEHE of a dummy predictor, which predicts a constant treatment effect by the difference-in-means from the training samples. For the feature importance,

25

Figure 4: Comparison of precision in estimation of heterogeneous effects (PEHE) of IntCF (blue), Forest DML (yellow), causal forest (green) and GRF (red), normalized to PEHE of a constant difference-in-means estimator. For each method, a band of one standard error is around the respective mean. In this figure, lower values are better. Top left shows interaction model, top right change point model, bottom left linear model, and bottom right mixed model. we considered a ROC-AUC statistic which measures how good the feature importance can distinguish between features which actually appear in the respective formula for treatment effect τ and those which do not. Each of the models with parameter s from 1 to 10 were used for simulations. For each model and parameter, 500 · s training samples and 100 000 separate evaluation samples were generated. The training samples were then used to train the machine learning models, which were subsequently evaluated on the evaluation samples. For each combination of data model, parameter s, and machine learning method, these steps were repeated 20 times. The results of these simulations for precision in estimation are shown in Figure 4. In most considered data models IntCF achieves the best prediction results, especially for the Interaction and Mixed model. For the Change point model, Forest DML and GRF, both of which include a double machine learning/orthogonalization step, are the best performing methods. Between the two methods that do not use this additional step,

26

Figure 5: ROC-AUC for feature importance as signal feature classifier for IntCF (blue), Forest DML (yellow), causal forest (green) and GRF (red). For each method, a band of one standard error is around the respective mean. In this figure, higher values are better. Top left shows interaction model, top right change point model, bottom left linear model, and bottom right mixed model. IntCF has a better precision than causal forest in all cases except one (Interaction model with s = 1). Simulation results regarding feature importance are shown in Figure 5. In all considered combinations of data model and complexity parameter s, the feature importance of IntCF results in a ROC-AUC of almost 1, while each other method achieves smaller values in most situations. This reflects a substantial improvement in interpretability, as IntCF is best at identifying the features actually relevant for treatment effect heterogeneity. This improvement is particularly notable given the relative simplicity of IntCF. As discussed in Section 3, IntCF retains the same structure as the original RF: predictions are obtained directly as the average of individual decision trees, with no additional correction step required, and feature importance can be read off directly from the tree structure. Causal forest shares this simple structure, but does not distinguish heterogeneity from bias when selecting splits, and consequently performs worse in our feature importance comparison. GRF and CausalForestDML, in contrast, rely on a substantially more involved procedure: both incorporate an additional orthogonalization (double machine learning) step, which

27

type

Heterogeneity score Bias score

0.5

Importance score

0.4

0.3

0.2

0.1

race

sex

exercise

active

education

smokeintensity

age

smokeyrs

wt71

0.0

Feature

Figure 6: Importance scores for heterogeneity (blue) and bias (orange) produced by IntCF for NHEFS data set. Features are sorted by heterogeneity score. For the bias scores, negative values where truncated at 0. alters the prediction mechanism and means that predictions can no longer be read off directly from the individual trees, unlike for IntCF and causal forest. Despite this added complexity, GRF and CausalForestDML do not achieve better feature importance scores than IntCF in our simulations. IntCF thus achieves the strongest interpretability results of all four methods while remaining, by construction, the structurally simplest.

5.3. Application on NHEFS Data set Finally, we demonstrate the application of our algorithm on a real-world data set: the National Health and Nutrition Examination Survey I Epidemiologic Followup Study (NHEFS)2 is a longitudinal study of a cohort first examined in 1971–75. This data set was also used previously to exemplify the handling of confounders and treatment effect heterogeneity [Hernan and Robins, 2020]. Like Hernan and Robins [2020], we consider data from 1566 smokers to estimate the effect of quitting smoking on weight change up to the follow-up examination in 1982–84, using the same 9 baseline covariates as Hernan and Robins [2020], including, e.g., the weight at the initial examination (wt71), the number of years of smoking (smokeyrs), and age (age). Notably, our method does not require these covariates to be labeled a priori as confounders or effect modifiers; instead, it attributes each split to bias correction or heterogeneity directly from the fitted tree. For the average treatment effect, we obtain an estimated weight gain of 3.2 kg due to quitting smoking, slightly lower than the 3.4 kg reported by Hernan and Robins [2020]. 2

More information about this study can be found at https://wwwn.cdc.gov/nchs/nhanes/nhefs/.

28

Additionally, we evaluate feature importance scores for the contribution to heterogeneity and bias of the predictions, shown in Figure 6. As can be seen in Figure 6, the feature with the highest heterogeneity importance is wt71, the weight at the initial examination, consistent with other methods that identify this feature as having the strongest influence on heterogeneity [see, e.g., Bénard and Josse, 2025]. In addition, our bias importance score identifies age as the most relevant confounder.

6. Discussion 6.1. Summary In this paper, we introduced the interpretable causal forest (IntCF), a modification of RF for estimating individual treatment effects. The only change relative to standard RF is the splitting criterion: by combining a heterogeneity criterion with a novel bias criterion, each split can be attributed to either genuine heterogeneity in the treatment effect or a correction for confounding bias. Together with a validation step that honestly re-evaluates these contributions at each node (Section 3.2), this enables direct interpretation of the resulting tree structure, for example through standard feature importance measures. We supported this splitting criterion theoretically by analyzing its population limit under a change point model tailored to causal inference, showing that it correctly attributes a split to bias correction precisely when the naive difference-in-means estimator is biased, and to heterogeneity precisely when the treatment effect is genuinely non-constant. In simulations, IntCF achieved prediction accuracy competitive with established treebased methods, while substantially outperforming all of them in recovering the features responsible for treatment effect heterogeneity, especially in the presence of confounders.

6.2. Simplicity as a Feature, not a Limitation A recurring theme of this paper is that IntCF requires no machinery beyond a modified splitting criterion: predictions are still obtained directly as the average of individual decision trees, exactly as in the original RF. This stands in contrast to generalized random forest and CausalForestDML, both of which rely on an additional orthogonalization (double machine learning) step that changes the prediction mechanism and severs the direct correspondence between the fitted trees and the resulting predictions. Causal forest shares IntCF’s simple structure, but does not distinguish heterogeneity from bias at the split level, and correspondingly falls short of IntCF’s interpretability in our simulations. That IntCF exceeds the interpretability of substantially more complex competitors, while remaining competitive with them in prediction accuracy, is in our view the central practical contribution of this paper: interpretability need not come at the cost of added model complexity.

29

6.3. Limitations and Future Work Several open questions remain. On the theoretical side, our analysis of the splitting criteria is currently limited to the population level; extending these results to finite samples, to more general data-generating processes beyond the change point model, and to the validation step from Section 3.2 are natural next steps. Likewise, our analysis of marginalization effects in Section 4.2 considered only a single confounder at a time; understanding how these effects interact in the presence of multiple, simultaneously confounding features warrants further investigation. Another open question is the consistency of the resulting IntCF estimator, which we did not establish in this paper. We expect that an adaptation of the consistency proof for causal forest from Wager and Athey [2018] should be feasible. Such a proof typically requires the conditional mean function E [Y | X = x] to be Lipschitz continuous, an assumption violated by the discontinuous change point model used in our theoretical analysis. However, as tree depth increases, the fraction of samples affected by any such discontinuity should vanish, suggesting that this is a technical rather than a fundamental obstacle.

Code Availability An implementation of IntCF, together with code to reproduce all analyses and figures in this paper, is publicly available at https://github.com/behr-group/intcf.

Acknowledgments The project was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation), project number 509149993, TRR 374.

References Abhineet Agarwal, Ana M. Kenney, Yan Shuo Tan, Tiffany M. Tang, and Bin Yu. Integrating random forests and generalized linear models for improved accuracy and interpretability, 2025. URL https://arxiv.org/abs/2307.01932. Joshua D. Angrist and Jörn-Steffen Pischke. Mastering ’metrics: the path from cause to effect. Princeton University Press, 2015. Susan Athey and Guido Imbens. Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27):7353–7360, 2016. doi: 10.1073/pnas.1510489113. Susan Athey, Julie Tibshirani, and Stefan Wager. Generalized random forests. The Annals of Statistics, 47(2):1148–1178, 2019. doi: 10.1214/18-AOS1709.

30

Vahid Balazadeh, Hamidreza Kamkari, Valentin Thomas, Benson Li, Junwei Ma, Jesse C. Cresswell, and Rahul G. Krishnan. CausalPFN: Amortized causal effect estimation via in-context learning. In Advances in Neural Information Processing Systems, volume 38, 2025. Sumanta Basu, Karl Kumbier, James B. Brown, and Bin Yu. Iterative random forests to discover predictive and stable high-order interactions. Proceedings of the National Academy of Sciences, 115(8):1943–1948, 2018. doi: 10.1073/pnas.1711236115. Keith Battocchi, Eleanor Dillon, Maggie Hei, Greg Lewis, Paul Oka, Miruna Oprescu, and Vasilis Syrgkanis. EconML: A Python Package for ML-Based Heterogeneous Treatment Effects Estimation. https://github.com/py-why/EconML, 2019. Version 0.x. Merle Behr, Yu Wang, Xiao Li, and Bin Yu. Provable boolean interaction recovery from tree ensemble obtained via random forests. Proceedings of the National Academy of Sciences, 119(22):e2118636119, 2022. doi: 10.1073/pnas.2118636119. Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001. doi: 10.1023/A: 1010933404324. Leo Breiman, Jerome H. Friedman, Richard A. Olshen, and Charles J. Stone. Classification and Regression Trees. Chapman and Hall/CRC, 1984. doi: 10.1201/9781315139470. Clément Bénard and Julie Josse. Variable importance for causal forests: breaking down the heterogeneity of treatment effects. Journal of Causal Inference, 13(1):20230062, 2025. doi: 10.1515/jci-2023-0062. Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21(1):C1–C68, 2018. doi: 10.1111/ectj.12097. Alicia Curth and Mihaela van der Schaar. Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 130, pages 1810–1818, 2021. Vincent Dorie, Jennifer Hill, Uri Shalit, Marc Scott, and Dan Cervone. Automated versus do-it-yourself methods for causal inference: Lessons learned from a data analysis competition. Statistical Science, 34(1):43–68, 2019. doi: 10.1214/18-STS667. Miguel A. Hernan and James M. Robins. Causal Inference: What If. Chapman & Hall/CRC, 2020. Jennifer L. Hill. Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics, 20(1):217–240, 2011. doi: 10.1198/jcgs.2010. 08162.

31

Guido W. Imbens and Donald B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, 2015. doi: 10. 1017/CBO9781139025751. Hao Jiang, Peng Qi, Jingying Zhou, Jack Zhou, and Sharath Rao. A short survey on forest based heterogeneous treatment effect estimation methods: Meta-learners and specific models. In 2021 IEEE International Conference on Big Data (Big Data), pages 3006–3012, 2021. doi: 10.1109/BigData52589.2021.9671439. Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel, and Bin Yu. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the National Academy of Sciences, 116(10):4156–4165, 2019. doi: 10.1073/pnas.1804597116. Robert J. LaLonde. Evaluating the econometric evaluations of training programs with experimental data. The American Economic Review, 76(4):604–620, 1986. Xiao Li, Yu Wang, Sumanta Basu, Karl Kumbier, and Bin Yu. A debiased MDI feature importance measure for random forests. In Advances in Neural Information Processing Systems, volume 32. 2019. Scott M. Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M. Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1):56–67, 2020. doi: 10.1038/s42256-019-0138-9. Charles F. Manski. Identification problems in the social sciences. Sociological Methodology, 23:1–56, 1993. doi: 10.2307/271005. Brady Neal, Chin-Wei Huang, and Sunand Raghupathi. RealCause: Realistic causal inference benchmarking, 2021. URL https://arxiv.org/abs/2011.15007. Miruna Oprescu, Vasilis Syrgkanis, and Zhiwei Steven Wu. Orthogonal random forest for causal inference. In Proceedings of the 36th International Conference on Machine Learning, volume 97, pages 4932–4941, 2019. Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12(85):2825–2830, 2011. Craig T. Ramey, Donna M. Bryant, Barbara H. Wasik, Joseph J. Sparling, Kaye H. Fendt, and Lisa M. La Vange. Infant health and development program for low birth weight, premature infants: program elements, family participation, and child intelligence. Pediatrics, 89(3):454–465, 1992. doi: 10.1542/peds.89.3.454. Paul R. Rosenbaum and Donald B. Rubin. The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1):41–55, 1983. doi: 10.1093/ biomet/70.1.41.

32

Donald B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5):688–701, 1974. doi: 10.1037/h0037350. Erwan Scornet and Giles Hooker. Theory of random forests. Annual Review of Statistics and Its Application, 13:99–121, 2026. doi: 10.1146/annurev-statistics-112723-034707. Uri Shalit, Fredrik D. Johansson, and David Sontag. Estimating individual treatment effect: generalization bounds and algorithms. In Proceedings of the 34th International Conference on Machine Learning, volume 70, pages 3076–3085, 2017. Claudia Shi, David Blei, and Victor Veitch. Adapting neural networks for the estimation of treatment effects. In Advances in Neural Information Processing Systems, volume 32. 2019. Assaf Shmuel, Oren Glickman, and Teddy Lazebnik. A comprehensive benchmark of machine and deep learning models on structured data for regression and classification. Neurocomputing, 655:131337, 2025. doi: 10.1016/j.neucom.2025.131337. Stefan Wager and Susan Athey. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113(523): 1228–1242, 2018. doi: 10.1080/01621459.2017.1319839. Zhengze Zhou and Giles Hooker. Unbiased measurement of feature importance in treebased methods. ACM Transactions on Knowledge Discovery from Data (TKDD), 15 (2):26:1–26:21, 2021. doi: 10.1145/3429445.

33

A. Proofs A.1. Proofs of Section 3 Proof of Lemma 1. In the following computation, the property Jparent = JL ⊔ JR is used in the second equation and Assumption (12) is used in the third equation:   X X 1 1 X  (τ̂parent − τi )2 − (τ̂L − τi )2 + (τ̂R − τi )2  N N i∈Jparent i∈JL i∈JR   X X 1  X 2 = τ̂parent − 2τ̂parent τi + τi2  N i∈Jparent i∈Jparent i∈Jparent     X X X X X X 1  − τ̂L2 − 2τ̂L τi + τi2  +  τ̂R2 − 2τ̂R τi + τi2  N i∈JL i∈JL i∈JL i∈JR i∈JR i∈JR     X X X 1  2 τi2  τi  + τi + = (nL + nR )τ̂parent − 2τ̂parent  N i∈Jparent i∈JR i∈JL   X X X 1  − τi2  nL τ̂L2 − 2τ̂L τi + nR τ̂R2 − 2τ̂R τi + N i∈JL i∈JR i∈Jparent   X 1  2 = (nL + nR )τ̂parent − 2τ̂parent (nL τL + nR τR ) + τi2  N i∈Jparent   X 1  − nL τ̂L2 − 2nL τ̂L2 + nR τ̂R2 − 2nR τ̂R2 + τi2  N i∈Jparent

 1 2 (nL + nR )τ̂parent − 2τ̂parent (nL τL + nR τR ) + nL τ̂L2 + nR τ̂R2 N 1 = (nL (τ̂parent − τ̂L )2 + nR (τ̂parent − τ̂R )2 ) N =

34

Proof of Lemma 2. By using assumption (13) direct computation yields: 1 (nL (τ̂parent − τ̂L )2 + nR (τ̂parent − τ̂R )2 ) N  2  2 ! nL τ̂L + nR τ̂R nL τ̂L + nR τ̂R 1 nL − τ̂L + nR − τ̂R = N nL + n R nL + n R     ! −nR τ̂L + nR τ̂R 2 nL τ̂L − nL τ̂R 2 1 nL + nR = N nL + n R nL + n R   n2L · nR nL · n2R 1 2 2 (−τ̂L + τ̂R ) + (τ̂L − τ̂R ) = N (nL + nR )2 (nL + nR )2   nL · n R nR nL = (τ̂L − τ̂R )2 + (τ̂L − τ̂R )2 N (nL + nR ) nL + nR nL + n R nL · n R = (τ̂L − τ̂R )2 N (nL + nR ) Proof of Lemma 3. By direct computation: 1 nL · n R (nL (τ̂parent − τ̂L )2 + nR (τ̂parent − τ̂R )2 ) − (τ̂L − τ̂R )2 N N (nL + nR )   1 nL · n R 2 2 2 = nL (τ̂parent − τ̂L ) + nR (τ̂parent − τ̂R ) − (τ̂L − τ̂R ) N nL + n R  1 2 2 = nL (τ̂parent − 2τ̂parent τ̂L + τ̂L2 ) + nR (τ̂parent − 2τ̂parent τ̂R + τ̂R2 ) N  nL · n R 2 2 (τ̂ − 2τ̂L τ̂R + τ̂R ) − nL + n R L  1 2 = (nL + nR )τ̂parent − 2τ̂parent (nL τ̂L + nR τ̂R ) N  n2L n2R nL · n R 2 2 + τ̂ + 2 τ̂L τ̂R + τ̂ nL + n R L nL + n R nL + n R R  2 ! n τ̂ + n τ̂ nL + n R n τ̂ + n τ̂ L L R R L L R R 2 = − 2τ̂ + τ̂parent N nL + n R nL + n R  2 nL τ̂L + nR τ̂R nL + n R = τ̂parent − N nL + n R

35

Proof of Lemma 4. At each node e we have τ̂eb = n1 2 2 X  X τ̂eb − τ̂ival τ̂ef − τ̂ival − i∈Ie

i∈Ie

=

val i∈Ie τ̂i , so

P

X

 X  (τ̂ef )2 − 2τ̂ef τ̂ival + (τ̂ival )2 − (τ̂eb )2 − 2τ̂eb τ̂ival + (τ̂ival )2

i∈Ie

i∈Ie









= n (τ̂ef )2 − (τ̂eb )2

 X + 2 τ̂eb − τ̂ef τ̂ival i∈Ie

  = n (τ̂ef )2 − (τ̂eb )2 + 2 τ̂eb − τ̂ef · n · τ̂eb   = n (τ̂ef )2 − (τ̂eb )2 + 2(τ̂eb )2 − 2τ̂ef τ̂eb   = n (τ̂ef )2 − 2τ̂ef τ̂eb + (τ̂eb )2  2 = n τ̂ef − τ̂eb , n τ̂ b +n τ̂ b

b and, using τ̂parent = LnLL +nRR R , we get 2 2 X  2 X  X  b τ̂Rb − τ̂ival τ̂Lb − τ̂ival − τ̂parent − τ̂ival −

=

X

i∈IR

i∈IL

i∈Iparent

b b (τ̂parent )2 − 2τ̂parent τ̂ival + (τ̂ival )2



i∈IR

X

−

  X (τ̂Rb )2 − 2τ̂Rb τ̂ival + (τ̂ival )2 (τ̂Lb )2 − 2τ̂Lb τ̂ival + (τ̂ival )2 −

i∈IR b 2 b = (nL + nR )(τ̂parent ) − 2(nL + nR )(τ̂parent )2 − nL (τ̂Lb )2 + 2nL (τ̂Lb )2 − nR (τ̂Rb )2 + 2nR (τ̂Rb )2 b = −(nL + nR )(τ̂parent )2 + nL (τ̂Lb )2 + nR (τ̂Rb )2 i∈IL

2 nL τ̂Lb + nR τ̂Rb nL + n R b 2 b (nL + nR )nL (τ̂L ) + (nL + nR )nR (τ̂R )2 − n2L (τ̂Lb )2 − 2nL nR τ̂Lb τ̂Rb − n2R (τ̂Rb )2 = nL + n R 2 nL · n R  b = τ̂ − τ̂Rb . nL + n R L

= nL (τ̂Lb )2 + nR (τ̂Rb )2 − (nL + nR )



Combining these results, we have 2 2 X  2 X  X  f τ̂Rf − τ̂ival = τ̂parent − τ̂ival − τ̂Lf − τ̂ival − i∈Iparent

i∈IL

i∈IR

2 n + n  2 n  2 n  2 nL · n R  b L R L R f b τ̂L − τ̂Rb + − τ̂parent − τ̂parent τ̂Lf − τ̂Lb − τ̂Rf − τ̂Rb . nL + n R N N N

36

A.2. Proof of Theorem 5 Remark 4. In the following statements and proofs, we will assume w.l.o.g. that the covariate X is uniformly distributed on [0, 1], also implying R(e) = [0, 1]. If its domain would be a different interval [a, b], replacing it by X−a b−a —as well as performing the same transformation for γ and k—would create an equivalent model with the same values for treatment effect estimates and (signed) splitting criteria. Lemma 7. For the univariate change-point model (27) with potential outcomes as in (28), assuming Assumption A3 holds, the bias of the difference-in-means estimate as defined in (29) is given by     E Y T =1 | T = 1 − E Y T =0 | T = 0 − E [τ (X)] = γ(1 − γ)q1

α1 + (1/2 − q0 − γq1 )β1 . (34) (q0 + γq1 )(1 − q0 − γq1 )

Proof. Because of Assumption A3, the conditional expectations are well defined and we have   (α1 − β1 /2)γ(1 − q0 − q1 ) E Y T =0 | T = 0 = α0 − β0 /2 + , 1 − q0 − γq1   (α1 + β1 /2)γ(q0 + q1 ) E Y T =1 | T = 1 = α0 + β0 /2 + q0 + γq1 and hence,     E Y T =1 | T = 1 − E Y T =0 | T = 0 (α1 + β1 /2)γ(q0 + q1 ) (α1 − β1 /2)γ(1 − q0 − q1 ) − =β0 + q0 + γq1 1 − q0 − γq1   (1 − γ)q1 (1 − γ)q1 + γβ1 1 + =β0 + γ(α1 − β1 /2) . (q0 + γq1 )(1 − q0 − γq1 ) q0 + γq1 Moreover, we have E [τ (X)] = β0 + γβ1 and hence,     E Y T =1 | T = 1 − E Y T =0 | T = 0 − E [τ (X)] γ(1 − γ)q1 γ(1 − γ)q1 =(α1 − β1 /2) + β1 (q0 + γq1 )(1 − q0 − γq1 ) q0 + γq1 α1 + (1/2 − q0 − γq1 )β1 =γ(1 − γ)q1 . (q0 + γq1 )(1 − q0 − γq1 ) Proof for Theorem 5. Define α0′ := α0 ,

α1′ := α1 − α1′′ ,

β0′ := β0 ,

α0′′ := 0,

α1′′ := −(1/2 − q0 − γq1 )β1 ,

β0′′ := 0,

β1′′ := β1 .

With these definitions, µ(x) = µB (x) + µH (x) and τ (x) = τ B (x) + τ H (x) follows directly.

37

For the second we difference between  H model,   Hconsider1 the  Hthe difference-in-means  H prediction E τ (X) = E µ (X) + 2 τ (X)| T = 1 − E µ (X) − 12 τ H (X) | T = 0 and the expectation of the treatment effect E τ H (X) . By Lemma 7 this is given by γ(1 − γ)q1

α1′′ + (1/2 − q0 − γq1 )β1′′ = 0, (q0 + γq1 )(1 − q0 − γq1 )

where for the evaluation we used α1′′ = −(1/2 − q0 − γq1 )β1 and β1′′ = β1 . So the difference-in-means estimator is in expectation equal to the treatment effect, i.e., is unbiased.

A.3. Proof of Theorem 6 Remark 5. We will continue assuming X ∼ U ([0, 1]), see Remark 4. Further we will assume that γ ∈ (0, 1) in (27). For values outside (0, 1), there is a model with β1 = q1 = α1 = 0 and γ ∈ (0, 1) which has almost surely the same propensities and expected outcomes, which therefore also result in the same expected predictions and (signed) splitting criteria. Lemma 8. For the univariate change-point model (27), assuming Assumption A3 holds, with potential outcomes as in (28) and a split point k ∈ (0, 1) define     τ̂L∞ := E Y T =1 | X ≤ k, T = 1 − E Y T =0 | X ≤ k, T = 0 , (35)     τ̂R∞ := E Y T =1 | X > k, T = 1 − E Y T =0 | X > k, T = 0 Then we have τ̂L∞ =

( β0 + β1

k ≤ γ,

1 β0 + (α1 − β1 /2) (kq0 +γqγ(k−γ)q + β1 1 )(k−kq0 −γq1 )



(k−γ)q0 1 − kq 0 +γq1



k>γ

and  (γ − k)(1 − γ)q1   β0 + (α1 − β1 /2)    ((1 − k)q0 + (γ − k)q1 )((1 − k)(1 − q0 ) − (γ − k)q1 )    (1 − γ)q0 τ̂R∞ = + β1 1 −    (1 − k)q0 + (γ − k)q1   β 0 Proof. For a splitting point k, we have (  T =0  α0 − β0 /2 + α1 − β1 /2 k ≤ γ, E Y | X ≤ k, T = 0 = (α1 −β1 /2)γ(1−q0 −q1 ) k > γ, α0 − β0 /2 + k−kq0 −γq1 (   α0 + β0 /2 + α1 + β1 /2 k ≤ γ, E Y T =1 | X ≤ k, T = 1 = (α1 +β1 /2)γ(q0 +q1 ) α0 + β0 /2 + k > γ, kq0 +γq1

38

k ≤ γ, k > γ.

and hence, τ̂L∞ =

( β0 + β1 1 β0 + (α1 − β1 /2) (kq0 +γqγ(k−γ)q + β1 1 )(k−kq0 −γq1 )



(k−γ)q0 1 − kq 0 +γq1

 k ≤ γ, k > γ.

Moreover, we have E Y T =0 | X > k, T = 0 = 



  E Y T =1 | X > k, T = 1 =

( −β1 /2)(γ−k)(1−q0 −q1 ) α0 − β0 /2 + (α1(1−k)(1−q 0 )−(γ−k)q1

k ≤ γ,

α0 − β0 /2 ( +β1 /2)(γ−k)(q0 +q1 ) α0 + β0 /2 + (α1(1−k)q 0 +(γ−k)q1

k > γ,

α0 + β0 /2

k ≤ γ, k > γ,

and hence,  (γ − k)(1 − γ)q1  β0 + (α1 − β1 /2)    ((1 − k)q0 + (γ − k)q1 )((1 − k)(1 − q0 ) − (γ − k)q1 )    ∞ (1 − γ)q0 τ̂R = + β1 1 −    (1 − k)q0 + (γ − k)q1   β 0

k ≤ γ, k > γ.

Note that all terms in the denominators are either treatment propensities for some subset of samples, scaled versions of these, or a product of multiple such propensities. By Assumption A3, these are positive, so all terms are defined. Proof of Theorem 6. Use the notation from (35) and     ∞ τ̂parent := E Y T =1 | T = 1 − E Y T =0 | T = 0   (1 − γ)q1 (1 − γ)q1 = β0 + γ(α1 − β1 /2) + γβ1 1 + , (q0 + γq1 )(1 − q0 − γq1 ) q0 + γq1 as calculated in Lemma 8. Note that we can write p Het± (k) = k(1 − k) · (τ̂L∞ − τ̂R∞ ), ∞ Bias± (k) = τ̂parent − kτ̂L∞ − (1 − k)τ̂R∞ . ∞ Further note, that α0 does not appear in any of τ̂parent , τ̂L∞ , τ̂R∞ , so it does not have any effect on Het± (k) or Bias± (k). Additionally, β0 appears in all three treatment effect estimates, but always as simple summand. So changing it does not have any effect ∞ on the differences τ̂L∞ − τ̂R∞ and τ̂parent − kτ̂L∞ − (1 − k)τ̂R∞ , which define Het± (k) and Bias± (k). To simplify notation, we will therefore, without loss of generality, set all of α0 , α0′ , α0′′ , β0 , β0′ , β0′′ zero for the following analysis of the signed splitting criteria. We proof the two statements (for bias and for heterogeneity) of the theorem separately.

39

Bias

After setting α0′ = β0′ = 0, as explained above, the bias model is given by µB (x) = α1′ · 1x≤γ , τ B (x) = 0, p(x) = q0 + q1 · 1x≤γ .

In this model, the estimated treatment effect tends to ∞ τ̂parent = γα1′

(1 − γ)q1 , (q0 + γq1 )(1 − q0 − γq1 )

while the true treatment effect is constant zero: τ B = 0. In the child nodes for a split at k, the estimates are ( 0 k ≤ γ, ∞ τ̂L = γ(k−γ)q1 ′ α1 (kq0 +γq1 )(k−kq0 −γq1 ) k > γ, ( (γ−k)(1−γ)q1 α1′ ((1−k)q0 +(γ−k)q ∞ 1 )((1−k)(1−q0 )−(γ−k)q1 ) τ̂R = 0

k ≤ γ, k > γ.

Therefore, the population signed splitting criteria are given by p k(1 − k)(τ̂L∞ − τ̂R∞ ), Het± B (k) = p   (γ−k)(1−γ)q1  k(1 − k) 0 − α′ k ≤ γ, 1 ((1−k)q0 +(γ−k)q1 )((1−k)(1−q0 )−(γ−k)q1 )   = p γ(k−γ)q1  k(1 − k) α′ k > γ, 1 (kq0 +γq1 )(k−kq0 −γq1 ) − 0 ( p (γ−k)(1−γ)q1 − k(1 − k)α1′ ((1−k)q0 +(γ−k)q k ≤ γ, 1 )((1−k)(1−q0 )−(γ−k)q1 ) = p (k−γ)γq1 ′ k(1 − k)α1 (kq0 +γq1 )(k−kq0 −γq1 ) k > γ, ∞ ∞ ∞ Bias± B (k) = τ̂parent − kτ̂L − (1 − k)τ̂R ( (1−γ)q

=

=

(γ−k)(1−γ)q

1 γα′1 (q +γq )(1−q1 −γq ) −k·0−(1−k)α′1 ((1−k)q +(γ−k)q )((1−k)(1−q k ≤ γ, 0 1 0 1 0 1 0 )−(γ−k)q1 ) (1−γ)q1 γ(k−γ)q1 ′ ′ γα1 (q0 +γq1 )(1−q0 −γq1 ) − kα1 (kq0 +γq1 )(k−kq0 −γq1 ) − (1 − k) · 0 k > γ, (   γ(1−γ)q (γ−k)(1−γ)q1 α′1 (q +γq )(1−q 1−γq ) −(1−k) ((1−k)q +(γ−k)q )((1−k)(1−q k ≤ γ, )−(γ−k)q )

α1′

0

1

0

1

0

1

γ(1−γ)q1 γ(k−γ)q1 (q0 +γq1 )(1−q0 −γq1 ) − k (kq0 +γq1 )(k−kq0 −γq1 )

0

1

k > γ,

± Note that Het± B (k) and BiasB (k) are both continuous and at k = γ we have

Het± B (γ) = 0 ′ Bias± B (γ) = α1

γ(1 − γ)q1 . (q0 + γq1 )(1 − q0 − γq1 )

40

Recall from Theorem 5 that µ(x) = µB (x) + µH (x) and τ (x) = τ B (x) + τ H (x) and hence,     E [τ (X)] = E τ B (X) + E τ H (X) ,       E µ(X)+ 21 τ (X) | T = 1 = E µB (X)+ 12 τ B (X) | T = 1 + E µH (X)+ 21 τ H (X) | T = 1 ,       E µ(X)− 21 τ (X) | T = 0 = E µB (X)− 12 τ B (X) | T = 0 + E µH (X)− 21 τ H (X) | T = 0 . From (30) it follows that     1 1 E µ(X) + τ (X) | T = 1 − E µ(X) − τ (X) | T = 0 − E [τ (X)] 2 2       1 1 = E µB (X) + τ B (X) | T = 1 − E µB (X) − τ B (X) | T = 0 − E τ B (X) . 2 2 By Lemma 7 the right-hand side of this equation is equal to Bias± B (γ). Because we assume that the difference-in-means estimator of the     original model is biased, i.e., E µ(X) + 12 τ (X) | T = 1 − E µ(X) − 12 τ (X) | T = 0 − E [τ (X)] = ̸ 0, it follows that ′ ̸= 0 and q ̸= 0. Moreover, as Bias± (0) = Bias± (1) = 0, it Bias± (γ) = ̸ 0 along with α 1 1 B B B follows that Bias± is non-constant and we will show that it has a global extremum at B k = γ. Considering the derivative of Bias± B for k > γ yields   d ′ γ(1 − γ)q1 γ(k − γ)q1 α −k dk 1 (q0 + γq1 )(1 − q0 − γq1 ) (kq0 + γq1 )(k − kq0 − γq1 )  (γ(k − γ)q1 + kγq1 )(kq0 + γq1 )(k − kq0 − γq1 ) = −α1′ (kq0 + γq1 )2 (k − kq0 − γq1 )2  kγ(k − γ)q1 (q0 (k − kq0 − γq1 ) + (kq0 + γq1 )(1 − q0 )) − (kq0 + γq1 )2 (k − kq0 − γq1 )2 (q 2 + 2q0 q1 − q0 − q1 )k 2 + 2γq12 k − γ 2 q12 . = −α1′ γ 2 q1 0 (kq0 + γq1 )2 (k − kq0 − γq1 )2 Note that Bias± B is contentiously differentiable in (γ, 1). Hence, for a potential extremum k of Bias± in (γ, 1), we would need B (q02 + 2q0 q1 − q0 − q1 )k 2 + 2γq12 k − γ 2 q12 (kq0 + γq1 )2 (k − kq0 − γq1 )2 ⇐⇒ 0 = (q02 + 2q0 q1 − q0 − q1 )k 2 + 2γq12 k − γ 2 q12 p −2γq12 ± (2γq12 )2 − 4(q02 + 2q0 q1 − q0 − q1 )(−γ 2 q12 ) ⇐⇒ k = , 2(q02 + 2q0 q1 − q0 − q1 ) 0 = −α1′ γ 2 q1

but the discriminant is (2γq12 )2 − 4(q02 + 2q0 q1 − q0 − q1 )(−γ 2 q12 ) = 4γ 2 q14 + 4γ 2 q02 q12 + 8γ 2 q0 q13 − 4γ 2 q0 q12 − 4γ 2 q13 = −4γ 2 q12 (q0 + q1 )(1 − q0 − q1 ) < 0.

41

Therefore, there are no local extrema for k ∈ (γ, 1). Similarly, there are also no local ± extrema within (0, γ). As Bias± B is continuous and non-constant with BiasB (k) = 0 at ± k = 0 and k = 1, BiasB (γ) must be a global extremum. Heterogeneity After setting α0′′ = β0′′ = 0, as explained above, the heterogeneity model is given by µH (x) = α1′′ · 1x≤γ , τ H (x) = β1′′ · 1x≤γ , p(x) = q0 + q1 · 1x≤γ , with α1′′ = −(1/2 − q0 − γq1 )β1′′ , such that the bias calculated in Lemma 7 is 0. For this model, the treatment effect obtained from the difference-in-means estimator ∞ tends to the correct average treatment effect, i.e., τ̂parent = E [τ (X)] = γβ1′′ and the estimates in the child nodes converge to ( ′′ β1  ∞  k ≤ γ, τ̂L = (kq0 +γq1 )(1−q0 )−(q0 +γq1 )γq1 ′′ β1 1 − (k − γ) (kq0 +γq1 )(k−kq0 −γq1 ) k > γ, ( (q0 +γq1 )(1−q0 −γq1 )−k(q0 +q1 )(1−q0 −q1 ) β1′′ (γ − k) ((1−k)q k ≤ γ, 0 +(γ−k)q1 )((1−k)(1−q0 )−(γ−k)q1 ) τ̂R∞ = 0 k > γ, which can be calculated by using α1′′ = −(1/2 − q0 − γq1 )β1′′ in the formulas given by Lemma 8. The population signed splitting criteria evaluate to p Het± k(1 − k)(τ̂L∞ − τ̂R∞ ), H (k) = p    k(1 − k) β ′′ − β ′′ (γ − k) (q0 +γq1 )(1−q0 −γq1 )−k(q0 +q1 )(1−q0 −q1 ) k ≤ γ, ((1−k)q0 +(γ−k)q1 )((1−k)(1−q0 )−(γ−k)q  1 1  1) = p  k(1 − k) β ′′ 1 − (k − γ) (kq0 +γq1 )(1−q0 )−(q0 +γq1 )γq1 − 0 k > γ, 1 (kq0 +γq1 )(k−kq0 −γq1 )  p   β ′′ k(1 − k) 1 − (γ − k) (q0 +γq1 )(1−q0 −γq1 )−k(q0 +q1 )(1−q0 −q1 ) k ≤ γ, 1 ((1−k)q0 +(γ−k)q1 )((1−k)(1−q0)−(γ−k)q1 )  = p β ′′ k(1 − k) 1 − (k − γ) (kq0 +γq1 )(1−q0 )−(q0 +γq1 )γq1 k > γ, 1 (kq0 +γq1 )(k−kq0 −γq1 ) ∞ ∞ ∞ Bias± H (k) = τ̂parent − kτ̂L − (1 − k)τ̂R  γβ ′′ − kβ ′′ − (1 − k)β ′′ (γ − k) (q0 +γq1 )(1−q0 −γq1 )−k(q0 +q1 )(1−q0 −q1 ) 1 1  1 ((1−k)q0 +(γ−k)q1 )((1−k)(1−q 0 )−(γ−k)q1 )  = (kq +γq )(1−q )−(q +γq )γq 0 1 0 0 1 1 ′′ ′′ γβ1 − kβ1 1 − (k − γ) (kq +γq )(k−kq −γq ) − (1 − k) · 0 0 1 0 1  k(1−γ)2 q12 −β ′′ (γ − k) k ≤ γ, 1 ((1−k)q0 +(γ−k)q1 )((1−k)(1−q0 )−(γ−k)q1 ) = (1−k)γ 2 q12 β ′′ (k − γ) k>γ 1

(kq0 +γq1 )(k−kq0 −γq1 )

± Note that Het± H (k) and BiasH (k) are both continuous and at k = γ we have p ′′ Het± (γ) = β γ(1 − γ), 1 H

Bias± H (γ) = 0.

42

k ≤ γ, k > γ,

If the treatment effect of the original model τ (x) is non-constant, it follows that 0 < γ < 1, ± ± β1 ̸= 0, and β1′′ ̸= 0. This implies Het± H (γ) ̸= 0 and, as HetH (0) = HetH (1) = 0, that ± ± HetH non-constant. We will now show that HetH (γ) is a local extremum. Considering the derivative of Het± H for k < γ yields p d ′′ Het± k(1 − k) H (k) = β1 dk  1 − 2k (q0 + γq1 )(1 − q0 − γq1 ) − k(q0 + q1 )(1 − q0 − q1 )  · + 2k(1 − k) ((1 − k)q0 + (γ − k)q1 )((1 − k)(1 − q0 ) − (γ − k)q1 ) p ′′ − β1 k(1 − k)(γ − k)  (q0 + q1 )(1 − q0 − q1 ) · − ((1 − k)q0 + (γ − k)q1 )((1 − k)(1 − q0 ) − (γ − k)q1 ) (q0 + γq1 )(1 − q0 − γq1 ) − k(q0 + q1 )(1 − q0 − q1 ) + (q0 + q1 ) ((1 − k)q0 + (γ − k)q1 )2 ((1 − k)(1 − q0 ) − (γ − k)q1 ) (q0 + γq1 )(1 − q0 − γq1 ) − k(q0 + q1 )(1 − q0 − q1 ) + (1 − q0 − q1 ) ((1 − k)q0 + (γ − k)q1 )((1 − k)(1 − q0 ) − (γ − k)q1 )2 1 − 2k (q0 + γq1 )(1 − q0 − γq1 ) − k(q0 + q1 )(1 − q0 − q1 )  + 2k(1 − k) ((1 − k)q0 + (γ − k)q1 )((1 − k)(1 − q0 ) − (γ − k)q1 ) and for k > γ   p d 1 − 2k (kq0 + γq1 )(1 − q0 ) − (q0 + γq1 )γq1 ± ′′ HetH (k) = β1 k(1 − k) − dk 2k(1 − k) (kq0 + γq1 )(k − kq0 − γq1 )  p q0 (1 − q0 ) + β1′′ k(1 − k)(k − γ) − (kq0 + γq1 )(k − kq0 − γq1 ) (kq0 + γq1 )(1 − q0 ) − (q0 + γq1 )γq1 + q0 (kq0 + γq1 )2 (k − kq0 − γq1 ) (kq0 + γq1 )(1 − q0 ) − (q0 + γq1 )γq1 + (1 − q0 ) (kq0 + γq1 )(k − kq0 − γq1 )2 1 − 2k (kq0 + γq1 )(1 − q0 ) − (q0 + γq1 )γq1  . − 2k(1 − k) (kq0 + γq1 )(k − kq0 − γq1 ) Note that Het± H (k) is continuously differentiable in (0, γ) and in (γ, 1). The limits of these derivatives at γ are p d q0 (1 − q0 ) + 2γ 2 q12 ′′ Het± (k) = β γ(1 − γ) lim 1 H 2γ(1 − γ)q0 (1 − q0 ) k→γ − dk p d (1 − q0 − q1 )(q0 + q1 ) + 2(1 − γ)2 q12 lim Het± (k) = −β1′′ γ(1 − γ) H 2γ(1 − γ)(q0 + q1 )(1 − q0 − q1 ) k→γ + dk All of γ, (1 − γ), q0 , (1 − q0 ), (q0 + q1 ) and (1 − q0 − q1 ) are assumed to be positive by Assumptions A3 and γ ∈ (0, 1), as well as q12 ≥ 0. Therefore, we have ± d d ′′ ′′ sgn(limk→γ − dk Het± H (k)) = sgn(β1 ) and sgn(limk→γ + dk HetH (k)) = − sgn(β1 ). Also ± ± ′′ note that sgn(HetH (γ)) = sgn(β1 ). Together, it follows that HetH (γ) has a local extremum at k = γ.

43

Record · ID 919414 · SHA-256 93f108b329fc0f7e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.