1
Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks
arXiv:2609.05073v1 [cs.LG] 4 Sep 2026
Abdessamed Qchohi, Jessica Moysen Cortes, and Matteo Zecchin
Abstract Conformal counterfactual inference enables network operators to use logged telemetry to reliably answer “what-if” questions about network operation. These answers typically take the form of prediction sets that contain, with a user-defined probability, the key performance indicators (KPIs) that would have been observed under alternative control actions. A key challenge is that logged telemetry may omit variables used by the controller, resulting in hidden confounding and invalidating the statistical guarantees of counterfactual analysis. In principle, this issue can be addressed using randomized telemetry, collected by assigning control actions independently of the network state. However, because such randomization may disrupt normal operation, randomized telemetry is typically scarce, causing counterfactual analysis based solely on it to produce uninformative prediction sets. To address these challenges, we propose ConfoundingValid Counterfactual Conformal Inference (CV-CCI), which combines abundant, potentially confounded observational telemetry with limited randomized data through the General Synthetic-Powered Inference (GESPI) principle. CV-CCI leverages observational data to improve efficiency while using randomized data to retain finite-sample coverage guarantees under arbitrary hidden confounding. Experiments on two representative radio access network (RAN) control tasks show that CV-CCI remains valid under hidden confounding while producing more efficient prediction sets than state-of-the-art confounding-valid baselines.
Index Terms Radio access networks, telemetry, prediction methods, inferential statistics, uncertainty Abdessamed Qchohi and Matteo Zecchin are with the Communication Systems Department, EURECOM, 06904 Sophia Antipolis, France (e-mail: {abdessamed.qchohi, matteo.zecchin}@eurecom.fr). Jessica Moysen Cortes is with Huawei Technologies Sweden AB, Sweden (e-mail: [email protected]).
2
I. I NTRODUCTION A. Motivation Modern radio access networks (RANs), including those based on Open RAN (O-RAN) architectures, increasingly rely on intelligent controllers to adapt network operation to changing traffic demands, channel conditions, and service requirements [1], [2]. Because these control decisions directly affect network key performance indicators (KPIs), operators may ask counterfactual questions such as: What KPI would have been observed had a different control action been taken? The practical value of answering such “what-if” questions depends critically on reliable uncertainty quantification. Operators need not only a prediction of the KPI under an alternative action, but also a faithful measure of its uncertainty, typically expressed through an interval or prediction set designed to contain the target outcome with a prescribed probability [3], [4]. A key difficulty, however, is that the telemetry available for counterfactual analysis may not contain all the information used by the controller to select its action. For example, the controller may rely on fine-grained measurements, lower-layer information, private internal states, or signals exchanged with other controllers, whereas the logged telemetry may contain only aggregated, delayed, or filtered observations [5]–[7]. If an omitted variable affects both the selected action and the resulting KPI, the logged data are subject to hidden confounding [8]. In that case, samples associated with a given action are not representative of what would have occurred had that action been assigned under the observed context, and prediction sets calibrated solely from observational telemetry may fail to satisfy their intended coverage requirements [9]. One possible solution to this problem consists of applying counterfactual analysis to randomized telemetry, namely, data obtained by assigning control actions independently of both the observed network context and any hidden information available to the controller. This randomization removes the dependence between the action and hidden confounders, thereby allowing valid counterfactual coverage to be recovered. In a live network, however, such interventions may be costly, suboptimal, or incompatible with operational constraints. As a result, randomized telemetry is typically scarce, and prediction sets calibrated exclusively on it may be uninformative and highly variable. These issues are illustrated through a scheduling example in Figure 1. At each scheduling frame, a network controller selects either round-robin (RR) or proportional-fair channel-aware (PFCA) scheduling based on the users’ backlog information and the available radio resources.
3
Fig. 1: A RAN controller selects between proportional-fair channel-aware (PFCA) and round-robin (RR) scheduling based on the UEs’ backlogs and available radio resources. The network operation is logged via telemetry that records the backlogs, selected scheduler, and residual backlog after scheduling, but omits resource availability. Counterfactual analysis can use abundant observational telemetry collected during normal operation and limited randomized telemetry in which the scheduler is selected independently of the network context. The goal is to construct a prediction set for the residual backlog that would have been observed under the alternative scheduler. Prediction intervals calibrated exclusively on observational telemetry (red) may fail to cover the ground-truth residual backlog (green), whereas those calibrated exclusively on randomized telemetry (blue) are valid but wide. CV-CCI combines both data sources to produce valid and more informative prediction sets (red-and-blue dashed).
The logged telemetry, however, records only the user-level context and omits the instantaneous radio-resource budget. The goal of reliable counterfactual analysis is to construct a prediction set that contains, with a prescribed probability, the residual backlog that would have been observed under the alternative scheduling policy. Because the unobserved resource budget affects both the policy selected by the controller and the resulting residual backlog, prediction sets constructed solely from observational data may fail to provide valid coverage for the counterfactual residual backlog, as shown in the right panel of Figure 1. Valid coverage can instead be recovered using randomized telemetry, collected by assigning the scheduling policy independently of both UEs’ backlog information and resource availability. Such randomization may degrade network performance, however, and can therefore be applied only sparingly. As illustrated in the right panel of Figure 1, the limited randomized data yield valid but wide prediction sets that provide little information about the counterfactual residual backlog.
4
The above discussion highlights a fundamental tension between validity and efficiency in counterfactual analysis under hidden confounding. Observational telemetry is abundant and informative but may fail to satisfy the desired coverage guarantee, whereas randomized telemetry supports valid inference but is scarce and costly to collect. Motivated by this tension, we propose Confounding-Valid Counterfactual Conformal Inference (CV-CCI), a conformal methodology that combines abundant, potentially confounded observational telemetry with limited randomized telemetry. Specifically, CV-CCI exploits observational data to improve the efficiency of the prediction sets while using randomized data as a validity guardrail against arbitrary hidden confounding. As illustrated in Figure 1, this combination retains finite-sample coverage while producing more informative and stable prediction sets than methods based solely on randomized telemetry. B. Related Work a) Counterfactual Conformal Inference: Counterfactual inference addresses the longstanding causal problem of estimating the potential outcomes that would have been observed under actions other than the one actually taken [10]. Since only the outcome under the realized action is observed, reliable counterfactual analysis requires quantifying uncertainty about the missing potential outcomes. Conformal prediction (CP) uses a calibration sample to convert any predictive model into a set predictor with marginal coverage guarantees, under the assumption that the calibration and test data are exchangeable [11], [12]. In counterfactual settings, however, action selection generally induces a distribution shift between the observed sample and the target interventional data, violating the exchangeability assumption. Weighted CP (WCP) accounts for this shift by weighting calibration samples based on propensity score estimates [13], enabling uncertainty quantification for counterfactual outcomes [14]. These guarantees rely on causal identification assumptions, including the absence of hidden confounding. Conformal sensitivity analyses address departures from this assumption by bounding the severity of hidden confounding and translating these bounds into adjusted uncertainty sets [15], [16]. Closely related work combines observational and randomized data through WCP using an estimated density ratio between the two distributions [9]. In case of high-dimensional telemetry, however, ratio-estimation errors may compromise coverage and reduce efficiency. b) Counterfactual Inference in Wireless Networks: Causal reasoning has been proposed as a foundation for generalizable and explainable AI-native wireless networks [17]. A central capability
5
is counterfactual, or “what-if,” analysis, which asks how network key performance indicators (KPIs) would change under an alternative configuration, control policy, or RAN application. Digital twins provide a natural platform for such analyses by maintaining a virtual representation of the physical network that is synchronized with network telemetry [18]–[20]. Operators can query the twin to compare candidate interventions before deployment, without disrupting the live network [21]. Closer to our setting, conformal counterfactual KPI estimation (CCKE) constructs prediction sets directly from live-network telemetry for the KPI that would have been observed under an alternative RAN application or action [3]. CCKE provides prediction sets with finitesample coverage under the assumption that the logged context captures all variables affecting action selection. Our work addresses the more general setting in which the controller also relies on unlogged network state, thereby introducing hidden confounding. c) Inference with Auxiliary Data: A growing literature develops inference procedures that combine a small dataset drawn from the target distribution with a larger dataset obtained from an auxiliary source. Prediction-powered inference follows this principle by combining labeled data with model predictions on abundant unlabeled data and using the labeled data to correct prediction errors [22]. A prediction-powered correction has also been used to improve the efficiency of conformal counterfactual inference under treatment imbalance, where treated calibration data are scarce [23]. Related approaches extend this principle to treatment-effect estimation across experiments [24], while other work combines randomized and observational studies for treatment-effect estimation [25], [26]. More recently, General Synthetic-Powered Inference (GESPI) has extended this paradigm to a broad class of statistical procedures, using the small target-distribution sample as a guardrail against arbitrary distributional mismatch in the auxiliary data [27]. CV-CCI specializes the GESPI principle to wireless counterfactual prediction, with randomized telemetry providing the target-distribution sample and potentially confounded observational telemetry serving as auxiliary data. C. Contributions We introduce Confounding-Valid Counterfactual Conformal Inference (CV-CCI) for constructing statistically valid and informative prediction sets for counterfactual network KPIs under hidden confounding. The main contributions are summarized as follows: •
We formulate reliable counterfactual KPI inference in wireless networks when the controller selects actions using information that is only partially available in the logged telemetry.
6
This setting captures the hidden confounding that arises from the mismatch between the controller’s information and the operator’s observations. •
We develop CV-CCI, a conformal counterfactual KPI estimation methodology that combines abundant, potentially confounded observational telemetry with limited randomized telemetry through a GESPI-based aggregation rule. We establish finite-sample marginal coverage under arbitrary hidden confounding, up to a user-specified tolerance, without requiring the hidden confounder to be observed or modeled.
•
We evaluate CV-CCI in representative RAN-control settings involving MAC-layer scheduling and handover decisions. The results show that CV-CCI maintains reliable coverage as the strength of hidden confounding increases, while producing more informative and less variable prediction sets than methods calibrated exclusively on randomized telemetry.
D. Organization The remainder of the paper is organized as follows. Section II introduces the system model and formalizes counterfactual KPI inference under hidden confounding. Section III reviews conformal prediction, weighted conformal prediction, counterfactual conformal inference, and GESPI. Section IV presents CV-CCI and establishes its theoretical guarantees. Section V evaluates the proposed methodology in MAC-layer scheduling and handover scenarios. Finally, Section VI concludes the paper. II. S ETTING AND P ROBLEM D EFINITION In this section, we formulate the problem of reliable counterfactual KPI estimation under hidden confounding in the context of RAN control. A. Setting We consider the closed-loop wireless RAN control system illustrated in Figure 2, which operates as follows. At each decision instant, a network controller observes the network context X and selects a control action A from a finite action space A based on X. After action A is applied, the controller records the resulting KPI Y . The action A may correspond to any network decision, such as a scheduling decision, a handover command, the selection of an application to deploy in an O-RAN architecture [28], or the choice of a set of network-configuration parameters. The
7
Fig. 2: Illustration of the considered counterfactual KPI estimation setting. The controller observes the full network context X = (X̃, U ), selects an action A under either the observational or randomized regime S, and records the resulting KPI Y . Telemetry contains (X̃, A, Y, S) but not the hidden context U . Using the recorded telemetry, the goal is to construct a prediction set for the potential KPI Y (a) associated with a target action a and a logged state X̃ = x̃.
KPI Y may quantify any performance metric of interest, such as latency, throughput, reliability, or energy efficiency. Following the potential-outcome framework [10], we define, for each action a ∈ A, the potential KPI that would be observed if action a were selected as Y (a) ∈ Y. Additionally, under the stable unit treatment value assumption (SUTVA) [29], the observed KPI corresponding to the selected action A is equal to the associated potential outcome, i.e., Y = Y (A).
(1)
Condition (1) implies that the reported KPI is consistent with the network decision that was actually executed. It also rules out possible perturbations in performance reporting that would cause the observed KPI to differ from the potential KPI associated with the selected action [3]. 1) Action-assignment regimes: We consider two mechanisms for assigning the action A and introduce the binary variable S ∈ {0, 1} to distinguish between the corresponding actionassignment regimes. The first regime, corresponding to S = 0, represents normal network operation. In this regime, the deployed controller selects action A according to an arbitrary, possibly optimized and stochastic policy that may depend on the full decision context X. We refer to this assignment regime as observational. Under the observational regime, A | X, S = 0 ∼ pobs (· | X),
(2)
8
where pobs (a | X) is the probability of selecting action a ∈ A given the decision context X. The second regime, corresponding to S = 1, represents randomized assignment, as in a randomized controlled trial (RCT) [8]. Under the randomized regime, the action is assigned independently of the decision context X according to a known policy prnd (·), A|S = 1 ∼ prnd (·),
A ⊥ X | S = 1.
(3)
Under the randomized policy prnd (·) every action has nonzero probability of being selected, for all a ∈ A,
prnd (a) > 0,
(4)
an assumption known as the positivity condition [8], [14]. We assume that the contexts and potential KPIs have the same distribution under the two regimes and that the regimes differ only in their action-assignment mechanisms. This assumption is satisfied when the choice of action-assignment regime is independent of the context and potential KPIs. For each regime s ∈ {0, 1}, the conditional joint distribution factorizes as PX,A,{Y (a)}a∈A |S=s = PX PA|X,S=s P{Y (a)}a∈A |X ,
(5)
where PX is the distribution of the decision context X, P{Y (a)}a∈A |X is the conditional distribution of the potential outcomes given X, and PA|X,S=s is the action-assignment mechanism under regime s, given by PA|X,S=s (a | x) =
pobs (a | x), s = 0, prnd (a),
(6)
s = 1.
2) Telemetry logging: The operation of the RAN is recorded through telemetry, with each decision instance logged as the tuple (X̃, A, Y, S),
(7)
where X̃ denotes the logged network state, A is the action selected by the controller, Y is the resulting KPI, and S is the action-assignment regime. Importantly, the logged state X̃ need not contain all the contextual information used by the controller. To make this mismatch explicit, we decompose the controller’s decision context as X = (X̃, U ),
(8)
9
where U collects the components of the controller’s information that are not recorded in the logged telemetry. Accordingly, the context distribution PX corresponds to the joint distribution of the logged and unobserved context components PX = PX̃,U = PX̃ PU |X̃ .
(9)
This formulation captures several practical situations. For example, the logged data may contain summaries of past network measurements obtained by averaging or pooling over time windows, whereas the controller may act on fine-grained instantaneous measurements [5]–[7]. The controller may also rely on internal states that are not exposed through standard interfaces, use protected attributes unavailable at higher layers, or interact with other applications that use information contained in U but absent from X̃. The telemetry available for counterfactual analysis consists of two independent datasets collected under the observational and randomized regimes. The observational dataset contains nobs independent and identically distributed samples from the observational regime, nobs Dobs = (X̃i , Ai , Yi , Si ) i=1 , with Si = 0.
(10)
Similarly, the randomized dataset contains nrnd independent and identically distributed samples from the randomized regime, nrnd Drnd = (X̃i , Ai , Yi , Si ) i=1 , with Si = 1.
(11)
The aggregate telemetry dataset comprises n = nobs + nrnd samples and is denoted as D = Dobs ∪ Drnd ,
(12)
In the following, we focus on the practically relevant regime nobs ≫ nrnd , in which observational telemetry collected during normal network operation is abundant, whereas randomized experimentation is limited by operational costs, safety constraints, or service-level requirements. B. Problem Definition Using the aggregate telemetry dataset D, the goal of reliable counterfactual KPI estimation is to answer what-if questions of the following form: given a logged network state X̃ = x̃, what KPI would have been observed had the controller selected an action a ∈ A?
10
U
X̃
A
Y
Fig. 3: Causal structure under hidden confounding. In the observational regime, action A is selected using the full context X = (X̃, U ), while the unobserved component U may also affect the potential KPIs Y (a)a∈A . Conditioning only on the logged state X̃ therefore does not, in general, block the dependence between A and the potential outcomes. Randomized assignment removes this dependence by making A independent of the context.
For a given target action a ∈ A, we answer such counterfactual queries by constructing a prediction set Γa (X̃, D) ⊆ Y
(13)
that includes the potential KPI Y (a) that would be observed if action a were selected under the logged context X̃. For a user-defined miscoverage level α ∈ (0, 1), the reliability is expressed via the marginal coverage guarantee Pr Y (a) ∈ Γa (X̃, D) ≥ 1 − α,
(14)
where the probability is taken jointly over the dataset D and an independent test unit distributed according to PX̃,U,Y (a) = PX̃ PU |X̃ PY (a)|X̃,U .
(15)
Condition (14) is a marginal coverage guarantee ensuring that the prediction set Γa (X̃, D) contains the target potential KPI Y (a) with probability at least 1 − α. It is important to highlight that, in the presence of hidden confounding, without additional assumptions, observational telemetry Dobs alone does not identify the target distribution PY (a)|X̃ and therefore cannot, in general, support nontrivial distribution-free prediction sets with the
11
desired target-population coverage (14). As illustrated in Figure 3, the unobserved variable U may affect both the selected action A and the potential outcomes {Y (a)}a∈A . Therefore, conditional ignorability with respect to the logged state, {Y (a)}a∈A ⊥ A | X̃, S = 0,
(16)
need not hold under the observational regime [9]. As a result, the observational outcome distribution and target potential-outcome distribution are not equal in general, PY |X̃=x̃,A=a,S=0 ̸= PY (a)|X̃=x̃ .
(17)
It follows that observational samples assigned to action a may not be representative of the potential KPI distribution for the target population. This motivates the use of randomized telemetry Drnd . In fact, under the randomized regime, the action A is independent of the decision context X = (X̃, U ) and PY |X̃=x̃,A=a,S=1 = PY (a)|X̃=x̃ .
(18)
Note that the coverage guarantee in (14) can be satisfied by a counterfactual analysis procedure that always returns the entire KPI space Y. Despite providing coverage, however, such a procedure is uninformative. Therefore, the quality of the counterfactual prediction sets Γa (X̃, D) is measured by their average size. To define this size for different types of KPIs, let µ be a reference measure on the KPI space Y; for instance, µ may be the Lebesgue measure for continuous KPIs and the counting measure for discrete KPIs. For each action a ∈ A, we define the inefficiency of the prediction set Γa (·) as h i Φ(Γa ) = E µ Γa (X̃, D) ,
(19)
where the expectation is taken over the calibration data D and over an independent test unit sampled according to (15). Among procedures satisfying the coverage guarantee in (14), prediction sets with smaller values of Φ(Γa ) correspond to more informative counterfactual answers, because they restrict the plausible values of Y (a) to a smaller region of Y. Finally, we note that potential-outcome prediction sets {Γa (X̃, D)}a∈A can also be used to compare the effect of different actions. For two candidate actions a, b ∈ A and a logged context X̃, the individual treatment effect [8], or action effect, is defined as ∆(a, b) = Y (a) − Y (b),
(20)
12
which quantifies the KPI change that would result from selecting action a instead of action b. Given prediction sets for Y (a) and Y (b), a natural prediction set for ∆(a, b) is obtained by taking the Minkowski difference of the two sets n o Γ∆(a,b) (X̃, D) = ya − yb : ya ∈ Γa (X̃, D), yb ∈ Γb (X̃, D) .
(21)
Indeed, if Γa (X̃, D) and Γb (X̃, D) satisfy (14) for actions a and b, respectively, then, by the union bound, Pr ∆(a, b) ∈ Γ∆(a,b) (X̃, D) ≥ 1 − 2α.
(22)
Thus, prediction sets for individual potential KPIs naturally induce valid prediction sets for action effects, enabling the comparison of candidate control decisions under uncertainty [14]. III. BACKGROUND In this section, we review the statistical tools underlying the proposed methodology. First, we introduce Conformal Counterfactual KPI Estimation (CCKE) [3], which uses weighted conformal prediction (WCP) [13], [14] to construct prediction sets for potential KPIs under conditional ignorability. We then review General Synthetic-Powered Inference (GESPI) [27], a general framework for statistical inference that combines abundant, potentially biased data with a limited amount of unbiased data to improve efficiency while limiting the worst-case degradation in statistical validity. A. Conformal Counterfactual KPI Estimation Conformal Counterfactual KPI Estimation (CCKE) [3] is a counterfactual inference method that targets the coverage requirement (14) under the assumption of no hidden confounding. For the model introduced in Section II, this assumption is verified when all variables influencing action selection are observed in the logged telemetry, i.e., X = X̃, so that the strong ignorability condition in (16) holds. As shown next, under this condition, CCKE exploits WCP to construct prediction sets for potential outcomes that satisfy the marginal coverage guarantee in (14). For the remainder of this subsection, we fix an arbitrary target action a ∈ A and, to simplify notation, we suppress the dependence on a whenever it is clear from the context. Define the subset of the observational telemetry samples Dobs for which the selected action is a as n o Da = (X̃i , Yi ) : (X̃i , Ai , Yi , Si ) ∈ Dobs , Ai = a .
(23)
13
CCKE first partitions Da into a training set Datr and a calibration set Dacal . The training set
Datr is used to fit lower and upper conditional quantile regressors, q̂α/2 (X̃) and q̂1−α/2 (X̃), which estimate the α/2 and 1 − α/2 conditional quantiles of Y (a) given X̃, respectively. For a context X̃, these regressors define the uncalibrated interval for the KPI Y (a) as h i Γ̃a (X̃) = q̂α/2 X̃ , q̂1−α/2 X̃ .
(24)
In general, the set predictor Γ̃a (·) in (24) is not calibrated and does not satisfy the desired marginal coverage guarantee. Therefore, CCKE uses the calibration set Dacal to adjust the predictor Γ̃a (·) by a data-dependent widening parameter γa (x̃). Let Iacal denote the index set of the samples in Dacal and for each calibration sample i ∈ Iacal , define the nonconformity score n
o Vi = V (X̃i , Yi ) = max q̂α/2 X̃i − Yi , Yi − q̂1−α/2 X̃i .
(25)
The nonconformity score Vi is the largest of the lower- and upper-tail residuals. It is positive when Yi lies outside the uncalibrated interval Γ̃a (X̃i ) and may be negative when Yi lies inside it. Based on the calibration nonconformity scores {Vi }i∈Iacal , CCKE sets the adjustment factor as the (1 − α)-quantile of a weighted empirical distribution of the calibration scores. The weighting accounts for the covariate shift between the calibration and test distributions. Specifically, the calibration covariates in Dacal are distributed according to the observational law PX̃|A=a,S=0 , whereas the guarantee in (14) is given for a test covariate distributed according to PX̃ . CCKE corrects this distributional mismatch using importance weights defined by the likelihood ratio 1 . (26) dPX̃|A=a,S=0 pobs (a | X̃ = x̃) Accordingly, contexts for which action a is rarely selected under the observational policy receive r(x̃) =
dPX̃
(x̃) ∝
larger weights, compensating for their under-representation in the action-specific calibration set. When the observational policy pobs (a | X̃) is unknown, the likelihood ratio in (26) must be estimated from data. For a given context X̃, the likelihood ratios are normalized jointly over the calibration samples and the test point, and each calibration sample i ∈ Iacal is assigned the weight r(X̃i ) , j∈Iacal r(X̃j ) + r(X̃)
(27)
r(X̃) . j∈Iacal r(X̃j ) + r(X̃)
(28)
wi (X̃) = P while the test point is assigned the weight w(X̃) = P
14
These weights define the weighted empirical distribution of the nonconformity scores X wi (X̃) δVi + w(X̃) δ+∞ , P̂VX̃ =
(29)
i∈Iacal
where δv denotes the Dirac delta function at v, and the additional mass at ∞ represents the unknown nonconformity score of the test point. The correction γa (X̃) is chosen as the (1 − α)-quantile of the distribution (29) γa (X̃) = Q1−α P̂VX̃ ,
(30)
where, for a distribution P , Qτ (P ) denotes the τ -quantile of P . For a network context X̃ and the target action a, CCKE outputs the prediction set for the potential outcome Y (a) h i ΓCCKE ( X̃) = q̂ ( X̃) − γ ( X̃), q̂ ( X̃) + γ ( X̃) . a a α/2 1−α/2 a
(31)
Assume that the ignorability condition in (16) holds, that the likelihood ratio in (26) is known exactly, and that there exists a constant c > 0 such that pobs (a | X̃ = x̃) ≥ c
(32)
for all x̃ in the support of PX̃ . Under the sampling assumptions of Section II, the prediction set ΓCCKE (·) satisfies the marginal coverage guarantee in (14) [3]. When the likelihood ratio is a estimated, the exact finite-sample guarantee may no longer hold, and the coverage error depends on the quality of the ratio estimate [14]. As shown in the motivating example in Figure 1, in the presence of hidden confounding, the ignorability condition in (16) is not satisfied in general, and the CCKE prediction set in (31) is no longer guaranteed to satisfy the marginal coverage requirement. A direct alternative is to apply CCKE to the randomized telemetry Drnd , which follows the target potential-outcome distribution. However, because nrnd ≪ nobs , the resulting prediction sets may be inefficient. B. General Synthetic-Powered Inference General Synthetic-Powered Inference (GESPI) [27] is a framework for statistically valid inference with auxiliary data. It is motivated by settings in which reliable real-world data from the target distribution are scarce, while large auxiliary datasets are easier to obtain, for example from simulators or deep generative models. Auxiliary data can improve statistical efficiency by
15
increasing the effective sample size, but they may also be biased or distributionally mismatched, thereby invalidating the guarantees of the inference procedure. GESPI addresses this issue by using the reliable data as a validity “guardrail” while exploiting the auxiliary data when they are informative. More specifically, let Drel denote reliable real-world data sampled from the target distribution P , and let Daux denote a larger auxiliary dataset sampled from a possibly different distribution Q. GESPI wraps around a base inference procedure Algα , such as a conformal prediction method or a hypothesis testing procedure, that is assumed to control a target risk at level α when applied to data sampled from the target distribution P . To formalize this notion, let ℓ(a, Z) denote the loss incurred by an inference output a when evaluated on an independent target-distribution sample Z ∼ P . The form of ℓ depends on the inference task. For confidence or prediction sets, it may correspond to the miscoverage loss; for multiple testing, it may measure a false-discovery criterion. The validity assumption on the base procedure is that, when applied to reliable data, its expected loss is at most α, i.e., R(Algα , P ) = E [ℓ (Algα (Drel ), Z)] ≤ α,
(33)
where the expectation is taken over both the reliable dataset Drel and the independent test sample Z ∼ P. The main principle behind GESPI is to leverage the validity of the base procedure while safely incorporating auxiliary data. To this end, GESPI aggregates three outputs of the base procedure: •
The reliable-data output Algα (Drel ), which uses only the reliable data at the nominal error level α.
•
The auxiliary-powered output Algα (Drel ∪ Daux ), which uses both the reliable and auxiliary data at the nominal error level α.
•
The guardrail output Algα+ϵ (Drel ), which uses only the reliable data at the relaxed error level α + ϵ. Here, ϵ > 0 is a user-specified tolerance parameter that determines the maximum additional error one is willing to allow in order to benefit from the auxiliary data.
To combine these three outputs, GESPI uses the natural ordering of the statistical objects returned by the base procedure. For example, if the base procedure returns confidence or prediction sets, then one set is more conservative than another if it contains it. If the base procedure returns rejection regions in a hypothesis-testing problem, then one rejection region is more aggressive
16
than another if it contains more rejected hypotheses. This ordering allows GESPI to aggregate the three outputs in a way that is adapted to the inference task. For base procedures whose outputs are prediction sets, GESPI combines the three sets as g = Alg (Drel ) ∩ Alg (Drel ∪ Daux ) ∪ Alg (Drel ) . Alg α α α+ϵ
(34)
The outer intersection ensures that the GESPI set is never larger than the prediction set obtained using only the reliable data at level α. The union with the guardrail set ensures that the final set always contains the reliable-data prediction set constructed at level α + ϵ. Consequently, the auxiliary-powered set can reduce the size of the reliable-data set only to the extent permitted by the guardrail. Under standard loss monotonicity assumptions, GESPI output in (34) satisfies the deterministic sandwich guarantee g ⊆ Alg (Drel ). Algα+ϵ (Drel ) ⊆ Alg α
(35)
g P ) ≤ α + ϵ, R(Alg,
(36)
and its risk satisfies
regardless of the auxiliary-data distribution Q. In the following section, we specialize this principle to the telemetry setting considered in this work. In particular, we use GESPI to construct counterfactual prediction sets by safely combining abundant observational telemetry, which may be affected by hidden confounding, with a smaller amount of randomized telemetry. IV. C ONFOUNDING -VALID C OUNTERFACTUAL C ONFORMAL I NFERENCE We now present Confounding-Valid Counterfactual Conformal Inference (CV-CCI), a general framework for counterfactual estimation of a wireless KPI under hidden confounding. As discussed in Section II, observational telemetry is abundant but may be confounded because the controller selects actions using the full context X = (X̃, U ), while the analyst observes only X̃. Consequently, CCKE applied to observational telemetry alone may fail to satisfy the marginal coverage guarantee in (14). In contrast, randomized telemetry is unconfounded because actions are randomized independently of both X̃ and U , but such data are scarce because randomized experimentation may require the network to take suboptimal actions. To exploit the large observational dataset
17
without losing the validity provided by randomized data, CV-CCI instantiates GESPI using as a base inference procedure the CCKE construction reviewed in Section III-A. For a target action a ∈ A, define the action-specific subset of randomized telemetry as n o Darnd = (X̃i , Yi ) : (X̃i , Ai , Yi , Si ) ∈ Drnd , Ai = a , (37) and the action-specific subset of observational telemetry as n o Daobs = (X̃i , Yi ) : (X̃i , Ai , Yi , Si ) ∈ Dobs , Ai = a .
(38)
The observational action-specific dataset Daobs is partitioned into a training set Daobs,tr and a
calibration set Daobs,cal . Let Iarnd denote the index set of randomized samples in Darnd , and let Iaobs,cal denote the index set of observational calibration samples in Daobs,cal .
As in CCKE, the training set Daobs,tr is used to fit lower and upper conditional quantile
regressors, q̂α/2 (·) and q̂1−α/2 (·), which define the uncalibrated interval Γ̃a (·) in (24). However, unlike CCKE, CV-CCI calibrates the prediction set Γ̃a (·) by applying GESPI to the WCP construction in Section III-A, treating the randomized telemetry Drnd as the reliable dataset and
the observational telemetry Dobs as the auxiliary dataset. Specifically, CV-CCI constructs three WCP prediction sets corresponding to the three components of the GESPI. For the fitted quantile regressors q̂α/2 (·) and q̂1−α/2 (·), define the calibration nonconformity score evaluated on the observational and on the randomized data as n o obs obs,cal V = V (X̃i , Yi ) : i ∈ Ia ,
(39)
and V
rnd
n o rnd = V (X̃i , Yi ) : i ∈ Ia .
(40)
CV-CCI defines two randomized-only WCP sets by adjusting the set predictor Γ̃a (x̃) using datadependent widening parameters computed from the weighted empirical distribution of the scores in V rnd . Since under the randomized regime actions are independent of (X̃, U ), conditioning on A = a does not change the distribution of X̃. Hence, the likelihood ratio between the distribution of selection covariates in Darnd and the test covariate distribution is constant. It follows that, for a test context X̃, the normalized WCP weights are equal to 1/(|Iarnd | + 1), and the resulting weighted empirical distribution of nonconformity scores for randomized telemetry is ! X 1 P̂Vrnd,X̃ = rnd δv + δ+∞ . |Ia | + 1 rnd v∈V
(41)
18
The empirical distribution P̂Vrnd,X̃ is used to define the randomized-only prediction set h i rnd rnd rnd Γa (X̃) = q̂α/2 (X̃) − γα (X̃), q̂1−α/2 (X̃) + γα (X̃) ,
(42)
where the widening parameter γαrnd (X̃) corresponds to the (1 − α)-quantile of P̂Vrnd,X̃ . Similarly, for a guardrail parameter ϵ > 0 such that α + ϵ < 1, the guardrail set predictor is defined as h i rnd rnd Γguard ( X̃) = q̂ ( X̃) − γ ( X̃), q̂ ( X̃) + γ ( X̃) , (43) α/2 1−α/2 a α+ϵ α+ϵ rnd where the widening parameter γα+ϵ (X̃) corresponds to the (1 − α − ϵ)-quantile of P̂Vrnd,X̃ .
The third set computed by CV-CCI is the auxiliary-powered pooled set. This set corresponds to a WCP set obtained by calibrating the base set predictor Γ̃a (X̃) using a weighted empirical distribution of both randomized and observational calibration scores. Observational calibration samples are weighted using the likelihood ratio between the test distribution PX̃ and the actionspecific observational calibration distribution PX̃|A=a,S=0 given by raobs (x̃) =
dPX̃ dPX̃|A=a,S=0
(x̃) ∝
1 , Pr(A = a | X̃ = x̃, S = 0)
(44)
where Pr(A = a | X̃ = x̃, S = 0) denotes the probability of selecting action a under logged context x̃ in the observational regime. In general, the likelihood ratios in (44) are unknown and must be estimated from data, for example via a probabilistic classifier trained to predict the selected action A from the logged state X̃. For randomized samples, the likelihood ratio is constant, and therefore, the pooled unnormalized weights are defined as β, i ∈ Iarnd , pool r̂a (X̃i ) = r̂obs (X̃ ), i ∈ I obs,cal , i a a
(45)
where β > 0 is a parameter that controls the relative weight assigned to randomized telemetry samples in the pooled auxiliary distribution; in our implementation, we set β = 1. Thus, for a test context X̃, the normalized pooled weights are wipool (X̃) = P
r̂apool (X̃i ) , P pool pool r̂ ( X̃ ) + ( X̃ ) + β obs,cal r̂a rnd a j j j∈Ia j∈Ia
(46)
for i ∈ Iarnd ∪ Iaobs,cal , and for the test covariate wpool (X̃) = P
β P pool (X̃j ) + j∈Iaobs,cal r̂apool (X̃j ) + β j∈I rnd r̂a a
.
(47)
19
These weights define the pooled weighted empirical distribution X pool X P̂Vpool,X̃ = wi (X̃) δV (X̃i ,Yi ) + wipool (X̃) δV (X̃i ,Yi ) + wpool (X̃) δ+∞ . i∈Iarnd
(48)
i∈Iaobs,cal
Based on the empirical distribution P̂Vpool,X̃ , CV-CCI computes the pooled conformal correction γαpool (X̃) = Q1−α P̂Vpool,X̃ , (49) and the corresponding auxiliary-powered prediction set h i pool pool Γpool ( X̃) = q̂ ( X̃) − γ ( X̃), q̂ ( X̃) + γ ( X̃) . α/2 1−α/2 a α α
(50)
This pooled set may be substantially more efficient than the randomized-only sets Γrnd and Γguard a a when the observational data are informative. Nevertheless, because the observational component may remain confounded through U , the prediction set in (50) alone does not provide the marginal coverage guarantee in (14). For this reason, CV-CCI obtains the final prediction set by applying the GESPI aggregation rule as rnd pool guard ΓCV-CCI ( X̃) = Γ ( X̃) ∩ Γ ( X̃) ∪ Γ ( X̃) . a a a a
(51)
Following the GESPI rationale, the guardrail set Γguard (X̃) limits the deterioration in marginal a coverage to the user-specified tolerance ϵ. Proposition 1 (Finite-sample coverage of CV-CCI). Fix an action a ∈ A and let α, ϵ > 0 satisfy α + ϵ < 1. Under the setting and sampling assumptions of Section II, the CV-CCI prediction set in (51) satisfies Γguard (x̃) ⊆ ΓCV-CCI (x̃) ⊆ Γrnd a a a (x̃) for every logged context x̃. Consequently, i h Pr Y (a) ∈ ΓCV-CCI ( X̃) ≥ 1 − α − ϵ, a
(52)
(53)
where the probability is taken jointly over the telemetry data and an independent test unit distributed according to (15). Proof. Since α + ϵ > α, the randomized conformal quantiles satisfy rnd γα+ϵ (x̃) ≤ γαrnd (x̃),
20
and hence Γguard (x̃) ⊆ Γrnd a a (x̃). Therefore, from (51), (x̃) ⊆ Γrnd (x̃) ⊆ ΓCV-CCI Γguard a (x̃). a a Under the randomized regime, the action-specific calibration samples and the test pair (X̃, Y (a)) are exchangeable. Thus, standard split conformal prediction gives h i Pr Y (a) ∈ Γguard ( X̃) ≥ 1 − α − ϵ. a The result follows from Γguard (X̃) ⊆ ΓCV-CCI (X̃). a a The guarantee in (53) establishes that CV-CCI safely combines abundant observational telemetry with scarce randomized telemetry while preserving the validity of counterfactual KPI estimation, up to the user-specified tolerance ϵ, under arbitrary hidden confounding and arbitrary errors in the estimated observational propensity scores. V. E XPERIMENTS In this section, we evaluate CV-CCI and compare it against state-of-the-art counterfactual conformal inference baselines on two representative RAN control tasks: medium access control (MAC)-layer scheduling [30] and handover [31]. The code for reproducing the experiments is available at https://github.com/abdessamedqchohi/Confounding-Valid-Conformal-CounterfactualInference. A. Baselines We compare CV-CCI to the following baselines: •
CCKE [3]. CCKE is the counterfactual conformal KPI estimator reviewed in Section III-A. For each action, it applies WCP to action-specific observational telemetry, using inverse propensity scores to correct for selection bias in the logged data. Its validity guarantee hinges on the conditional ignorability assumption, and therefore may not satisfy the coverage requirement (14) under hidden confounding.
21
•
Weighted Split Conformal Prediction with Density-Ratio Estimation (wSCP-DR) [9]. wSCP-DR addresses hidden confounding by estimating the ratio between the densities of randomized and observational samples assigned to action a, i.e. dPX̃,Y |S=1,A=a (x̃, y). r(x̃, y) = dPX̃,Y |S=0,A=a
(54)
This ratio (54) is estimated by training a probabilistic classifier to distinguish observational from randomized samples. wSCP-DR then applies WCP to both the observational and randomized telemetry, correcting for hidden confounding by reweighting the samples using the estimated density ratios. We evaluate the two-stage variants proposed in [9]: wSCP-DR Inexact and wSCP-DR Exact. Both methods first construct WCP intervals by combining observational and randomized data through density-ratio weighting. They then learn an interval predictor by regressing the lower and upper interval endpoints on the logged context, enabling prediction intervals to be generated for new samples. wSCP-DR Inexact directly uses the regressed endpoints as the final prediction interval, making it computationally efficient but without finite-sample coverage guarantees. wSCP-DR Exact instead applies a second split-conformal calibration step to the learned interval predictor, restoring finite-sample validity. This additional calibration requires reserving part of the already limited randomized data, which may reduce statistical efficiency and lead to wider prediction intervals. •
Guardrail. This baseline corresponds to the guardrail component of CV-CCI. The base quantile regressors are trained using observational telemetry, while randomized samples are used for calibration at an error level α + ϵ. The guardrail predictor satisfies the marginal coverage requirement at the level 1 − α − ϵ even under hidden confounding. This baseline isolates the benefit obtained by the auxiliary observational component of CV-CCI.
B. Medium Access Control-Layer Scheduling 1) Setup: We consider a resource-allocation problem [30], similar to that studied in [3], for an orthogonal frequency-division multiplexing system with K UEs. At each resource-allocation frame, the controller selects one of two scheduling policies, A ∈ {RR, PFCA} ,
(55)
where RR denotes round-robin scheduling and PFCA denotes proportional-fair channel-aware scheduling.
22
Fig. 4: At each scheduling frame, the MAC scheduler selects an action A ∈ {RR, PFCA} based on the full context X = (X̃, U ), where X̃ contains the logged UE backlogs, CQIs, and packet delays, and U represents the instantaneous PRB budget assigned to the slice by a slicing application. After the selected scheduling policy is applied, the network produces the residualbacklog KPI Y . The PRB budget is available to the scheduler, but is not recorded in the telemetry used for counterfactual inference. Because it may affect both the action selection and the resulting KPI, it acts as a hidden confounder.
The complete decision context comprises the UE backlog sizes {bk }K k=1 , the average packet de-
K lays in the backlogs {dk }K k=1 , the channel quality indicators (CQIs) {ck }k=1 , and the instantaneous
number of available physical resource blocks NPRB . Accordingly, X = [b1:K , c1:K , d1:K , NPRB ] ∈ R3K+1 .
(56)
The logged context contains the UE-level measurements, but not the instantaneous PRB budget X̃ = [b1:K , c1:K , d1:K ] ∈ R3K .
(57)
Hence, the hidden component of the decision context is the total number of PRBs available for allocation, i.e., U = NPRB . As illustrated in Figure 4, this setting models a hierarchical RAN architecture in which a slice-level resource-management application determines the available PRB budget without exposing it through the scheduler’s telemetry interface. The PRB budget nevertheless affects both the preferred scheduling policy and the resulting KPI, and may therefore act as a hidden confounder. The target KPI for counterfactual inference is the vector of residual UE backlog sizes at the end of the scheduling frame under the alternative scheduling policy A ∈ {RR, PFCA}. Specifically,
23
let Yk (A) denote the residual backlog of UE k at the end of the scheduling frame under policy A. The inferential target is then the vector Y (A) = [Y1 (A), . . . , YK (A)] ∈ RK .
(58)
To construct a prediction set for the multidimensional target KPI Y (A), we adopt an approach similar to that in [3]. Specifically, all methods are instantiated using a multidimensional quantile regressor, and the nonconformity score is defined as the maximum of the componentwise nonconformity scores. Calibrating this maximum score yields a prediction set with marginal simultaneous coverage for the entire KPI vector. Under observational telemetry, the probability of selecting RR scheduling is modeled as pobs (RR | X̃, NPRB ) = λ pU (NPRB ) + (1 − λ) pX̃ (X̃),
(59)
where the mixture coefficient λ ∈ [0, 1] controls the relative influence of the unobserved PRB budget and the observed logged context. The first component captures the dependence on the hidden context. Specifically, denoting the sigmoid function by σ(·), we define pU (NPRB ) = σ NPRB − N̄PRB ,
(60)
where N̄PRB is a reference PRB level. Equation (60) implies that the RR scheduler is increasingly favored as the number of available PRBs increases. Following [3], the second component is constructed from an estimate q̂ktx (ck ) of the throughput achievable by UE k under RR scheduling, obtained from its CQI ck . Specifically, we define pX̃ (X̃) = σ − max bk − q̂ktx (ck ) , (61) k
where bk − q̂ktx (ck ) represents the estimated residual backlog of UE k under RR scheduling. Consequently, pX̃ (X̃) decreases as the maximum residual backlog increases, reflecting a preference for PFCA when RR is expected to leave a large backlog for at least one UE. The parameter λ directly controls the degree of hidden confounding. When λ = 0, the scheduling policy depends only on the observed context X̃, so the no-hidden-confounding assumption holds. As λ increases, the policy relies progressively more on the unobserved PRB budget. In the limiting case λ = 1, the controller ignores the logged context entirely and bases its decision solely on the hidden variable NPRB , resulting in the strongest level of hidden confounding.
24
20
1.0
18 16
Interval Size
Coverage
0.9
0.8
0.7
CV-CCI CCKE wSCP-DR exact wSCP-DR inexact Guardrail 1 - α = 0.900 1 - α - = 0.875
14 12 10 8
0.6
6 0.5
0
0.25
0.5
0.75
1
λ
0
0.25
0.5
0.75
1
λ
Fig. 5: Empirical coverage (left) and prediction-set efficiency (right) for the RR potential outcome as a function of the confounding strength λ. The dashed line indicates the nominal coverage level 1 − α = 0.90, while the dotted line indicates the guardrail level 1 − α − ϵ = 0.875. Boxplots summarize the results over nfolds = 100 folds.
Randomized telemetry is generated using a randomized policy under which RR and PFCA are chosen with the same probability, i.e., 1 prnd (RR) = prnd (PFCA) = . 2
(62)
2) Results: We evaluate the benchmarked methods using telemetry generated by the Nokia Wireless Suite Simulator [32]. Specifically, we consider a scenario with K = 8 UEs, where the number of available PRBs is uniformly distributed as NPRB ∼ U{4, . . . , 11}. For each UE k ∈ {1, . . . , K}, the CQI ck is generated according to the standard 3GPP pathloss model specified in Table 7.2.3-1 of TS 36.213 Rel-11 [33]. The initial backlog bk is independently sampled from a geometric distribution with mean 20 000 bits, while the initial delay dk , measured in transmission time intervals (TTIs), is sampled uniformly from {0, . . . , 50}.
We consider an observational telemetry dataset Dobs containing nobs = 7000 samples. Of these,
ntr = 5000 samples are used to train the conditional quantile regressors and propensity models, while the remaining ncal obs = 2000 samples are reserved for calibration. In addition, we consider a limited randomized telemetry dataset Drnd comprising nrnd = 50 samples. Using these datasets, we evaluate all methods with a target miscoverage level of α = 0.1 and a guardrail parameter of ϵ = 0.025.
25
Figure 5 reports empirical coverage and efficiency of the prediction sets for the RR potential outcome produced by the benchmarked methods as the confounding strength varies over λ ∈
{0, 0.25, 0.5, 0.75, 1}. The empirical coverage is measured as the fraction of the nte = 1000
test instances for which the residual backlogs {Yik (a)}K k=1 lie within their respective predicted intervals {Γka (·)}K k=1 . For action a, it is defined as
nte Y K n o X 1 k k da = Cov 1 Y (a) ∈ Γ ( X̃ ) . i i a nte i=1 k=1
(63)
Efficiency is measured by the average size of the predicted backlog intervals {Γka (·)}K k=1 , i.e., nte K 1 X 1 X k b Φa = te µ Γa (X̃i ) , n i=1 K k=1
(64)
where µ(·) denotes the Lebesgue measure. These results highlight the effect of hidden confounding on the benchmarked methods. When λ = 0, and the decision policy does not depend on the hidden PRB budget NPRB , CCKE, wSCP-DR Exact, CV-CCI, and the Guardrail attain coverage close to or above the nominal level 1 − α = 0.9, whereas wSCP-DR Inexact already exhibits substantial undercoverage. As λ increases, the coverage of the methods operating under the no-hidden confounding assumption deteriorates markedly. In particular, the empirical coverage of CCKE decreases from approximately 0.91 at λ = 0 to approximately 0.78 at λ = 1. The coverage of wSCP-DR Inexact follows a similar trend, decreasing from approximately 0.79 to approximately 0.67. Although both methods produce comparatively narrow prediction sets, their efficiency is accompanied by increasingly severe undercoverage. Hence, the reduction in interval width should not be interpreted as an improvement in statistical efficiency. By contrast, CV-CCI, Guardrail, and wSCP-DR Exact maintain empirical coverage above the nominal level for all values of λ. However, while the coverage achieved by CV-CCI and Guardrail remains close to the nominal level, wSCP-DR Exact exhibits noticeable overcoverage, resulting in substantially wider prediction sets. This loss in efficiency stems from the additional data split required to restore finite-sample validity. The resulting reduction in the effective calibration sample is particularly detrimental in the setting considered here, where randomized data are severely limited. Overall, these results show that CV-CCI and the Guardrail achieve the most favorable balance between robustness to hidden confounding and prediction-set efficiency. To further assess the benefit of incorporating observational telemetry, Table I compares the empirical failure rates of CV-CCI and Guardrail for the PFCA potential outcome. The empirical
26
TABLE I: Empirical failure rates of CV-CCI and Guardrail for the PFCA potential outcome. Confounding (λ)
CV-CCI
Guardrail
0.00
30%
47%
0.25
23%
38%
0.50
24%
41%
0.75
28%
48%
1.00
26%
41%
failure rate is defined as the fraction of the nfolds = 100 folds in which empirical coverage falls below the guardrail level 1 − α − ϵ = 0.875, i.e., nfolds n o 1 X b d a,i < 1 − α − ϵ , Fa = 1 Cov nfolds i=1
(65)
d a,i denotes the empirical coverage for action a in fold i. Whereas average coverage where Cov summarizes performance across all data splits, the empirical failure rate measures how often a method fails to attain the prescribed guardrail coverage level. This distinction is practically relevant because, although marginal coverage cannot be guaranteed for every realization of the limited randomized calibration dataset, practitioners would prefer a method that consistently attains the desired coverage level across different realizations of the calibration dataset. The results in Table I show that CV-CCI consistently achieves a lower empirical failure probability than Guardrail across all considered confounding levels. In the absence of hidden confounding, i.e., at λ = 0, Guardrail falls below the coverage level 1 − α − ϵ in 47% of the folds, compared with 30% for CV-CCI. Under maximum confounding, i.e., at λ = 1, the corresponding failure probabilities are 41% for Guardrail and 26% for CV-CCI. We next investigate the effect of the user-specified tolerance ϵ, which determines the maximum additional miscoverage allowed beyond the nominal miscoverage level α when observational telemetry is incorporated as auxiliary data. Although the coverage guarantee in (53) holds for any admissible value of ϵ, its choice directly affects both the empirical coverage and the width of the resulting prediction intervals. We fix the hidden confounding strength at λ = 0.25 and evaluate CV-CCI and Guardrail for ϵ ∈ {0.01, 0.025, 0.05, 0.075, 0.1}, keeping all other experimental parameters unchanged. Figure 6 reports the empirical coverage and prediction-interval efficiency for the PFCA potential outcome Y (PFCA).
27
1.0 20.0
Interval Size
Coverage
0.9 0.8 0.7
CV-CCI Guardrail 1 − α = 0.90
15.0 12.5 10.0
0.6 0.5
17.5
7.5 0.01
0.025
0.05
0.075
0.1
0.01
0.025
0.05
0.075
0.1
Fig. 6: Empirical coverage (left) and efficiency (right) of CV-CCI and Guardrail for the PFCA potential outcome Y (PFCA) as a function of the miscoverage tolerance ϵ.
Both prediction-set methods satisfy the marginal coverage guarantee at level 1−α −ϵ. However, CV-CCI maintains empirical coverage close to the nominal level 1 − α = 0.90 for all considered values of ϵ, whereas the empirical coverage of Guardrail closely tracks the lower guarantee 1 − α − ϵ and therefore decreases as ϵ increases. Moreover, the variability in the empirical coverage of Guardrail increases with ϵ, as reflected by the widening interquartile range. These results show that, by incorporating observational telemetry, CV-CCI mitigates the variability arising from the limited randomized calibration sample while maintaining an empirical coverage level close to the nominal level across the full range of ϵ. C. Handover 1) Setup: We consider the wireless handover scenario illustrated in Figure 7. A UE is initially connected to a serving base station (BS1) and moves within a deployment area that also contains a neighboring base station (BS2). As the UE moves, it periodically reports the received signal strength (RSS) from both BS1 and BS2. At each reporting instant, the network determines whether a handover opportunity exists based on past RSS measurements. When such an opportunity arises, the network decides whether the UE should remain connected to BS1 (A = 0) or be handed over to BS2 (A = 1) based on the traffic load at BS2. Our objective is to answer counterfactual queries about the throughput that would be achieved under either handover decision. Accordingly, for each action a ∈ {0, 1}, we define the potential outcome Y (a) as the average throughput achieved over a look-ahead horizon of H TTIs if action a were taken.
28
Fig. 7: A serving base station (BS1) proposes a handover using the logged RSS histories −τ X̃ = [RSS−τ 1 , RSS2 ], while a neighboring base station (BS2) accepts or rejects the request based
on its instantaneous load U = ℓ2 . The resulting action A ∈ {0, 1} determines whether the UE remains connected to BS1 or is handed over to BS2, and the network produces the corresponding throughput KPI Y . Telemetry records (X̃, A, Y ) but not the target-cell load U . Because U may influence both the handover decision and the resulting throughput, it acts as a hidden confounder in counterfactual KPI inference.
A handover opportunity is declared if the signal from BS2 has been stronger than that from BS1 throughout the previous τ reporting steps. Conditional on such an opportunity, the handover decision is modeled as a two-stage process. First, BS1 decides whether to initiate a handover request based on the latest τ RSS measurements from both cells. Let σ(·) denote the sigmoid +1 +1 function, and let RSS−τ = [RSSt−τ , . . . , RSSt1 ] and RSS−τ = [RSSt−τ , . . . , RSSt2 ] denote 1 1 2 2
the latest τ RSS measurements from BS1 and BS2, respectively. The proposal probability is −τ pprop RSS−τ = σ ∆RSS − mp , 1 , RSS2 where
(66)
τ −1
1X t−i ∆RSS = RSSt−i − RSS 2 1 τ i=0
(67)
is the average RSS advantage of BS2 over BS1 during the observation window, and mp is a reference RSS margin. If BS1 proposes a handover, BS2 decides whether to accept the request based on its instantaneous load ℓ2 . The acceptance probability is pacc (ℓ2 ) = σ λ τℓ − ℓ2
,
(68)
29
where τℓ is a reference load threshold and λ controls the influence of the BS2 load on the acceptance decision. Therefore, under normal network operation, the probability that a handover is executed is −τ −τ −τ pobs A = 1 | RSS−τ pacc (ℓ2 ). 1 , RSS2 , ℓ2 = pprop RSS1 , RSS2
(69)
We assume that BS1 has access to RSS measurements required for handover evaluation, but not to the instantaneous load of BS2. This reflects practical operational deployments, where neighboring cells exchange measurement reports to support mobility management, while detailed scheduler state and resource utilization remain local to each base station. As a result, the logged context available to the BS1 is e = RSS−τ , RSS−τ ∈ R2τ , X 1 2
(70)
whereas the handover decision is made using the full context −τ 2τ +1 X = RSS−τ . 1 , RSS2 , ℓ2 ∈ R
(71)
Therefore, the load ℓ2 acts as the hidden confounder, namely U = ℓ2 , while the parameter λ controls the strength of hidden confounding. When λ = 0, the acceptance decision is randomized and independent of ℓ2 , so the no-hidden-confounding assumption holds. As λ increases, the acceptance decision depends increasingly on the unobserved load of BS2, resulting in progressively stronger hidden confounding. Finally, we assume that randomized telemetry is generated by randomizing the handover decision independently of the observed and hidden contexts, i.e., 1 prnd (A = 0) = prnd (A = 1) = . 2
(72)
2) Results: We evaluate the benchmarked methods on the wireless handover scenario described in Section V-C1 using SionnaRT ray tracing for the physical- and link-layer simulation [34]. We consider an RSS history of length τ = 5 and a prediction horizon of H = 10 TTIs. The observational telemetry dataset comprises nobs = 7000 samples, of which ntr = 5000 are used to train the conditional quantile regressors and propensity models, while the remaining ncal obs = 2000 are reserved for calibration. In addition, we generate a randomized telemetry dataset containing nrnd = 50 samples. The benchmarked methods are evaluated over nfolds = 100 independent folds using nte = 1000 test samples. We set the target miscoverage level to α = 0.1 and the guardrail parameter to ϵ = 0.025.
30
5.0 1.0 4.5 4.0
Interval Size
Coverage
0.9 0.8 0.7
CV-CCI CCKE wSCP-DR exact wSCP-DR inexact Guardrail 1 - α = 0.900 1 - α - = 0.875
3.5 3.0 2.5
0.6 0.5
2.0 0
0.25
0.5
0.75
1
λ
0
0.25
0.5
0.75
1
λ
Fig. 8: Empirical coverage (left) and prediction-interval efficiency (right) for the handover treatment outcome (Y (1)) as a function of the confounding strength λ. The dashed line indicates the nominal coverage level 1 − α = 0.90, while the dotted line indicates the guardrail level 1 − α − ϵ = 0.875. Boxplots summarize the results over nfolds = 100 folds.
We evaluate the benchmarked methods for the potential outcome Y (1) by varying the confounding strength through the parameter λ in the handover acceptance probability (68), considering the values λ ∈ {0, 2, 4, 6, 8}. As discussed in Section V-C1, increasing λ strengthens the dependence of the observational handover policy on the unobserved target-cell load ℓ2 , thereby increasing the degree of hidden confounding. For each value of λ, Figure 8 reports the empirical coverage and efficiency computed according to (63) and (64), with K = 1 since the target potential outcome is scalar. The results confirm the advantage of the proposed CV-CCI methodology, which consistently achieves the best trade-off between statistical validity and prediction-set efficiency. In particular, when λ = 0 and hidden confounding is absent, all methods except wSCP-DR Inexact achieve an empirical coverage above the nominal level 1 − α = 0.9. However, as the confounding strength increases, the coverage of CCKE and wSCP-DR Inexact progressively deteriorates, reaching an empirical coverage level of 0.75 and 0.71 for λ = 8, respectively. Although these methods produce comparatively narrow prediction intervals, they consistently violate the coverage requirement across all values of λ > 0. By contrast, CV-CCI, Guardrail, and wSCP-DR Exact maintain empirical coverage above the nominal level for all values of λ. However, while the coverage achieved by CV-CCI and Guardrail remains close to the nominal level, wSCP-DR Exact exhibits noticeable overcoverage, resulting in substantially wider prediction intervals due to its data inefficiency. Overall, CV-CCI achieves the
31
TABLE II: Empirical failure rates of CV-CCI and Guardrail for the no-handover potential outcome. Confounding (λ)
CV-CCI
Guardrail
0
9%
18%
2
7%
28%
4
12%
25%
6
10%
18%
8
13%
40%
most favorable trade-off between robustness to hidden confounding and prediction-set efficiency, matching the robustness of Guardrail while producing substantially narrower prediction intervals than wSCP-DR Exact. Although CV-CCI and Guardrail provide similar empirical mean coverage and efficiency, they do so with markedly different variability. This is illustrated in Table II, which reports the empirical failure probability of CV-CCI and Guardrail for the potential outcome Y (0). The empirical failure probability is defined as the fraction of the nfolds = 100 folds for which the empirical coverage falls below the guardrail level 1 − α − ϵ = 0.875, and it is defined as in (65). The results show that CV-CCI consistently achieves a lower empirical failure probability than Guardrail across all confounding levels. In the absence of hidden confounding (λ = 0), Guardrail violates the guardrail level in 18% of the folds, compared with only 9% for CV-CCI. As the confounding strength increases, this gap becomes even more pronounced. Under the strongest confounding level, Guardrail fails to satisfy the prescribed coverage in 40% of the folds, whereas CV-CCI reduces this probability to only 13%. In Figure 9 we examine whether the reliability advantage of CV-CCI persists across different choices of the tolerance parameter ϵ. Holding the confounding strength fixed at λ = 2, we vary ϵ over {0.01, 0.025, 0.05, 0.075, 0.1} and evaluate CV-CCI and Guardrail for the no-handover outcome Y (0). All other experimental parameters remain unchanged. The empirical coverage of Guardrail decreases as ϵ increases, remaining close to its theoretical guarantee of 1 − α − ϵ. Its performance also becomes less stable at larger values of ϵ, as indicated by the increasing dispersion across folds. CV-CCI is considerably less sensitive to the choice of ϵ, and its empirical coverage remains close to the nominal level 1 − α. These results highlight the role of the observational component in CV-CCI, which not only reduces variability for a fixed value of ϵ,
1.0
0.35
0.9
0.30
Interval Size
Coverage
32
0.8
0.7
0.6
CV-CCI Guardrail 1 − α = 0.90
0.25
0.20 0.01
0.025
0.05
0.075
0.1
0.01
0.025
0.05
0.075
0.1
Fig. 9: Empirical coverage (left) and efficiency (right) of CV-CCI and Guardrail for the nohandover potential outcome Y (0) as a function of the miscoverage tolerance ϵ.
as observed in Table II, but also enables CV-CCI to achieve empirical coverage closer to the nominal level across different choices of ϵ. VI. C ONCLUSIONS This paper studied reliable counterfactual inference in wireless networks when the available telemetry does not contain all the information used by the controller. This mismatch gives rise to hidden confounding and can invalidate the coverage guarantees of counterfactual methods that rely on observational telemetry alone. To address this problem, we introduced Confounding-Valid Counterfactual Conformal Inference (CV-CCI), a methodology that combines abundant, potentially confounded observational telemetry with limited randomized data to construct counterfactual KPI prediction sets with finite-sample coverage guarantees, up to a user-specified tolerance. Experiments on MAC-layer scheduling and handover demonstrated the practical effectiveness of CV-CCI. Across different levels of hidden confounding, CV-CCI maintained the prescribed coverage while producing informative prediction sets. By contrast, methods relying primarily on observational telemetry increasingly violated the coverage requirement as the strength of confounding increased, whereas alternative valid methods produced substantially wider prediction sets or exhibited greater variability because of the limited randomized sample size. Future work will investigate extensions of the proposed framework to sequential and online counterfactual inference under evolving network conditions [35], interference-aware settings involving interacting network controllers [36], and active intervention strategies that adaptively collect the most informative data while minimizing network disruption [37].
33
R EFERENCES [1] L. Bonati, S. D’Oro, M. Polese, S. Basagni, and T. Melodia, “Intelligence and learning in O-RAN for data-driven NextG cellular networks,” IEEE Communications Magazine, vol. 59, no. 10, pp. 21–27, 2021. [2] M. Polese, L. Bonati, S. D’Oro, S. Basagni, and T. Melodia, “ColO-RAN: Developing machine learning-based xapps for Open RAN closed-loop control on programmable experimental platforms,” IEEE Transactions on Mobile Computing, vol. 22, no. 10, pp. 5787–5800, 2023. [3] Q. Hou, S. Park, M. Zecchin, Y. Cai, G. Yu, and O. Simeone, “What if we had used a different app? Reliable counterfactual KPI analysis in wireless systems,” IEEE Transactions on Cognitive Communications and Networking, vol. 11, no. 5, pp. 3529–3543, 2025. [4] O. Simeone, S. Park, and M. Zecchin, “Conformal calibration: Ensuring the reliability of black-box AI in wireless systems,” IEEE Communications Magazine, 2026. [5] A. Lacava, L. Bonati, N. Mohamadi, R. Gangula, F. Kaltenberger, P. Johari, S. D’Oro, F. Cuomo, M. Polese, and T. Melodia, “dApps: Enabling real-time AI-based Open RAN control,” Computer Networks, vol. 269, p. 111342, 2025. [6] A. S. Abdalla, P. S. Upadhyaya, V. K. Shah, and V. Marojevic, “Toward next generation open radio access networks: What O-RAN can and cannot do!” IEEE Network, vol. 36, no. 6, pp. 206–213, 2022. [7] H. Wen, P. Porras, V. Yegneswaran, and Z. Lin, “A fine-grained telemetry stream for security services in 5G open radio access networks.”
New York, NY, USA: Association for Computing Machinery, 2022, p. 18–23.
[8] M. A. Hernán and J. M. Robins, Causal inference.
CRC Boca Raton, FL, 2010.
[9] Z. Chen, R. Guo, J.-F. Ton, and Y. Liu, “Conformal counterfactual inference under hidden confounding,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, pp. 397–408. [10] D. B. Rubin, “Estimating causal effects of treatments in randomized and nonrandomized studies.” Journal of educational Psychology, vol. 66, no. 5, p. 688, 1974. [11] V. Vovk, A. Gammerman, and G. Shafer, Algorithmic learning in a random world.
Springer, 2005.
[12] G. Shafer and V. Vovk, “A tutorial on conformal prediction,” Journal of Machine Learning Research, vol. 9, pp. 371–421, 2008. [13] R. J. Tibshirani, R. Foygel Barber, E. Candes, and A. Ramdas, “Conformal prediction under covariate shift,” Advances in neural information processing systems, vol. 32, 2019. [14] L. Lei and E. J. Candès, “Conformal inference of counterfactuals and individual treatment effects,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 83, no. 5, pp. 911–938, 2021. [15] M. Yin, C. Shi, Y. Wang, and D. M. Blei, “Conformal sensitivity analysis for individual treatment effects,” Journal of the American Statistical Association, vol. 119, no. 545, pp. 122–135, 2024. [16] Y. Jin, Z. Ren, and E. J. Candès, “Sensitivity analysis of individual treatment effects: A robust conformal inference approach,” Proceedings of the National Academy of Sciences, vol. 120, no. 6, p. e2214889120, 2023. [17] C. K. Thomas, C. Chaccour, W. Saad, M. Debbah, and C. S. Hong, “Causal reasoning: Charting a revolutionary course for next-generation AI-native wireless networks,” IEEE Vehicular Technology Magazine, vol. 19, no. 1, pp. 16–31, 2024. [18] L. U. Khan, Z. Han, W. Saad, E. Hossain, M. Guizani, and C. S. Hong, “Digital twin of wireless systems: Overview, taxonomy, challenges, and opportunities,” IEEE Communications Surveys & Tutorials, vol. 24, no. 4, pp. 2230–2254, 2022. [19] X. Lin, L. Kundu, C. Dick, E. Obiodu, T. Mostak, and M. Flaxman, “6G digital twin networks: From theory to practice,” IEEE Communications Magazine, vol. 61, no. 11, pp. 72–78, 2023. [20] C. Ruah, O. Simeone, and B. M. Al-Hashimi, “A bayesian framework for digital twin-based control, monitoring, and data collection in wireless systems,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 10, pp. 3146–3160, 2023.
34
[21] E. Ak, B. Canberk, V. Sharma, O. A. Dobre, and T. Q. Duong, “What-if analysis framework for digital twins in 6G wireless network management,” in 2024 International Wireless Communications and Mobile Computing (IWCMC). IEEE, 2024, pp. 232–237. [22] A. N. Angelopoulos, S. Bates, C. Fannjiang, M. I. Jordan, and T. Zrnic, “Prediction-powered inference,” Science, vol. 382, no. 6671, pp. 669–674, 2023. [23] A. Farzaneh, M. Zecchin, and O. Simeone, “Synthetic counterfactual labels for efficient conformal counterfactual inference,” arXiv preprint arXiv:2509.04112, 2025. [24] R. Cadei, I. Demirel, P. De Bartolomeis, L. Lindorfer, S. Cremer, C. Schmid, and F. Locatello, “Prediction-powered causal inferences,” Advances in Neural Information Processing Systems, vol. 38, pp. 73 950–73 979, 2026. [25] E. T. Rosenman, G. Basse, A. B. Owen, and M. Baiocchi, “Combining observational and experimental datasets using shrinkage estimators,” Biometrics, vol. 79, no. 4, pp. 2961–2973, 2023. [26] L. Wu and S. Yang, “Integrative R-learner of heterogeneous treatment effects combining experimental and observational studies,” in Conference on Causal Learning and Reasoning.
PMLR, 2022, pp. 904–926.
[27] M. Bashari, Y. Lee, R. M. Lotan, E. Dobriban, and Y. Romano, “General synthetic-powered inference,” in Forty-third International Conference on Machine Learning, 2026. [28] M. Polese, L. Bonati, S. D’oro, S. Basagni, and T. Melodia, “Understanding O-RAN: Architecture, interfaces, algorithms, security, and research challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 2, pp. 1376–1411, 2023. [29] D. B. Rubin, “Formal mode of statistical inference for causal effects,” Journal of statistical planning and inference, vol. 25, no. 3, pp. 279–292, 1990. [30] E. Yaacoub and Z. Dawy, “A survey on uplink resource allocation in OFDMA wireless networks,” IEEE Communications Surveys & Tutorials, vol. 14, no. 2, pp. 322–337, 2011. [31] A. Ahmed, L. M. Boulahia, and D. Gaiti, “Enabling vertical handover decisions in heterogeneous wireless networks: A state-of-the-art and a classification,” IEEE Communications Surveys & Tutorials, vol. 16, no. 2, pp. 776–811, 2013. [32] “Nokia wireless suite,” https://github.com/nokia/wireless-suite, 2021. [33] 3GPP TS 36.21 .213 version 11.11.0 Release 1 11. [34] J. Hoydis, S. Cammerer, F. Ait Aoudia, M. Nimier-David, L. Maggi, G. Marcus, A. Vem, and A. Keller, “Sionna,” 2022, https://nvlabs.github.io/sionna/. [35] E. Gauthier, F. Bach, and M. Jordan, “Backward conformal prediction,” Advances in Neural Information Processing Systems, vol. 38, pp. 103 897–103 929, 2026. [36] P. B. del Prever, S. D’Oro, L. Bonati, M. Polese, M. Tsampazi, H. Lehmann, and T. Melodia, “PACIFISTA: Conflict evaluation and management in open RAN,” IEEE Transactions on Mobile Computing, vol. 24, no. 10, pp. 10 590–10 605, 2025. [37] T. Zrnic and E. Candes, “Active statistical inference,” in Proceedings of the 41st International Conference on Machine Learning, vol. 235.
PMLR, 21–27 Jul 2024, pp. 62 993–63 010.