Conceptio › Archive › arXiv CS
arXiv CSopen access

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning Weiwei Wang1 , Yuqiang Li2,∗ , Xianyi Wu2,∗ and Bingyi Jing3,4, ∗ 1

Department of Statistics and Data Science, Southern University of Science and Technology

arXiv:2609.20389v1 [stat.ML] 17 Sep 2026

2 3

School of Statistics & KLATASDS-MOE, East China Normal University

School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen 4

Shenzhen Loop Area Institute

Abstract Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and risk-aware decision-making. A comprehensive way to unify these tasks is to estimate the sampling distribution of the evaluation error. Existing approaches, however, often suffer from limited robustness, scalability, or finite-sample validity. In this paper, we propose a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs). Unlike classical bootstrap methods that rely on resampling complete episodes, the proposed method regenerates trajectories from an estimated MDP and can therefore accommodate a much broader range of offline data formats, including complete trajectories, transition-level observations, and trajectory fragments. This flexibility further improves finite-sample statistical efficiency. We establish bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value. Extensive simulations show that the proposed method accurately captures the sampling distribution of the OPE estimator, yielding tighter confidence intervals and more accurate variance estimates in most settings. Keywords: Bootstrap, confidence interval, Markov decision process, offline policy evaluation, reinforcement learning, uncertainty quantification. ∗ ∗

Corresponding author, email: [email protected] E-mail addresses: Weiwei Wang, weiweiwang [email protected]; Yuqiang Li, [email protected]; Xianyi Wu,

[email protected]

1

1

Introduction Reinforcement learning (RL) addresses sequential decision-making problems that are commonly

modeled as Markov decision processes (MDPs), with the goal of learning policies that maximize cumulative rewards (Sutton and Barto, 2018). Over recent decades, RL has achieved notable success across a broad range of domains, including healthcare, education, robotics and autonomous driving (Luckett et al., 2020; Riedmann et al., 2025; Han et al., 2023; Li et al., 2026). In many such applications, however, online interaction with the environment is impractical due to safety, cost, or time constraints. This has fueled the development of offline reinforcement learning (ORL), where policies are learned solely from pre-collected offline datasets and large-scale historical data can be effectively leveraged (Levine et al., 2020). Within this paradigm, offline policy evaluation (OPE) is of central importance: before deploying a policy in a safety-critical setting, one must assess not only its expected performance but also the uncertainty associated with that assessment. Most existing OPE studies focus on point estimation, e.g., Fonteneau et al. (2010), Dudı́k et al. (2011), Thomas et al. (2015), Thomas and Brunskill (2016), Jiang and Li (2016), Liu et al. (2018), Farajtabar et al. (2018), Xie et al. (2019), Yin and Wang (2020), Duan et al. (2020), Tang and Wiens (2023), Wang et al. (2024), Liu et al. (2025) and Liu et al. (2026), among many others. However, point estimates alone are insufficient in practice, where a policy that appears superior may in fact underperform once uncertainty is accounted for. In high-stakes applications, reliable decision-making requires principled uncertainty quantification, such as confidence intervals and variance estimates, in order to distinguish genuinely superior policies from those that only appear favorable due to statistical noise. Despite substantial progress in offline policy evaluation, high-confidence OPE remains relatively underdeveloped. Early work mainly focused on confidence intervals for the expected policy value. For instance, Thomas et al. (2015) proposed a high-confidence estimator based on importance sampling (IS) (Precup et al., 2000) and concentration inequalities. However, IS typically requires a known behavior policy and suffers from the “Curse of Horizon” (Liu et al., 2018), often resulting in overly conservative bounds. Hanna et al. (2017) combined Efron’s bootstrap (Efron, 1979) with model-based and weighted doubly robust estimators, but without consistency guarantees. Kostrikov and Nachum (2020) proposed bootstrapping independent transitions, i.e., (s, a, r, s′ ), and established asymptotic guarantees under sufficient coverage assumptions; however, Hao et al. (2021) later showed that transition-level resampling may fail to faithfully recover the error distribution in OPE. To address 2

this issue, Hao et al. (2021) integrated Efron’s bootstrap with fitted Q-evaluation (FQE) and derived theoretical confidence guarantees. Shi et al. (2021) further developed a deeply debiased procedure for constructing asymptotic confidence intervals under smoothness assumptions on the Q-function. Another line of work formulates interval estimation as an optimization problem (Feng et al., 2020, 2021; Dai et al., 2020). While effective in specific settings, these methods are tailored to confidence interval construction and do not directly yield a distributionally consistent approximation that can be reused for broader inferential tasks, such as variance estimation. More recently, Shi et al. (2024) studied confidence interval estimation for infinite-horizon offpolicy evaluation in the presence of unobserved confounding. Dann et al. (2023) investigated simultaneous high-confidence evaluation of multiple target policies in the on-policy setting. Luo et al. (2026) introduced the first simultaneous inference framework for off-policy evaluation, extending pointwise confidence intervals to uniformly valid simultaneous confidence regions over continuous or infinite initial-state spaces. In parallel, conformal approaches have recently been developed for contextual bandits, finite-horizon MDPs, and more general sequential decision-making problems under distribution shift (Taufiq et al., 2022; Foffano et al., 2023; Zhang et al., 2023). While these methods offer attractive finite-sample coverage guarantees, their primary inferential goal is coverage calibration. In this paper, we study uncertainty quantification for OPE in finite-horizon, time-inhomogeneous tabular MDPs. Our goal is to go beyond point estimation and develop a trajectory-regeneration bootstrap framework for both confidence interval construction and variance estimation. To this end, we propose a model-based bootstrap (MB) procedure that regenerates trajectories from an estimated MDP and consistently approximates the sampling distribution of the policy value estimator while preserving the Markovian structure of the data. Unlike the classical episode bootstrap, which fundamentally relies on resampling complete observed trajectories, the proposed procedure operates at the level of the estimated MDP and can therefore exploit a much broader class of offline data, including complete trajectories, transition-level observations, and trajectory fragments. This substantially broadens the scope of bootstrap-based inference and makes the proposed framework particularly appealing for data-constrained and partially logged offline RL settings. Our main contributions are summarized as follows. (1) We establish conditional distributional consistency of the proposed bootstrap procedure for both Monte Carlo estimation in the on-policy case and Plug-in estimation (a minimax-optimal 3

OPE method) in the off-policy case. These results provide a mathematically rigorous justification for model-based bootstrap inference in offline RL and fill an important theoretical gap in the current literature, where uncertainty quantification remains much less developed than point estimation. (2) Building on this distributional theory, we show that the proposed procedure yields asymptotically valid bootstrap confidence intervals and consistent variance estimation. Hence, the framework provides a unified statistical inference framework that goes substantially beyond point estimation, enabling principled uncertainty quantification for offline policy evaluation. (3) We conduct simulation studies in two representative tabular RL environments: Time-varying MDP (a nonstationary MDP with a fixed horizon) and the Cliff-walking environment (a stationary MDP with a random horizon). The experiments assess the accuracy of the proposed model-based bootstrap in approximating the sampling distribution, constructing confidence intervals, and estimating variance. The results show that the proposed method faithfully captures the error distribution, achieves a favorable balance between coverage accuracy and interval efficiency, and provides accurate variance estimation in most settings. The remainder of the paper is organized as follows. Section 2 introduces the model setup and reviews several commonly used OPE estimators. Section 3 presents the model-based bootstrap procedure, develops its theoretical properties, and describes the construction of confidence intervals and variance estimators in both on-policy and off-policy settings. Section 4 reports simulation results, and Section 5 concludes the paper. All proofs are deferred to the Appendix.

2

Preliminaries Let ∆(S) denote the set of all probability distributions over a set S and for a positive integer

H, define [H] := {0, 1, . . . , H − 1}. We consider a finite-horizon, time-inhomogeneous MDP M = (S, A, {Ph }h∈[H] , {Rh }h∈[H] , d0 , H), where S and A are finite state and action spaces, d0 ∈ ∆(S) is the initial state distribution, and H < ∞ is the step horizon. For each step h ∈ [H], Ph (·|s, a) ∈ ∆(S) stands for the transition kernel, and Rh : S × A × S → [0, 1] indicates the reward function with

rh (s, a, s′ ) the immediate random reward received upon the transition (s, a) 7→ s′ at step h.

A policy π = {πh }h∈[H] is a collection of decision rules with πh (·|s) ∈ ∆(A). A trajectory

under a policy π is denoted by ξ = (s0 , a0 , r0 , . . . , sH−1 , aH−1 , rH−1 , sH ), whose distribution Pπ is 4

an integration of s0 ∼ d0 , ah ∼ πh (·|sh ), and sh+1 ∼ Ph (·|sh , ah ) for all h ∈ [H]. Denote by PH−1 G (ξ) = h=0 rh its return, and EM,π (or Eπ if no ambiguity arises) the corresponding expectation.

Denote the state- and action-value functions of π at step h by  X  X H−1 H−1 π π Vh (s) := Eπ rh′ |sh = s and Qh (s, a) := Eπ rh′ |sh = s, ah = a , ′ ′ h =h

h =h

respectively, and the overall performance by vπ := Eπ [G (ξ)]. Goal.

Given an offline dataset oon n n (i) (i) (i) (i) D = ξ (i) = s0 , a0 , r0 , . . . , sH

i=1

,

of n independent trajectories collected by the target policy π or possibly a different behavior policy µ = {µh }h∈[H] , we aim to construct: • a confidence interval CI(δ) with coverage probability 1−δ, δ ∈ (0, 1), such that P (vπ ∈ CI(δ)) ≥ 1 − δ and d (v̂π (D)). • an estimate of the variance of the OPE v̂π (D), i.e., Var

Offline Policy Evaluation.

This study considers the following typical on-policy Monte Carlo

(MC) and off-policy Plug-in evaluation methods: • On-policy MC method evaluates a policy by averaging its returns. Specifically, given a data set D, it is defined as

1 π v̂MC =

n

n X

G(ξ (i) ).

i=1

The MC estimator is unbiased and consistent. Moreover, by the classical central limit theorem,  √ π n (v̂MC − vπ ) ⇒ N 0, σπ2 ,

as

n → ∞,

(1)

where ⇒ denotes converging in distribution, and σπ2 := Varπ [G(ξ)], for whose definition please refer to Lemma A.3 in Appendix. • Off-policy Plug-in estimate is defined as Q̂πh (s, a) =

(i)  (i) (i) (i) π i=1 1{sh = s, ah = a} rh + V̂h+1 (sh+1 ) , Pn (i) (i) = a} = s, a 1{s i=1 h h

Pn

5

whenever the denominator is positive, and set V̂Hπ (s) ≡ 0 and V̂hπ (s) =

X

a∈A

πh (a | s)Q̂πh (s, a),

for h = H − 1, . . . , 0. The Plug-in value estimator is π v̂Plug-in =

X

d0 (s)V̂0π (s).

(2)

s∈S

π is equivalent to the fitted Q-evaluation (FQE), which is Hao et al. (2021) has proved that v̂Plug-in

known to achieve minimax-optimality in several settings, including tabular MDPs, linear function approximation, and certain nonlinear function classes (e.g., Duan et al. (2020); Hao et al. (2021); Zhang et al. (2022)). Lemma 2.1 (Asymptotic normality of the Plug-in estimator) The Plug-in estimator is √ n-consistent and asymptotically normal: √ The asymptotic variance is 2 = σPlug-in

H−1 X h=0

 π 2 n v̂Plug-in − vπ ⇒ N (0, σPlug-in ).

Eµ

"

dπh (sh , ah ) dµh (sh , ah )

2

π (sh+1 ) | sh , ah Var rh + Vh+1

where dµh (s, a) := Pµ (sh = s, ah = a).

# 

,

Remark 2.1 The MC is generally statistically inefficient, as its asymptotic variance σπ2 does not attain the Cramér-Rao (C-R) lower bound; see Lemma A.4 in Appendix. Conversely, Lemma 2.1 establishes the asymptotic efficiency of the Plug-in estimator. Relatedly, Hao et al. (2021) established analogous efficiency results for linear time-homogeneous MDPs.

3

Model-based Bootstrap for OPE For OPE, in addition to point estimation, it is also critical to quantify uncertainty via, e.g.,

confidence intervals and variance estimation. A comprehensive way to unify these tasks is to estimate the distribution of the evaluation error v̂π − vπ , denoted by G. A standard approach to approximating G is Efron’s bootstrap. Prior works (e.g., Hanna et al. (2017); Kostrikov and Nachum (2020); Hao et al. (2021)) have shown that bootstrap methods can 6

provide asymptotically valid confidence intervals and serve as a promising technique for OPE. However, Hao et al. (2021)’s numerical experiments demonstrated that resampling should be performed at the trajectory level, because resampling at the transition level ignores the strong dependence between transitions within a trajectory. We refer to these two approaches as bootstrap by episodes (BE) and bootstrap by transitions (BT), respectively. Specifically, given D, the BE procedure generates a bootstrap dataset D ∗ by sampling trajectories

π,∗ with replacement. Applying the MC estimator to D ∗ yields the BE-MC estimator v̂BE-MC . Its

distributional consistency √

 π,∗ π n v̂BE-MC − v̂MC ⇒ N (0, σπ2 ),

can be readily established by Lemma A.1 in Appendix. Although BE is valid, it is purely nonparametric and may suffer from high variability, especially in small samples. Furthermore, Hao et al. (2021) proved that the BE-FQE off-policy method (i.e., applying the FQE to D ∗ ) in linear episodic homogeneous MDPs is distributionally consistent, and established the consistency of bootstrap confidence interval as well as bootstrap variance estimation. To improve statistical efficiency and relax the complete-trajectory requirement, we propose a model-based bootstrap (MB) procedure that leverages the underlying Markovian structure to estimate G more efficiently. Compared with the nonparametric BE procedure, the proposed MB method can reduce variance and improve finite-sample efficiency, particularly in limited-data settings. We first introduce the model-based bootstrap procedure in Sec. 3.1, and then present the corresponding confidence interval and variance estimation procedures in Sec. 3.2.

3.1

Model-based bootstrap procedure

Specifically, given the offline dataset D, we construct the empirical MDP  c = S, A, {P̂h }h∈[H] , {R̂h }h∈[H] , d0 , H , M

(3)

where the transition kernel is estimated by

P̂h (s′ | s, a) =

nh (s, a, s′ ) , nh (s, a)

with nh (s, a, s′ ) and nh (s, a) denoting the empirical counts of (s, a, s′ ) and (s, a) at step h, respectively. For each (h, s, a, s′ ), let o n (i) (i) (i) (i) (i) (i) (i) Dh (s, a, s′ ) = (sh , ah , rh , sh+1 ) : (sh , ah , sh+1 ) = (s, a, s′ ), i = 1, . . . , n 7

Observed offline data

Regenerated trajectory

Pool local transitions

Trajectory 1

s0

s1

s0

s3

Trajectory 2

s0

s2

s1

s2

s3

s4

s0

s2

s3

s4 A new trajectory.

Fragment

s2

s3

Figure 1: Illustration of the efficient use of offline data by the proposed model-based bootstrap. Observed trajectories and fragments are first pooled to estimate the local transition dynamics of the MDP, after which new trajectories can be regenerated from the estimated model. In this example, the regenerated trajectory s0 → s2 → s3 is supported by local transition information scattered across different samples, even though the full trajectory is not directly observed.

be the collection of observed transitions (s, a) 7→ s′ at step h, and define R̂h (s, a, s′ ) as the sample

mean of the rewards in Dh (s, a, s′ ). Obviously, for every (h, s, a, s′ ) visited with positive probability, the empirical estimators satisfy p

p

P̂h (s′ | s, a) − → Ph (s′ | s, a),

R̂h (s, a, s′ ) − → Rh (s, a, s′ ).

c is available, bootstrap can be performed using transition-level or fragment-level data as Once M well.

c π̄) is denoted by For any policy π̄, a model-based bootstrap trajectory generated from (M,  ∗ ξ ∗ = s∗0 , a∗0 , r0∗ , . . . , s∗H−1 , a∗H−1 , rH−1 , s∗H ,

where s∗0 ∼ d0 , a∗h ∼ π̄h (· | s∗h ), s∗h+1 ∼ P̂h (· | s∗h , a∗h ), and rh∗ is sampled from the empirical reward  distribution associated with the observed transitions in Dh s∗h , a∗h , s∗h+1 . In this way, we generate a bootstrap dataset

e ∗ = {ξ ∗(j) }n , D π̄ j=1

c π̄), ξ ∗ ∼ (M,

which is then used to conduct statistical inference for the target policy π. The advantage of regenerating trajectories from an estimated MDP is not merely computational. By consolidating compatible local dynamics across samples, the proposed procedure makes more efficient use of the information contained in the offline dataset than methods that resample complete observed episodes, as illustrated in Fig. 1. 8

Remark 3.1 The regenerated trajectories considered in this paper are closely connected to the synthetic trajectories in Wang et al. (2024), essentially reflecting the same mechanism from a different viewpoint. Specifically, Wang et al. (2024) breaks observed trajectories into transitions and then restitch them on a “sampling-with-replacement” basis, with the primary goal of improving point estimation accuracy. By contrast, our work adopts the regenerated-trajectory viewpoint to study uncertainty quantification. In what follows, we investigate this idea separately in the on-policy and off-policy settings. 3.1.1

On-policy bootstrap

e ∗ is generated under the target policy π. The For on-policy evaluation, the bootstrap dataset D π

corresponding expectation is denoted by

∗ v̂π := EM,π c [G(ξ )] .

eπ∗ yields the model-based bootstrap estimator v̂ π,∗ Applying the MC estimator to D MB-MC .

In the sequel, we study the asymptotic properties of the estimator and show how it can be used

for CI construction and variance estimation. Theorem 3.1 (On-policy validity of MB-MC) Suppose that the dataset D = {ξ (i) }ni=1 consists of n i.i.d. trajectories generated under the target policy π. Then the following statements hold: as n → ∞, 1. Expectation consistency: p

v̂π − → vπ . 2. Variance consistency: p

→ σπ2 , σ̂π2 −

where

 G(ξ ∗ ) σ̂π2 := VarM,π c

and

 σπ2 := Varπ G(ξ) .

3. Conditional asymptotic normality: conditional on D,

 √  π,∗ 2 − v̂ n v̂MB-MC π ⇒ N (0, σπ ).

4. Distributional consistency: sup P t∈R

√

  p √ π π,∗ → 0. n(v̂MB-MC − v̂π ) ≤ t | D − P n(v̂MC − vπ ) ≤ t − 9

Remark 3.2 Theorem 3.1 shows that the proposed MB-MC procedure consistently approximates the asymptotic sampling distribution of the MC estimator in the on-policy setting; see (1). Besides, in the absence of distribution shift, the MB correctly reproduces the asymptotic fluctuations of the MC estimator. A key feature is that the bootstrap estimator is centered at the expectation value induced by the estimated MDP, rather than the original MC estimator itself. Denote the lower δ-quantile of MB-MC estimate error distribution by   π,∗ π qδ,MB-MC = inf y ∈ R : P v̂MB-MC − v̂π ≤ y|D ≥ δ .

Based on these bootstrap quantiles, we construct a bootstrap confidence interval h i π π π π CIMB-MC (δ) = v̂MC − q1−δ/2,MB-MC , v̂MC − qδ/2,MB-MC .

As a result of Theorem 3.1, the coverage probability of the empirical confidence interval of vπ con-

structed above converges to the nominal level. Corollary 3.1 (Asymptotic coverage validity) As n → ∞, P (vπ ∈ CIMB-MC (δ)) → 1 − δ. 3.1.2

Off-policy bootstrap

In off-policy inference, the central question is no longer whether to bootstrap, but how to generate π bootstrap trajectories under distribution shift. We build on the asymptotic normality of v̂Plug-in , and

focus on how bootstrap trajectory generation affects the resulting inference procedure. c be the estimated Let the offline dataset D be collected under a known behavior policy µ, and let M

MDP learned from D as shown in (3). The target remains the value vπ of a target policy π, whose π as shown in (2). original point estimate is given by v̂Plug-in

The bootstrap dataset for off-policy is collected by the behavior policy µ, that is eµ∗ = {ξµ∗(j) }nj=1 , D

c µ). ξµ∗ ∼ (M,

e ∗ , yielding the bootstrap Apply the Plug-in estimate with target policy π to the regenerated dataset D µ π,∗ . estimator v̂MB-(Plug-in)

Assumption 3.1 (Sufficient data coverage) For every (h, s, a) on the target policy support, one has dµh (s, a) := Pµ (sh = s, ah = a) > 0. 10

Theorem 3.2 (Distributional consistency) Suppose that Assumption 3.1 holds, then the following statement holds:

Consequently, it implies sup P t∈R

√

 √  π,∗ 2 π ). D ⇒ N (0, σPlug-in n v̂MB-(Plug-in) − v̂Plug-in

 π,∗ π − v̂Plug-in n v̂MB-(Plug-in) ≤t



D −P

√

  p π → 0. n v̂Plug-in − vπ ≤ t −

It follows that the behavior-driven model-based bootstrap yields asymptotically valid confidence intervals and consistent bootstrap variance estimation for vπ . Corollary 3.2 (Asymptotic coverage validity) Define the conditional bootstrap quantile  o n  π,∗ π π ≤y|D ≥δ . − v̂Plug-in qδ,MB-(Plug-in) = inf y ∈ R : P v̂MB-(Plug-in) Then the confidence interval i h π π π π − qδ/2,MB-(Plug-in) − q1−δ/2,MB-(Plug-in) , v̂Plug-in CIMB-(Plug-in) (δ) = v̂Plug-in satisfies

 P vπ ∈ CIMB-(Plug-in) (δ) → 1 − δ.

Estimated-behavior-driven bootstrap.

We now consider the practically important setting in

which the offline dataset is generated by a single but unknown behavior policy. In this case, we first estimate the behavior policy from the observed data by µ̂h (a | s) = nh (s, a)/nh (s), and then regenerate a bootstrap dataset e ∗ = {ξ ∗(j) }nj=1 , D µ̂ µ̂

c µ̂). ξµ̂∗ ∼ (M,

c and M, we also need to Here, beyond controlling the discrepancy between the estimated MDP M p

establish the consistency of the behavior policy estimator: µ̂h (a|s) − → µh (a|s). Once the regenerated

c µ̂) is shown to consistently approximate its population counterpart under episode law under (M, (M, µ), the remaining bootstrap validity argument follows the same line as in the known-behavior case.

11

3.2

Offline policy evaluation inference

In this section, we describe how to conduct policy evaluation inference based on the output of above theoretical results. More specifically, the implementation is described in Algorithm 1. • Confidence interval. Compute the δ/2 and 1−δ/2 quantiles of the empirical error distribution  π , q̂ π ε(1) , · · · , ε(B) , denoted as q̂δ/2 1−δ/2 , respectively. The bootstrap confidence interval is h

i π π v̂π − q̂1−δ/2 , v̂π − q̂δ/2 .

(4)

• Variance estimation. To estimate the variance of MC and Plug-in estimators, we calculate the sample variance as

where ε̄ = B1

d (v̂π ) = Var

PB

b=1 ε(b) .

B 2 1 X ε(b) − ε̄ , B−1

(5)

b=1

Algorithm 1: Model-based Bootstrap (i)

(i)

(i)

(i)

Input: Dataset D = {ξ (i) = (s0 , a0 , r0 , . . . , sH )}ni=1 , target policy π, behavior policy µ, confidence level δ ∈ [0, 1], and bootstrap size B. d π ). Output: (1 − δ) confidence interval for vπ and variance estimate Var(v̂ Compute the original estimate v̂π = F(D); // F: Plug-in estimate over D for b = 1, . . . , B do c π) for on-policy, or from (M, c µ) for off-policy; e ∗ = {ξ ∗(j) }n from (M, Generate D j=1 (b) (b) e∗ ) ; Compute v̂ π,∗ = U(D // U: MC for on-policy, Plug-in for off-policy (b)

(b)

π,∗ Set ε(b) = v̂(b) − v̂π . end return the bootstrap confidence interval in (4) and the variance estimator in (5).

4

Simulation Studies In this section, we conduct numerical experiments to evaluate the performance of the proposed

model-based bootstrap (MB) offline policy evaluation method in two representative tabular RL environments: Time-varying MDP (a nonstationary MDP with a fixed horizon) and the Cliff-walking environment (a stationary MDP with a random horizon). The experiments are designed to achieve three main objectives: (1) to compare the accuracy of the error distribution of MC and Plug-in estimate obtained by bootstrapping episodes (BE) and model-based bootstrap (MB) to characterize 12

the true error distributions (see Sec. 4.2); (2) to examine the effectiveness and tightness of the resulting confidence intervals (see Sec. 4.3); (3) to assess the accuracy of the MC and Plug-in variance estimators (see Sec. 4.4).

4.1

Simulation environments

1. The Time-varying MDP consists of two states {s0 , s1 }, and two actions {a1 , a2 }. At each time step, s0 remains unchanged regardless of the chosen action. On the other hand, state s1 transitions to either itself or s0 based on time-varying probabilities. These probabilities are    2/H, if p < 0.5  1, if ph < 0.5 h Ph (s0 |s1 , a1 ; ph ) = and Ph (s1 |s1 , a2 ; ph ) = ,  1,  1 − 2/H, if ph ≥ 0.5 if ph ≥ 0.5

determined by a sequence of i.i.d random numbers ph ∈ U [0, 1] , h ∈ [H], as in Yin and Wang

(2020) and Wang et al. (2024). Here, we modify the immediate reward 1 and 0 to some uniform distributions to make the environment more stochastic. The horizon is set to H = 10. The behavior policy selects both actions with equal probability at each state, while the target policy assigns probability 1/2 to both actions at s0 and 1/4 to a1 and 3/4 to a2 at s1 . 2. Cliff-walking environment is a typical tabular MDP with random horizons as shown in Fig. 2. In this environment, the agent begins at the initial state and will be terminated when it reaches the terminal state or falls off the cliff. At each state, four actions are available, namely A = {up, down, left, right}. We modify the environment in order to make it more stochastic. Specifically, we introduce randomness in state transitions. Given a state-action pair (s, a), the agent moves to the next state indicated by the action with a probability 1 − ǫ and the four neighborhoods with an equal probability ǫ/4. Transferring beyond the boundary makes the agent stay where it is. A reward of −50 is given for falling off the cliff, and a reward of −1 is given for any other transition. We set the parameter ǫ to 0.4 to control the stochasticity of the environment. The target policy is a near-optimal policy trained by Q-learning, whereas the behavior policy is a 0.1-greedy policy used to generate the offline data.

13

Initial

The Cliff

Terminal

Figure 2: Cliff-walking environment.

4.2

Comparison of error distribution estimation

This section investigates how the two resampling ways – BE and MB – approximate the true error distributions of MC and Plug-in policy evaluators, respectively. The ground-truth error distributions are obtained via Monte Carlo simulation with a sample size of 10,000, and the number of bootstrap replications is set to 2,000. Fig. 3 (a) and (b) present the MC and Plug-in on-policy estimate error distributions, respectively, and Fig. 3 (c) reports the off-policy Plug-in error distributions. The results demonstrate that both BE- and MB-based bootstrap distributions closely match the true distribution of the evaluation error v̂π − vπ .

4.3

Comparison of confidence intervals

In this section, we evaluate the performance of MB in confidence interval construction in the above two RL environments. We study the empirical coverage probability and interval width with different sample sizes. The confidence intervals are constructed at significance levels δ ∈ {0.25, 0.1, 0.05}. The number of bootstrap resampling and replications are all set to be 100. The simulation results shown in Tab. 1–4 lead to the following observations: (1) Tab. 1 and 2 report the empirical coverage probabilities and average interval widths in the Time-varying environment. Across both on-policy and off-policy settings, all methods exhibit decreasing interval widths as the sample size increases, confirming that the uncertainty of policy value estimation diminishes with more trajectories. In the on-policy setting, MB-(Plug-in) performs comparably to BE-(Plug-in) and BT-(Plugin), while often producing shorter intervals with empirical coverage close to the nominal level, especially at the practically relevant 0.90 and 0.95 confidence levels. The MC-based methods, BE-MC and MB-MC, show similar but slightly inferior behavior, as MC is not asymptotically optimal. 14

True Distribution

Bootstrap by Episode

(a1)

350

300

300

300

250

250

250

200

200

200

150

150

150

100

100

100

50

50

0

−0.100 −0.075 −0.050 −0.025 0.000

0.025

0.050

0.075

50

0 −0.100 −0.075 −0.050 −0.025 0.000

True Distribution

(a2)

Model-based Bootstrap

350

350

350

300

300

250

250

200

200

150

150

100

100

50

50

−0.6

−0.4

−0.2

0.0

0.050

0

0.075

−0.100 −0.075 −0.050 −0.025 0.000

Bootstrap by Episode

350

0

0.025

0.2

0.4

0.050

0.075

Model-based Bootstrap 400

300

200

100

0

0.6

0.025

−0.6

−0.4

−0.2

0.0

0.2

0.4

0 −0.8

0.6

−0.6

−0.4

−0.2

0.0

0.2

0.4

0.6

0.8

(a) MC estimation error distributions True Distribution 250 200

(b1)

Bootstrap by Episode

350

300

250

250

200

200

150

150

150

100

100

100

50

50

0 −0.08

0

−0.06

−0.04

−0.02

0.00

0.02

0.04

0.06

50

−0.08 −0.06 −0.04 −0.02

True Distribution

0.00

0.02

0.04

0

0.06

−0.06

−0.04

Bootstrap by Episode

−0.02

0.00

0.02

0.04

0.06

Model-based Bootstrap 250

250

250 200

200

(b2)

Model-based Bootstrap

300

200 150

150

150

100

100

50

50

0

0

−0.6

−0.4

−0.2

0.0

0.2

0.4

0.6

100 50

−0.6

−0.4

−0.2

0.0

0.2

0.4

0.6

0

−0.6

−0.4

−0.2

0.0

0.2

0.4

0.6

(b) Plug-in estimation error distributions True Distribution 300

300

250

250

200

200

150

150

150

100

100

100

50

50

0

−0.075 −0.050 −0.025 0.000

0.025

0.050

50

0 −0.08 −0.06 −0.04 −0.02

0.075

True Distribution

250

0.00

0.02

0.04

0.06

0

−0.2

0.0

0.2

0.4

0.6

0

0.04

0.06

0.08

100 50

50

−0.4

0.02

150

100

−0.6

0.00

200

150

50

−0.08 −0.06 −0.04 −0.02

Model-based Bootstrap

200

100

0

250

250

150

0.08

Bootstrap by Episode

300

200

(c2)

Model-based Bootstrap

300

200

250

(c1)

Bootstrap by Episode

350

350

−0.8

−0.6

−0.4

−0.2

0.0

0.2

0.4

0.6

0

−0.6

−0.4

−0.2

0.0

0.2

0.4

0.6

(c) Off-policy Plug-in estimation error distributions Figure 3: Simulation results of MC and Plug-in estimation error distributions in the Time-varying MDP and Cliff-walking environment.

15

In the off-policy setting, the advantage of MB-(Plug-in) is more pronounced. Although distribution shift makes inference more challenging and generally leads to wider intervals, MB-(Plug-in) yields substantially shorter intervals than BE-(Plug-in) and BT-(Plug-in), while maintaining acceptable coverage accuracy, especially when the sample size is small or moderate. For instance, at the 0.95 level with n = 100, the average interval width of MB-(Plug-in) is 1.3641, compared with 1.5261 for BE-(Plug-in) and 1.7652 for BT-(Plug-in). This indicates that MBFQE achieves higher finite-sample statistical efficiency. (2) Tab. 3 and 4 summarize the results for the Cliff-walking environment. In the on-policy setting, all methods show decreasing interval widths as the sample size increases. MB-(Plug-in) achieves coverage close to the nominal levels while producing competitive interval widths, especially at the 0.90 and 0.95 levels. In the off-policy setting, distribution shift leads to substantially wider confidence intervals, particularly for small sample sizes. BE-(Plug-in) and BT-(Plug-in) are more conservative, with BT-(Plug-in) producing especially wide intervals when n is small. MB-(Plug-in) yields much shorter intervals and thus shows a clear efficiency advantage, although it may suffer from finitesample under-coverage under severe distribution shift. Its coverage improves as the sample size increases and approaches the nominal level. All in all, the extensive numerical simulation studies confirm the effectiveness and advantages of MB method in offline policy interval estimation.

4.4

Comparison of variance estimation

In this section, we investigate the performance of variance estimation for the MC and Plug-in policy evaluators in both on-policy and off-policy settings. The ground-truth variance, Var(v̂π ), is approximated via a large-scale Monte Carlo procedure. For each method, we evaluate the variance d π ) − Var(v̂π ) using B = 300 bootstrap resamples over 100 independent repliestimation error Var(v̂

cations. The simulation results are reported in Tab. 5–7. In these tables, we summarize the errors in the form of median [Q1 , Q3 ], where Q1 and Q3 denote the first and third quartiles, respectively. For completeness, the corresponding mean (SD) results are provided in the Appendix. The simulation results reveal several clear patterns: (1) First, across all settings, the variance estimation error generally decreases as the sample size 16

Table 1: Time-varying (on-policy): empirical coverage probability and average interval width. Each entry is reported as Coverage / Width. Method

Level

n = 50

n = 100

n = 200

n = 500

n = 1000

BE-MC

0.75 0.90 0.95

0.78 / 0.8561 0.91 / 1.2166 0.95 / 1.4495

0.76 / 0.6235 0.91 / 0.8972 0.96 / 1.0634

0.81 / 0.4338 0.93 / 0.6207 0.97 / 0.7395

0.77 / 0.2753 0.89 / 0.3927 0.95 / 0.4701

0.71 / 0.1979 0.90 / 0.2814 0.94 / 0.3329

BE-(Plug-in)

0.75 0.90 0.95

0.78 / 0.8101 0.89 / 1.2208 0.93 / 1.4935

0.74 / 0.4765 0.88 / 0.6874 0.95 / 0.8230

0.69 / 0.3280 0.89 / 0.4695 0.93 / 0.5588

0.74 / 0.2054 0.86 / 0.2943 0.91 / 0.3507

0.69 / 0.1454 0.87 / 0.2069 0.95 / 0.2496

BT-(Plug-in)

0.75 0.90 0.95

0.81 / 0.9035 0.90 / 1.3676 0.96 / 1.6827

0.74 / 0.4845 0.88 / 0.6987 0.98 / 0.8480

0.70 / 0.3302 0.89 / 0.4711 0.94 / 0.5584

0.73 / 0.2066 0.88 / 0.2960 0.94 / 0.3534

0.64 / 0.1460 0.86 / 0.2089 0.96 / 0.2497

MB-(Plug-in)

0.75 0.90 0.95

0.66 / 0.6734 0.87 / 0.9777 0.96 / 1.1837

0.74 / 0.4600 0.85 / 0.6598 0.95 / 0.7868

0.71 / 0.3241 0.87 / 0.4694 0.94 / 0.5622

0.72 / 0.2063 0.88 / 0.2946 0.92 / 0.3508

0.65 / 0.1456 0.86 / 0.2070 0.95 / 0.2462

MB-MC

0.75 0.90 0.95

0.79 / 0.8689 0.92 / 1.2443 0.94 / 1.4804

0.72 / 0.6137 0.92 / 0.8826 0.95 / 1.0545

0.83 / 0.4341 0.91 / 0.6219 0.98 / 0.7412

0.76 / 0.2791 0.89 / 0.3975 0.94 / 0.4725

0.74 / 0.1948 0.90 / 0.2778 0.94 / 0.3303

Table 2: Time-varying (off-policy): empirical coverage probability and average interval width. Each entry is reported as Coverage / Width. Method

Level

n = 50

n = 100

n = 200

n = 500

n = 1000

BE-(Plug-in)

0.75 0.90 0.95

0.79 / 1.5465 0.88 / 2.3545 0.91 / 2.9800

0.76 / 0.8706 0.89 / 1.2610 0.92 / 1.5261

0.82 / 0.5587 0.89 / 0.8203 0.95 / 0.9950

0.76 / 0.3214 0.87 / 0.4612 0.96 / 0.5556

0.77 / 0.2233 0.91 / 0.3234 0.95 / 0.3884

BT-(Plug-in)

0.75 0.90 0.95

0.81 / 1.7004 0.87 / 2.5893 0.93 / 3.2895

0.79 / 0.9427 0.91 / 1.4160 0.93 / 1.7652

0.77 / 0.5826 0.93 / 0.8510 0.95 / 1.0623

0.74 / 0.3214 0.91 / 0.4629 0.97 / 0.5582

0.78 / 0.2267 0.91 / 0.3271 0.94 / 0.3884

MB-(Plug-in)

0.75 0.90 0.95

0.77 / 1.1785 0.89 / 1.7392 0.89 / 2.1339

0.72 / 0.7737 0.86 / 1.1235 0.93 / 1.3641

0.75 / 0.5217 0.88 / 0.7479 0.94 / 0.8956

0.75 / 0.3232 0.92 / 0.4592 0.92 / 0.5478

0.75 / 0.2292 0.90 / 0.3288 0.95 / 0.3944

17

Table 3: Cliff-walking (on-policy): empirical coverage probability and average interval width. Each entry is reported as Coverage / Width. Method

Level

n = 10

n = 50

n = 100

n = 200

n = 500

BE-MC

0.75 0.90 0.95

0.66 / 12.2985 0.79 / 17.4210 0.88 / 20.8415

0.72 / 5.9936 0.90 / 8.4850 0.93 / 10.1985

0.73 / 4.1815 0.89 / 6.0345 0.93 / 7.2514

0.75 / 3.0496 0.88 / 4.3553 0.94 / 5.1888

0.75 / 1.8777 0.89 / 2.7034 0.93 / 3.2301

BE-(Plug-in)

0.75 0.90 0.95

0.64 / 11.9818 0.81 / 17.0041 0.88 / 20.0601

0.68 / 5.8675 0.87 / 8.3158 0.95 / 10.0290

0.71 / 4.2237 0.88 / 6.0918 0.94 / 7.2759

0.73 / 2.9391 0.88 / 4.2192 0.93 / 5.0268

0.80 / 1.8997 0.89 / 2.7557 0.95 / 3.2815

BT-(Plug-in)

0.75 0.90 0.95

0.93 / 36.9530 0.99 / 79.1637 1.00 / 117.3243

0.80 / 6.9459 0.94 / 12.9712 0.97 / 26.6507

0.76 / 4.4046 0.91 / 6.5181 0.96 / 8.4290

0.73 / 2.9856 0.87 / 4.3052 0.95 / 5.2033

0.76 / 1.8883 0.90 / 2.6942 0.94 / 3.2238

MB-(Plug-in)

0.75 0.90 0.95

0.64 / 12.4012 0.78 / 17.5708 0.89 / 20.8501

0.73 / 5.8929 0.89 / 8.4319 0.94 / 10.0468

0.76 / 4.1429 0.88 / 5.9971 0.95 / 7.2251

0.72 / 2.9835 0.90 / 4.2695 0.95 / 5.0948

0.76 / 1.9303 0.92 / 2.7220 0.95 / 3.2301

MB-MC

0.75 0.90 0.95

0.66 / 12.8410 0.87 / 17.9860 0.91 / 21.6290

0.70 / 5.9955 0.90 / 8.5994 0.93 / 10.3744

0.72 / 4.2637 0.90 / 6.1361 0.93 / 7.3029

0.76 / 2.9879 0.91 / 4.3061 0.96 / 5.1521

0.76 / 1.8801 0.89 / 2.7084 0.96 / 3.2329

Table 4: Cliff-walking (off-policy): empirical coverage probability and average interval width. Each entry is reported as Coverage / Width. Method

Level

n = 50

n = 100

n = 200

n = 500

n = 1000

BE-(Plug-in)

0.75 0.90 0.95

0.72 / 178.5161 0.90 / 470.0272 0.93 / 943.9822

0.84 / 11.5095 0.97 / 37.3836 0.99 / 66.5477

0.63 / 3.2889 0.86 / 4.6406 0.94 / 5.5635

0.74 / 2.0457 0.89 / 2.9168 0.93 / 3.5171

0.75 / 1.4331 0.90 / 2.0659 0.94 / 2.4708

BT-(Plug-in)

0.75 0.90 0.95

0.78 / 331.1729 0.92 / 986.3986 0.99 / 2370.0668

0.85 / 13.8215 0.94 / 68.6383 1.00 / 263.9416

0.68 / 3.2364 0.87 / 4.6177 0.90 / 5.5442

0.75 / 2.0699 0.91 / 2.9364 0.95 / 3.4890

0.81 / 1.4709 0.90 / 2.0879 0.95 / 2.5196

MB-(Plug-in)

0.75 0.90 0.95

0.37 / 73.6656 0.58 / 279.7528 0.72 / 687.4008

0.64 / 4.3549 0.82 / 12.0482 0.86 / 95.9276

0.63 / 3.1060 0.81 / 4.4200 0.88 / 5.5729

0.73 / 2.0727 0.89 / 2.9721 0.95 / 3.5011

0.77 / 1.4524 0.88 / 2.0848 0.94 / 2.5026

18

increases, which is consistent with the asymptotic variance consistency. (2) For on-policy MC estimation (Tab. 5), MB-MC consistently outperforms or matches BE-MC in both environments, indicating that the model-based bootstrap approach yields more accurate variance estimates. For on-policy Plug-in estimation (Tab. 6), MB-(Plug-in) achieves the best overall performance, while BE-(Plug-in) is generally competitive. By contrast, BT-(Plug-in) is highly unstable in the Cliff-walking environment when the sample size is small, although its performance improves as n grows. In the Time-varying MDP, the differences among the three methods become much smaller, but MB-(Plug-in) remains the most accurate or nearly the most accurate throughout. (3) For off-policy Plug-in estimation (Tab. 7), the contrast between the two environments is even more pronounced. In the Time-varying MDP, MB-(Plug-in) attains the smallest error across nearly all sample sizes, demonstrating clear robustness and accuracy in the off-policy setting. In the Cliff-walking environment, all methods incur substantial variance estimation errors in small samples, with BT-(Plug-in) being particularly unstable and exhibiting extremely large errors. Although BE-(Plug-in) and MB-(Plug-in) are also inaccurate when n is small, they remain far more stable than BT-(Plug-in). Overall, the results suggest that the MB-based method consistently outperforms the alternatives in most scenarios, with particularly pronounced advantages in the small-sample regime. Table 5: Variance estimation error of on-policy MC estimate under different sample sizes. Entries are reported as median [Q1, Q3]. Sample size n 50 100 200 500 1000 10 50 100 200 500

BE-MC MB-MC Panel A: Time-varying MDP

0.022 [0.010, 0.032] 0.011 [0.005, 0.019] 0.010 [0.004, 0.016] 0.005 [0.002, 0.008] 0.004 [0.002, 0.006] 0.003 [0.001, 0.004] 0.001 [0.001, 0.002] 0.001 [0.000, 0.001] 0.001 [0.000, 0.001] 0.000 [0.000, 0.001] Panel B: Cliff-walking environment 8.392 [4.745, 13.168] 6.841 [2.977, 10.833] 0.807 [0.375, 1.274] 0.652 [0.278, 1.058] 0.282 [0.135, 0.487] 0.236 [0.128, 0.375] 0.122 [0.060, 0.224] 0.093 [0.044, 0.166] 0.048 [0.023, 0.071] 0.040 [0.023, 0.069]

19

Table 6: Variance estimation error of on-policy Plug-in estimate under different sample sizes. Entries are reported as median [Q1, Q3]. Sample size n 50 100 200 500 1000 10 50 100 200 500

BE-(Plug-in) BT-(Plug-in) Panel A: Time-varying MDP 0.046 [0.021, 0.081] 0.070 [0.043, 0.118] 0.005 [0.002, 0.008] 0.005 [0.003, 0.009] 0.002 [0.001, 0.002] 0.002 [0.001, 0.002] 0.001 [0.000, 0.001] 0.000 [0.000, 0.001] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] Panel B: Cliff-walking environment

MB-(Plug-in) 0.011 [0.005, 0.016] 0.003 [0.001, 0.005] 0.001 [0.000, 0.002] 0.000 [0.000, 0.001] 0.000 [0.000, 0.000]

7.127 [2.972, 11.298] 853.455 [505.645, 1294.019] 5.739 [2.645, 10.286] 0.761 [0.381, 1.279] 30.153 [5.655, 117.543] 0.550 [0.217, 0.936] 0.257 [0.160, 0.419] 0.451 [0.194, 1.390] 0.246 [0.119, 0.384] 0.144 [0.067, 0.217] 0.123 [0.056, 0.209] 0.117 [0.049, 0.204] 0.047 [0.022, 0.072] 0.041 [0.025, 0.076] 0.035 [0.018, 0.067]

Table 7: Variance estimation error of off-policy Plug-in estimate under different sample sizes. Entries are reported as median [Q1, Q3]. Sample size n 50 100 200 500 1000 50 100 200 500 1000

BE-(Plug-in) BT-(Plug-in) Panel A: Time-varying MDP

MB-(Plug-in)

0.148 [0.062, 0.341] 0.308 [0.135, 0.527] 0.064 [0.029, 0.098] 0.033 [0.015, 0.057] 0.057 [0.035, 0.085] 0.016 [0.006, 0.030] 0.009 [0.005, 0.015] 0.012 [0.006, 0.023] 0.004 [0.003, 0.010] 0.002 [0.001, 0.003] 0.002 [0.001, 0.003] 0.002 [0.001, 0.003] 0.001 [0.000, 0.001] 0.001 [0.000, 0.001] 0.001 [0.000, 0.001] Panel B: Cliff-walking environment 552.9 [551.8, 554.1] 130294.6 [57771.7, 283078.6] 554.7 [445.6, 555.6] 39.72 [39.30, 40.14] 15409.64 [4213.79, 33701.15] 40.42 [39.89, 40.72] 4.211 [4.027, 4.364] 324.875 [4.166, 2510.311] 4.362 [4.238, 4.483] 0.051 [0.032, 0.088] 0.052 [0.023, 0.083] 0.053 [0.027, 0.084] 0.021 [0.008, 0.035] 0.024 [0.011, 0.039] 0.028 [0.016, 0.042]

20

5

Conclusions This paper studies uncertainty quantification for offline policy evaluation using a model-based

bootstrap (MB) approach, with guarantees of asymptotic distributional consistency. Extensive simulations demonstrate that the proposed MB procedure accurately captures the sampling distribution of the policy value estimator, yields confidence intervals with reliable coverage and competitive tightness, and provides accurate variance estimation. The simulation studies conducted in two representative tabular RL environments lead to the following conclusions: (1) both the MC and Plug-in estimate error distributions obtained by MB can correctly characterize their true distributions in both on-policy and off-policy scenarios; (2) in most cases, the confidence intervals constructed by BE and MB exhibit similar behavior, with empirical coverage converging to the nominal level as the sample size increases; however, in small-sample and off-policy regimes, MB consistently outperforms BE; (3) variance estimates of MC and Plug-in estimate based on MB are generally more accurate than those based on BE and BT, particularly when the sample size is limited; (4) both confidence interval construction and variance estimation based on BT perform substantially worse than BE- and MB-based methods, with deficiencies being most pronounced in small-sample settings. One limitation of the current method is that it relies on the support-overlap assumption, requiring the support of the target policy to be contained in that of the behavior policy. Extending it to settings with limited support overlap is an important direction for future work. Overall, the model-based bootstrap resampling method holds promise for uncertainty quantification in offline policy evaluation scenarios. Its weaker requirements on data formats make it more applicable to real-world applications.

References Bickel, P. J. and Freedman, D. A. (1981). Some asymptotic theory for the bootstrap. The Annals of Statistics 9(6), 1196–1217. Van der Vaart, A. W. (2000). Asymptotic statistics. Volume 3. Cambridge University Press. Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvári, C. and Schuurmans, D. (2020). CoinDICE: Off-policy confidence interval estimation. In Advances in Neural Information Processing Systems 33, 9398–9411. Dann, C., Ghavamzadeh, M. and Marinov, T. V. (2023). Multiple-policy high-confidence policy evaluation. In International Conference on Artificial Intelligence and Statistics, 9470–9487.

21

Duan, Y., Jia, Z. and Wang, M. (2020). Minimax-optimal off-policy evaluation with linear function approximation. In International Conference on Machine Learning, 2701–2709. Dudı́k, M., Langford, J. and Li, L. (2011). Doubly robust policy evaluation and learning. arXiv preprint arXiv:1103.4601. Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics 7(1), 1–26. Farajtabar, M., Chow, Y. and Ghavamzadeh, M. (2018). More robust doubly robust off-policy evaluation. In International Conference on Machine Learning, 1447–1456. Feng, Y., Ren, T., Tang, Z. and Liu, Q. (2020). Accountable off-policy evaluation with kernel Bellman statistics. In International Conference on Machine Learning, 3102–3111. Feng, Y., Tang, Z., Zhang, N. and Liu, Q. (2021). Non-asymptotic confidence intervals of off-policy evaluation: Primal and dual bounds. arXiv preprint arXiv:2103.05741. Foffano, D., Russo, A. and Proutiere, A. (2023). Conformal off-policy evaluation in Markov decision processes. In 2023 62nd IEEE Conference on Decision and Control, 3087–3094. Fonteneau, R., Murphy, S. A., Wehenkel, L. and Ernst, D. (2010). Model-free monte carlo-like policy evaluation. In International Conference on Artificial Intelligence and Statistics, 217–224. Han, D., Mulyana, B., Stankovic, V. and Cheng, S. (2023). A survey on deep reinforcement learning algorithms for robotic manipulation. Sensors 23(7), 3762. Hanna, J., Stone, P. and Niekum, S. (2017). Bootstrapping with models: Confidence intervals for off-policy evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence 31, 4933–4934. Hao, B., Ji, X., Duan, Y., Lu, H., Szepesvári, C. and Wang, M. (2021). Bootstrapping fitted Q-evaluation for off-policy inference. In International Conference on Machine Learning, 4074–4084. Jiang, N. and Li, L. (2016). Doubly robust off-policy value evaluation for reinforcement learning. In International Conference on Machine Learning, 652–661. Kostrikov, I. and Nachum, O. (2020). Statistical bootstrapping for uncertainty estimation in off-policy evaluation. arXiv preprint arXiv:2007.13609. Levine, S., Kumar, A., Tucker, G. and Fu, J. (2020). Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643. Li, Y., Tian, M. Zhu, D., Zhu, J., Lin, Z., Xiong, Z. and Zhao, X. (2026). Drive-r1: Bridging reasoning and planning in vlms for autonomous driving with reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence 40, 6708–6716. Liu, P., Zhao, L., Agarwal, S., Liu, J., Huang, A., Amortila, P. and Jiang, N. (2026). Model selection for off-policy evaluation: New algorithms and experimental protocol. In Advances in Neural Information Processing Systems 38,

22

42942–42976. Liu, Q., Li, L., Tang, Z. and Zhou, D. (2018). Breaking the curse of horizon: Infinite-horizon off-policy estimation. In Advances in Neural Information Processing Systems 31, 5356–5366. Liu, S. D., Chen, C. and Zhang, S. (2025). Efficient multi-policy evaluation for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence 39, 18951–18959. Luckett, D. J., Laber, E. B., Kahkoska, A. R., Maahs, D. M., Mayer-Davis, E. and Kosorok, M. R. (2020). Estimating dynamic treatment regimes in mobile health using v-learning. Journal of the American Statistical Association 115(530), 692–706. Luo, T., Fan, X. and Wu, W. (2026). Simultaneous statistical inference for off-policy evaluation in reinforcement learning. In Advances in Neural Information Processing Systems 38, 70396–70430. Precup, D., Sutton, R. S. and Singh, S. (2000). Eligibility traces for off-policy policy evaluation. In International Conference on Machine Learning, 759–766. Riedmann, A., Schaper, P. and Lugrin, B. (2025). Reinforcement learning in education: A systematic literature review. International Journal of Artificial Intelligence in Education, 1–55. Shi, C., Wan, R., Chernozhukov, V. and Song, R. (2021). Deeply-debiased off-policy interval estimation. In International Conference on Machine Learning, 9580–9591. Shi, C., Zhu, J., Shen, Y., Luo, S., Zhu, H. and Song, R. (2024). Off-policy confidence interval estimation with confounded Markov decision process. Journal of the American Statistical Association 119(545), 273–284. Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction. MIT Press. Tang, S. and Wiens, J. (2023). Counterfactual-augmented importance sampling for semi-offline policy evaluation. In Advances in Neural Information Processing Systems 36, 11394–11429. Taufiq, M. F., Ton, J. F., Cornish, R., Teh, Y. W. and Doucet, A. (2022). Conformal off-policy prediction in contextual bandits. In Advances in Neural Information Processing Systems 35, 31512–31524. Thomas, P. and Brunskill, E. (2016). Data-efficient off-policy policy evaluation for reinforcement learning. In International Conference on Machine Learning, 2139–2148. Thomas, P., Theocharous, G. and Ghavamzadeh, M. (2015). High-confidence off-policy evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence 29, 3000–3006. Wang, W., Li, Y. and Wu, X. (2024). Off-policy evaluation for tabular reinforcement learning with synthetic trajectories. Statistics and Computing 34(1), 41. Xie, T., Ma, Y. and Wang, Y. X. (2019). Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling. In Advances in Neural Information Processing Systems 32, 9665–9675.

23

Yin, M. and Wang, Y. X. (2020). Asymptotically efficient off-policy evaluation for tabular reinforcement learning. In International Conference on Artificial Intelligence and Statistics, 3948–3958. Zhang, R., Zhang, X., Ni, C. and Wang, M. (2022). Off-policy fitted Q-evaluation with differentiable function approximators: Z-estimation and inference theory. In International Conference on Machine Learning, 26713–26749. Zhang, Y., Shi, C. and Luo, S. (2023). Conformal off-policy prediction. In International Conference on Artificial Intelligence and Statistics, 2751–2768.

24

Record · ID 978449 · SHA-256 d674fb7105610840
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.