Bet on Features: Anytime-Valid and Feature-Aware Auditing of
arXiv:2607.11653v1 [cs.LG] 13 Jul 2026
Conditional Quantile Forecasters Ivane Antonov∗
Sohom Mukherjee∗
Julius-Maximilians-Universität Würzburg
Julius-Maximilians-Universität Würzburg
Richard Pibernik
Yo Joong Choe
Julius-Maximilians-Universität Würzburg
INSEAD
Zaragoza Logistics Center
Abstract Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We develop a distribution-free and game-theoretic testing framework for continuously auditing black-box conditional quantile forecasters with non-i.i.d. losses, such that the resulting evidence process is powerful against predictably chosen alternatives specified by the features available to the auditor. We first formalize notions of conditional quantile calibration when different sets of features are available to the auditor, establishing that the coarseness of the auditor’s information set determines the hardness of the testing problem. We then identify the sets of alternatives for which the auditor can achieve power, and focusing on contextual bets linear in the features, we derive finite-time detection guarantees for such alternatives, all without an i.i.d. assumption. The resulting evidence processes are interpretable at the feature level, as they quantify fine-grained, “feature-aware” evidence for miscalibration. We empirically validate these methods on simulated and real data, finding that a popular time series forecaster (Chronos-2) is highly miscalibrated w.r.t. multiple relevant features.
1
Introduction “Most of the literature implicitly assumes homoskedastic errors even when this is clearly violated, and proceed by merely testing for correct unconditional coverage. Consequently, I set out to build a consistent framework for conditional interval forecast evaluation.” —Christoffersen (1998)
In his seminal work, Christoffersen (1998) named the requirement that has informed forecast evaluation for nearly three decades: tests must be sensitive to conditional miscalibration, because ∗
Equal contribution.
1
failures that average out marginally can be predictable and exploitable in context. This perspective is increasingly relevant as large black-box models (Ansari et al., 2025; Liu et al., 2025) are used for the probabilistic forecasting of covariate-rich time series in real-world applications (e.g., Yang et al., 2025). Modern deployments add a second statistical difficulty: forecasts need to be monitored continuously. Practitioners inspect calibration as outcomes arrive, stop once evidence accumulates, or intervene after a suspected distribution shift (Hoga and Demetrescu, 2023). Fixed-horizon backtests are not designed for this workflow and can lose their nominal Type-I guarantees under optional stopping. This necessitates incorporating ideas from the safe anytime-valid inference paradigm (Ville, 1939; Shafer et al., 2011; Vovk and Wang, 2021; Ramdas et al., 2023; Grünwald et al., 2024). An audit must, therefore, satisfy two requirements at once: it must be conditional in a sense rich enough for covariate-driven failures, and it must be anytime-valid. In particular, for decision-making problems under asymmetric cost, e.g., inventory management (Cao and Shen, 2019), public health response (Doms et al., 2018), and financial risk (Engle and Manganelli, 2004), the relevant action is often based on the conditional quantile (as opposed to the conditional mean) based on past history (Koenker and Bassett Jr, 1978). Persistent bias in the quantile forecasts (too high or too low relative to the conditional law) translates into downstream decision error, necessitating conditional quantile calibration. In econometrics and risk management, this translates to the popular problem of backtesting value-at-risk (Christoffersen and Pelletier, 2004). Existing frameworks satisfy at most one of the above requirements. In this paper, we develop an anytime-valid, distribution-free audit for black-box conditional quantile forecasters, while making the auditor’s information explicit in the calibration notion. This information-aware viewpoint of calibration is essential in operational settings. The forecaster may have access to rich internal context, while an external auditor may observe only a coarser view. For instance, a wholesaler auditing a retailer’s demand forecasts may see aggregate sales and past forecast errors, but not the retailer’s promotion calendar. Conversely, an internal auditor may have access to promotions, prices, calendar features, and lagged errors. These situations can not be modelled by a single null hypothesis: the calibration property being tested, and the alternatives against which the audit can have power, are determined by the information available to the auditor. Figure 1 shows this gap on Rossmann store-sales data (Cukierski, 2015) using Chronos-2 (Ansari et al., 2025) forecasts: marginal audits rarely reject, while promotion- and Saturday-aware audits reveal feature-specific miscalibration and accumulate anytime-valid evidence. Our framework formalizes this idea through a sequential betting game. At each time, the 2
forecaster issues a quantile forecast, the outcome is observed, and the auditor records whether the outcome fell below the forecast. Before seeing the next outcome, the auditor chooses a bet using only its own monitoring information. If the forecast is calibrated relative to that information, every legal betting strategy produces a nonnegative test martingale, and hence an e-process. Thus large accumulated wealth is valid evidence against calibration even under continuous monitoring and data-dependent stopping (Ville, 1939; Shafer et al., 2011; Vovk and Wang, 2021; Ramdas et al., 2023; Grünwald et al., 2024). The construction uses only issued forecasts and realized outcomes; it does not require a parametric model, stationarity, independence, or moment assumptions on the outcomes. This information-indexed view separates validity from power. Coarser audits are safe, but they can be blind: a marginal audit may remain anytime-valid even when the forecaster is systematically wrong in contexts visible only through features the auditor does not observe. Thus the auditor’s available information determines both the calibration null being tested and the alternatives the audit can detect. To gain power, the auditor must bet on predictable structure visible in its own information. Scalar bets can detect simple aggregate bias, but feature-specific errors require featureaware bets. We therefore introduce contextual betting: the auditor supplies a predictable feature dictionary—such as calendar indicators, prices, promotions, lagged hits, or lagged forecasts—and an online learner adapts the betting direction from the observed stream. The learned weights are interpretable, identifying which features expose conditional miscalibration. Theoretically, we show that if full-feature miscalibration has persistent predictable edge along some linear feature direction, then any contextual learner with a pathwise regret bound yields finite-time detection and stopping-time guarantees. A more general cumulative-edge result in the appendix allows time-varying and intermittent feature-aligned edge. Our work is related to the literature on safe sequential inference for quantiles, forecast calibration, elicitable functionals, and risk backtesting. We highlight the main connections in the following, and refer the reader to Appendix A for additional related work. Howard and Ramdas (2022) construct time-uniform confidence sequences for population quantiles from i.i.d. observations, while Mineiro and Howard (2023) build time- and value-uniform confidence bands for CDFs under nonstationarity; both are estimation problems, whereas we audit a conditional quantile forecast and ask which features make its errors detectable. Arnold et al. (2023) develop sequentially valid tests for probabilistic forecast calibration, closely related to PIT calibration, which requires distributional forecasts rather than a single quantile. Methodologically, Casgrain et al. (2024) provide a broad supermartingale 3
framework for elicitable and identifiable functionals, and our audit can be viewed as a quantile-hit specialization; however, our null is indexed by the auditor’s monitoring information, leading to a study of feature-aware validity and power. The e-backtesting framework of Wang et al. (2025) is possibly the most closely related to our work: their VaR e-statistic corresponds to one endpoint of our betting interval. The main differences are that e-backtesting conditions on the full market information filtration and studies power in i.i.d. settings, whereas we study information-indexed calibration and prove finite-time power under non-i.i.d., history-dependent streams. We summarize our major contributions in the following: 1. Information-indexed anytime-valid calibration audits. We formulate continuous auditing of black-box conditional quantile forecasts as a sequential testing-by-betting game indexed by the auditor’s monitoring information (Section 2). Predictable no-bankruptcy bets generate test martingales, hence e-processes, that are valid under arbitrary non-i.i.d. streams and optional stopping (Corrolary 2.6). We prove a hierarchy of calibration nulls, show that validity transfers from coarser to richer nulls, and show that coarser valid audits can nevertheless be powerless against richer feature-conditional violations (Corr. 2.7, Prop. 2.8). 2. Feature-aware contextual betting with finite-time power. We develop contextual betting strategies that learn linear bets over a predictable feature dictionary available to the auditor (Section 3). For non-i.i.d. alternatives with persistent feature-aligned predictable edge, a regretto-power theorem converts any pathwise OCO regret bound into finite-time detection and stopping-time guarantees (Theorem 3.2). A more general cumulative-edge alternative, allowing time-varying or intermittent edge, is analyzed in Appendix C.3 (Theorem C.8). 3. Empirical evidence for feature-specific miscalibration. We validate our proposed featureaware audit in three settings (Section 4). On a controlled two-state Markov hit process, we exhibit the marginal-versus-conditional null gap. Under a perfectly calibrated Negative-Binomial oracle, empirical rejection rates respect the Ville-prescribed test level across quantiles (Type-I validity). Replacing the oracle with Chronos-2 forecasts on synthetic and Rossmann store-sales data, feature-aware skeptics detect conditional miscalibration that marginal audits miss.
4
Figure 1: Chronos-2 quantile forecasts on Rossmann store-sales data can pass marginal audits while failing feature-aware audits. Left: across stores and quantile levels, marginal audits rarely reject, whereas promotion- and Saturday-aware audits reveal substantial conditional miscalibration. Right: for one store, marginal wealth stays nearly flat, while feature-aware wealth grows and crosses anytime-valid rejection thresholds.
2
Information-aware anytime audits
Operational audit with asymmetric information. We formalize conditional quantile forecast calibration auditing as a sequential testing-by-betting game. A central theme of our work is information asymmetry. In operational forecasting, the forecaster may have access to rich internal context (e.g., promotions, store identifiers, recent demand), while the auditor may observe only a coarser view. As motivated before, a wholesaler auditing a retailer’s demand quantile forecasts may see past forecast errors and aggregate sales, but not the retailer’s promotion flag. A coarser audit may fail to detect miscalibration that is visible only through the presence of a missing feature. This feature-aware notion of evidence is the organizing principle of the paper. Conditional quantile forecast calibration game. Let (Ω, F, F) be a filtered measurable space, with F = (Ft )t≥0 . All laws considered below are probability measures on (Ω, F). We index the filtration so that Ft−1 is the richest pre-outcome information available before observing the outcome Yt . For rounds t = 1, 2, . . ., we have the following sequence of events. The covariates Ct which is Ft−1 -measurable is revealed. For fixed α ∈ (0, 1), the black-box forecaster issues a Ft−1 -measurable point forecast q̂t ∈ R. The forecast is intended as a conditional α-quantile forecast for Yt . Finally, the outcome Yt is revealed, which is Ft -measurable. Moreover, we will use a predictable feature dictionary ϕt ∈ Rd , Ft−1 -measurable, to denote the full collection of audit-relevant pre-outcome features (or analogously context), i.e., including Ct and
5
possibly further information like past forecasts and hit sequences. A more nuanced formulation of the conditional quantile calibration game involving the forecaster, the skeptic (game-theoretic terminology for the auditor), and the nature can be found in Appendix B.1. To define the null hypothesis for conditional quantile calibration, we first set up some notations. Define the uncentered hit and centered hit by Bt := 1{Yt ≤ qbt }, and Zt := Bt − α. Since q̂t is Ft−1 -measurable and Yt is Ft -measurable, Zt is Ft -measurable and takes values in {−α, 1 − α}. Monitoring information structure. What is crucial to our further discussion is the different information sets available to the auditor. A monitoring information structure is a sequence H = (Ht )t≥0 such that, before the outcome at time t is observed, Ht−1 ⊆ Ft−1 , and after the outcome is observed the hit Zt is included in the monitored information. We consider three marg monitoring levels. First, is the marginal information denoted by Ht−1 . This could include, for
example, the past hit record or some aggregate thereof. Second, we have the single-feature (or w . This could include, for example, a single coordinate scalar-feature) information denoted by Ht−1 ϕ ϕt,j of the covariate. Finally, we have the full-feature monitoring information denoted by Ht−1 . By marg w ⊆ Hϕ construction, Ht−1 ⊆ Ht−1 t−1 ⊆ Ft−1 , for every t.
Definition 2.1 (H-conditional quantile calibration null1 ). For any monitoring information structure H, define the information-indexed conditional quantile calibration null as the composite class P0 (H) := {P : P (Yt ≤ qbt | Ht−1 ) = α, P -a.s., for every t ≥ 1} . Equivalently, P ∈ P0 (H) iff EP [Zt | Ht−1 ] = 0, P -a.s., for every t ≥ 1. The equivalence is due to the bounded martingale difference sequence (MDS) property of (Zt )t≥1 with respect to H, and formally proved in Proposition C.3. Thus P0marg := P0 (Hmarg ) is marginal sequential calibration, P0w := P0 (Hw ) is calibration with respect to one scalar feature, and P0ϕ := P0 (Hϕ ) is calibration with respect to the full feature dictionary. The nulls follow a hierarchy as outlined in the following Proposition. Proposition 2.2 (Null hierarchy). Let H and G be two monitoring information structures such that Ht−1 ⊆ Gt−1 , for every t ≥ 1. Then P0 (G) ⊆ P0 (H). Consequently, P0ϕ ⊆ P0w ⊆ P0marg . Thus full-feature calibration is the strongest requirement and marginal calibration is the weakest requirement. 1
For a general conditional law, q̂t is a valid conditional α-quantile iff P (Yt < q̂t | Ht−1 ) ≤ α ≤ P (Yt ≤ q̂t | Ht−1 ) P -a.s. The exact conditional hit coverage null above (Definition 2.1) is a stronger requirement. The two notions coincide under the no-atom condition P (Yt = q̂t | Ht−1 ) = 0 P -a.s. for every t.
6
at round t
Reality
Yt
P0marg
validity transfers inward
Forecaster
Zt = 1{Yt ≤ q̂t } − α
P0ϕ
P0w
⊆
Ft−1
q̂t
Ht−1
Skeptic
λt
Mt = Mt−1 (1 + λt Zt ) marg P0marg \ P0w : Ht−1 -aware test has no power w P0w \ P0ϕ : Ht−1 -aware test has no power
informs λt+1
(a) One game round (Def. B.1): Forecaster announces q̂t ∈ Ft−1 ; Reality reveals Yt ; Skeptic, with information Ht−1 ⊆ Ft−1 , bets λt ∈ Λα on hit Zt = 1{Yt ≤ q̂t } − α, and wealth updates as Mt = Mt−1 (1 + λt Zt ).
(b) Hierarchy of calibration nulls (Proposition 2.2). Refining the monitoring information structure shrinks the null: P0ϕ ⊆ P0w ⊆ P0marg . Coarser information tests might fail to gain power (Proposition 2.8).
Figure 2: Left: one round of the conditional quantile calibration game with information asymmetry and feature-aware bets. Right: the induced hierarchy of nulls, indexed by the monitoring subfiltration. In particular, if we reject the single-feature calibration null P0w , w.r.t. any feature, then we readily reject the full-feature calibration null P0ϕ (see Figure 2b). Moreover, rejecting P0w for a particular feature would give us strictly more information, in that we additionally know conditional on which feature the forecaster is miscalibrated. We can view the testing problem studied in existing work (e.g., Casgrain et al., 2024; Wang et al., 2025) as the case where Ht ≡ Ft , ∀t. The corresponding null hypothesis is at least as small as P0ϕ , and thus it is an easier testing problem for which there are more possibilities for rejection. Test supermartingales and e-processes. We now recall the notions of sequential testing needed to convert the bounded MDS property above into anytime-valid evidence. The following definitions follow the modern e-value and e-process literature (Vovk and Wang, 2021; Ramdas et al., 2022, 2023; Grünwald et al., 2024; Waudby-Smith and Ramdas, 2024; Ramdas and Wang, 2025). Definition 2.3 (Test supermartingales and e-processes). A nonnegative, H-adapted process (St )t≥0 is a test supermartingale for P0 (H) if S0 ≤ 1 and, for every P ∈ P0 (H), EP [St | Ht−1 ] ≤ St−1 , P -a.s., for all t ≥ 1. An e-process for P0 (H) is a nonnegative, H-adapted process (Mt )t≥0 such that, for every P ∈ P0 (H), there exists a test supermartingale (StP )t≥0 with Mt ≤ StP P -a.s. for all t ≥ 0. Every test supermartingale is therefore an e-process. Definition 2.4 (Skeptic’s wealth process and betting domain). An H-predictable betting strategy is a sequence (λt )t≥1 such that λt is Ht−1 -measurable and takes values in Λα = [−1/(1 − α), 1/α]. The wealth process for a skeptic with such bets is Mt := 7
i=1 (1 + λi Zi ), where M0 := 1.
Qt
The domain Λα is the maximal no-bankruptcy interval, i.e., the largest interval ensuring 1+λt Zt ≥ 0. Since Zt ∈ {−α, 1 − α}, nonnegativity is equivalent to 1 + λ(1 − α) ≥ 0 and 1 − αλ ≥ 0, which simplifies to Λα . Under the null, these multipliers have conditional mean one: for every predictable λt taking values in Λα and every P ∈ P0 (H), EP [1 + λt Zt | Ht−1 ] = 1 + λt EP [Zt | Ht−1 ] = 1. Hence Et := 1 + λt Zt is a conditional e-variable (Definition C.1), and the product wealth process is a test martingale under every null law. We formally state this well-known construction (e.g., WaudbySmith and Ramdas, 2024) adapted to this problem as a Theorem below; by Ville’s inequality (Ville, 1939), we immediately have an anytime-valid test. Theorem 2.5 (Validity of predictable betting under a monitoring null). Let (λt )t≥1 be any Hpredictable betting strategy taking values in Λα , define (Mt ) as in Def. 2.4. If P ∈ P0 (H), then (Mt )t≥0 is a nonnegative P -martingale with respect to the monitoring filtration H. Consequently, (Mt ) is a test martingale, and hence an e-process, for P0 (H). Corollary 2.6 (Ville’s inequality). Under the assumptions of Theorem 2.5, for every P ∈ P0 (H)
and every c > 0, P supt≥0 Mt ≥ c
≤ 1c . Equivalently, for any γ ∈ (0, 1), the rejection time
τγ := inf{t ≥ 1 : Mt ≥ 1/γ}, with inf ∅ := ∞, satisfies supP ∈P0 (H) P (τγ < ∞) ≤ γ. Validity transfers inward while power is lost under coarsening. For our problem, we are particularly interested in how the validity of tests can transfer from coarser (larger) to richer (smaller) nulls. Intuitively, this is true because Type I error (or equivalently, e-process validity) holding uniformly over a null set also holds uniformly over its subset. This is formalized in the following Corollary. Corollary 2.7 (Validity transfers from coarser to richer nulls). Suppose that two monitoring filtrations H = (Ht )t≥0 and G = (Gt )t≥0 satisfy Ht−1 ⊆ Gt−1 for every t. Then every Hpredictable betting process that is valid for the larger null P0 (H) is automatically valid for the smaller, richer null P0 (G). In particular, marginal-valid tests are valid under P0w and P0ϕ , and one-feature-valid tests are valid under P0ϕ . While the previous corollary shows validity, it does not guarantee anything about the power of the coarser information test against violations of the richer null. A coarser test can remain perfectly valid under a richer null while being blind to its violations. The next proposition formalizes this point.
8
Proposition 2.8 (Coarser valid tests can miss richer-null violations). Let H and G satisfy Ht−1 ⊆ Gt−1 for every t. Suppose a law Q satisfies Q ∈ P0 (H) \ P0 (G). Then Q violates the richer G-conditional calibration null, but every H-predictable no-bankruptcy wealth process remains a nonnegative martingale under Q. Consequently, for every γ ∈ (0, 1), Q(τγ < ∞) ≤ γ, where τγ = inf{t ≥ 1 : Mt ≥ 1/γ}. Moreover, for every finite horizon T , EQ [log MT ] ≤ 0, i.e., a coarser auditor has no positive systematic log-growth under Q, even though Q violates calibration at the richer information level.
3
Feature-aware power and stopping times
The preceding section established the validity aspect of the audit: for any monitoring information structure H, every H-predictable no-bankruptcy betting strategy yields an anytime-valid e-process for P0 (H), with no distributional assumptions. We also noted that power of a test is informationsensitive: as Proposition 2.8 shows, a skeptic using a coarser information set can remain perfectly valid while having no positive systematic log-growth against violations that are visible only after conditioning on richer features. Equivalently, wealth grows only when the chosen bets align with conditional deviations visible through the skeptic’s own information. This section characterizes alternatives when a full-feature skeptic has finite-time power. An alternative law Q may be nonstationary and history-dependent; what matters is whether the fullϕ feature conditional miscalibration mϕt := EQ [Zt | Ht−1 ] has a component that can be exploited by
predictable feature-aware bets. If the conditional hit probability were known, the skeptic could use the Kelly-optimal stake that maximizes conditional expected log-growth. In practice this probability is unknown, and the relevant miscalibration may be a feature-specific pattern rather than a scalar bias. We therefore consider contextual bets of the form λt (θ) = ⟨θ, ϕt ⟩, and use online learning to adapt θt from the observed stream. Thus the full-feature skeptic can use the entire predictable feature dictionary to learn profitable betting directions, while the learned weights also serve as interpretable diagnostics for which features expose the forecaster’s conditional miscalibration. Alternatives induced by the hierarchy. We first describe how alternatives are induced by the hierarchy of monitoring information. Fix an ambient model class P of probability laws on (Ω, F). For any monitoring information H, define the corresponding alternative as Q(H) := P \ P0 (H). In 9
particular, Qmarg := P \ P0marg , Qw := P \ P0w , Qϕ := P \ P0ϕ . Because the nulls are nested, the alternatives are nested in the reverse direction: Qmarg ⊆ Qw ⊆ Qϕ . Hence full-feature alternatives are the broadest: a forecaster may violate full-feature calibration while still satisfying every coarser calibration null available to a weaker auditor. Conditional means at different information levels. For any law Q and any monitoring information H, write mH t (Q) := EQ [Zt | Ht−1 ]. When the law is clear from context, we suppress Q. marg ϕ w For the three canonical levels, write mmarg := EQ [Zt | Ht−1 ], mw t := EQ [Zt | Ht−1 ], mt := EQ [Zt | t ϕ Ht−1 ].
Then Q ∈ Qa if and only if there exists at least one time t such that mat ̸= 0 on an event of positive marg Q-probability, where a ∈ {marg, w, ϕ}. Moreover, by the tower property, mmarg = EQ [mw t | Ht−1 ], t ϕ w mw t = EQ [mt | Ht−1 ]. These identities quantify information loss: coarser skeptics observe only
conditional averages of the miscalibration visible at richer information levels. The remainder of the section develops contextual betting strategies that exploit the full-feature signal mϕt by learning feature-aligned bets online. Full-feature contextual bets.
ϕ Assume the full-feature vector ϕt ∈ Rd is Ht−1 -measurable
and satisfies ∥ϕt ∥2 ≤ R a.s. for all t, for a known constant R > 0. Define the feasible set K := {θ ∈ Rd : ∥θ∥2 ≤ 1/(2R)}. For θ ∈ K, define the contextual betting fraction λt (θ) := ⟨θ, ϕt ⟩. Since |Zt | ≤ 1 and ∥ϕt ∥2 ≤ R, Cauchy–Schwarz yields |λt (θ)Zt | ≤ ∥θ∥2 ∥ϕt ∥2 |Zt | ≤ 12 . Hence 1 + λt (θ)Zt ≥ 1/2 > 0, so every predictable sequence θt ∈ K induces the valid no-bankruptcy wealth process Mtctx :=
t Y
(1 + ⟨θi , ϕi ⟩Zi ) .
i=1
For a fixed comparator θ ∈ K, define the one-step loss gt (θ) := − log(1+⟨θ, ϕt ⟩Zt ) and the full-feature ϕ 2 quadratic edge bQ t (θ) := ⟨θ, ϕt ⟩mt − ⟨θ, ϕt ⟩ . Next, we define the linear predictable-edge alternative
Qϕ,lin κ,t0 (Φ). For this alternative class, the full-feature miscalibration is exploitable by a linear betting direction in the dictionary Φ. The skeptic can learn linear bets of the form λt (θ) := ⟨θ, ϕt ⟩ via OCO to obtain finite-time power and stopping-time bounds, against this alternative class. We state the uniform-edge case in the main text because it yields a clean stopping-time bound in Theorem 3.2; the more general cumulative-edge regret-to-power theorem, which allows nonlinear envelopes and intermittent edge, is relegated to the Appendix (see Thm. C.8 and Remark C.10). Definition 3.1 (Linear full-feature predictable-edge alternative). Fix κ > 0, t0 ∈ N, and a 10
predictable feature sequence Φ = (ϕt )t≥1 . The linear full-feature predictable-edge alternative is the class n
o
Q ⋆ ⋆ Qϕ,lin κ,t0 (Φ) := Q : ∃θ ∈ K such that bt (θ ) ≥ κ Q-a.s. for every t ≥ t0 .
Regret-to-power for full-feature contextual betting.
(1)
ϕ Let an online learner output Ht−1 -
measurable vectors θt ∈ K. With Mtctx and gt defined as before, we can define the regret against a fixed comparator θ ∈ K by RTctx (θ) :=
t=1 {gt (θt ) − gt (θ)}. We have the following result.
PT
Theorem 3.2 (Power against the linear full-feature alternative). Assume that the online learner satisfies the pathwise regret bound RTctx (θ) ≤ rT for every θ ∈ K and every T ≥ 1, where (rT )T ≥1 is deterministic. Let Q ∈ Qϕ,lin κ,t0 (Φ) with κ ∈ (0, 1]. Define
nctx γ (r) := inf n ≥ t0 : log(1/γ) + 2(t0 − 1) + rT ≤
κT for every T ≥ n , 2
ctx ctx > T ) ≤ exp(−κ2 T /2). with inf ∅ := ∞. If nctx γ (r) < ∞, then for every T ≥ nγ (r), Q(τγ
Consequently,
Q τγctx < ∞ = 1 h
i
EQ τγctx ≤ nctx γ (r) +
(Power) 4 κ2
(Stopping time)
ctx Remark 3.3 (Interpreting the certification time nctx γ (r)). The quantity nγ (r) is a deterministic
certification time, which is different form the stopping time: it is the first horizon after which the linear predictable edge guaranteed by the alternative dominates the Ville threshold, the pret0 penalty, and the learner’s regret. Writing Aγ,t0 := log(1/γ) + 2(t0 − 1), we may equivalently n
o
κT ctx express nctx γ (r) = inf n ≥ t0 : Aγ,t0 + rT ≤ 2 for every T ≥ n . Thus nγ (r) < ∞
n
t0 such that supT ≥n rT − κT 2
o
⇐⇒
∃n ≥
≤ −Aγ,t0 . A simple sufficient condition is lim supT →∞ rTT < κ2 . In
particular, any online learner with sublinear regret rT = o(T ) has nctx γ (r) < ∞. This is why standard OCO algorithms with sublinear adversarial regret can be used as contextual skeptics: once their learning cost is asymptotically smaller than the cumulative predictable edge, Theorem 3.2 enters the exponential-power regime. While our theoretical framework so far is agnostic to the choice of the OCO algorithm, we summarize the update rules for common OCO algorithms (namely, OGD, FTLR, and ONS) in Table B.1 of the Appendix B.2. The experimental analysis in the following section uses various variants of these standard OCO algorithms. 11
Remark 3.4 (Beyond uniform edge alternative (1)). The uniform-edge alternative in Definition 3.1 is a special case of a more general cumulative-edge alternative studied in Appendix C.7. The cumulative alternative Qϕcum (Dκ,t0 ; Φ) is defined by requiring that, for some comparator θ⋆ ∈ K, the cumulative predictable edge
Q ⋆ t=1 bt (θ ) is lower bounded by a deterministic envelope DT . We
PT
ϕ κ,t0 ; Φ) ⊆ Qϕ (ref. Prop. C.6, Remark C.10). The corresponding then have Qϕ,lin κ,t0 (Φ) ⊆ Qcum (D
regret-to-power theorem (Thm. C.8) shows that power only requires cumulative feature-aligned signal to dominate the Ville threshold and the learner’s regret; it need not be uniformly positive at every time.
4
Experiments
Our experiments are organized around the validity–power separation developed above. We ask three questions. First, can a coarser skeptic be valid but blind when the hit process is non-i.i.d. and the violation is visible only through richer information? Second, under a correctly specified oracle forecaster, do the betting processes respect the Ville-prescribed Type-I level? Third, when the oracle is replaced by a black-box quantile forecaster, do feature-aware skeptics detect conditional miscalibration missed by marginal audits? Throughout, for a test level γtest = 0.05, we record a rejection if τγtest ≤ T , equivalently if the running wealth crosses the anytime-valid threshold 1/γtest = 20 by the end of the horizon. Sticky-coin: marginal vs. conditional null. We first remove the forecasting layer and feed the skeptic a controlled hit process. Specifically, we use the two-state Markov chain of Christoffersen and Pelletier (2004), reparametrized so that the hit indicator Bt = 1{Yt ≤ q̂t } has marginal probability α and one-step autocorrelation δ. In this experiment, the pair (Yt , q̂t ) is not modeled separately; instead, the centered hit Zt = Bt − α is generated directly from (α, δ). This isolates the effect of monitoring information on power. For each cell of the (α, δ) grid, we run N Monte Carlo trials of length T and average the rejection time τ̄ . Figure 3 shows the gap predicted by Proposition 2.8: the intercept-only marginal skeptic is censored at T across the valid region, because it cannot exploit one-step dependence, whereas the Zt−1 -aware skeptic rejects quickly across most of the grid. Validity under the null. We next instantiate the full forecast-and-bet protocol under a calibrated synthetic oracle. The data are N =1,000 count series over T =1,460 days from the Negative-Binomial panel described in Appendix D.1. At each quantile level α ∈ {0.01, 0.02, . . . , 0.99},
12
Figure 3: Mean rejection time τ̄ over the (α, δ) grid for the AdaGrad OCO skeptic (N =50 runs, T =1,460, γtest =0.05). Censored cells, where no rejection occurs by time T , appear dark blue; faster rejections appear light blue; grey marks invalid parameter pairs. The marginal skeptic is censored across the valid region, while the Zt−1 -aware skeptic rejects quickly. the oracle issues the corresponding Negative-Binomial forecast under the true parameters. Thus any rejection in this experiment is a false alarm for the intended calibration benchmark. For each pair of series and quantile level, we run the betting audit with the same feature configurations used below and reject when τγtest ≤ T . Figure 4 shows that the empirical rejection rates track the nominal 5% Ville level across quantile levels and skeptic configurations. This confirms the main validity claim in a fully specified forecast-and-outcome setting before replacing the oracle by a black-box forecaster. Experiments with a quantile forecaster. Finally, we replace the calibrated oracle with Chronos-2 forecasts (Ansari et al., 2025). We evaluate two panels. The first reuses the synthetic
rejection rate (% of series)
Negative-Binomial series from the validity experiment, but now the forecasted quantiles are produced 10%
10%
8%
8% AdaGrad
ONS
6%
6%
4%
4%
2%
2%
0% 0
0.2
0.4 0.6 0.8 α (quantile level tested) full marginal
z_prev sin_w
0% 0
0.2
cos_w promo
log_price sin_y
1
0.4 0.6 0.8 α (quantile level tested)
1
cos_y g
Figure 4: Type-I rejection rate versus quantile level α at γtest = 0.05 for the AdaGrad and ONS skeptics on N =1,000 oracle series (T =1,460). The grey band is a one-sided 99% Monte Carlo tolerance band. SGD and FTRL show the same pattern and are reported in Appendix D.2 (Figure 6).
13
by Chronos-2 rather than by the true conditional law. The second is the Rossmann store-sales dataset (Cukierski, 2015), where the true conditional law is unknown. In both cases, the goal is no longer to measure Type-I error under a known null, but to ask which monitoring views expose forecast miscalibration. For each quantile level, we run marginal, single-feature, and full-feature skeptics using only information available before Yt is observed. The marginal skeptic uses an intercept-only betting direction; a single-feature skeptic uses one coordinate of the predictable feature dictionary, such as promotion status, Saturday, or log-price; and the full-feature skeptic uses the full dictionary. Figure 5 reports per-feature rejection rates. On both datasets, feature-aware skeptics reject far more often than the marginal skeptic. The strongest signals are promotion status on both panels, Saturday on Rossmann, and log-price on the synthetic data, illustrating that the learned evidence is not only more powerful but also feature-diagnostic. Rossmann · chronos2 · sliding_365
Synthetic · chronos2 · sliding_365
1,115 stores · ftrl (lr0.1-l11.0-shrink0.5)
1,000 SKUs · ftrl (lr0.1-l10.1-shrink1.0)
rejection rate
100%
75%
50%
25%
5%
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
α (quantile level) marginal
promo
is_saturday
log_price
full
Figure 5: Per-feature rejection rate versus quantile level α at γtest = 0.05 for FTRL betting on Chronos-2 forecasts. Left: Rossmann store-sales data with N =1,115 stores. Right: synthetic Negative-Binomial panel with N =1,000 series. Each highlighted curve is one skeptic configuration; faded grey curves are the remaining single-feature skeptics. The dashed line marks the Ville 5% reference level from the calibrated-oracle experiment (Figure 4). Feature-aware skeptics reveal conditional miscalibration that is largely missed by marginal audits, most visibly through promo, is_saturday, and log_price.
5
Conclusion
We developed an information-aware testing-by-betting framework for auditing black-box conditional quantile forecasters. For any monitoring information available to the auditor, predictable nobankruptcy bets generate anytime-valid e-processes for the corresponding calibration null, without
14
assuming independence, stationarity, parametric structure, or moment conditions on the outcomes. The resulting null hierarchy separates validity from power: coarser audits remain safe, but can be blind to violations visible only through richer features. To gain feature-specific power, we introduced contextual betting, in which an online learner adapts linear bets over the auditor’s predictable feature dictionary. Under persistent, or more generally cumulative, feature-aligned predictable edge, pathwise regret bounds translate into finite-time detection and stopping-time guarantees. Empirically, the method controls Type-I error under a calibrated oracle and detects conditional miscalibration missed by marginal audits on controlled non-i.i.d. streams and on Chronos-2 forecasts for synthetic and Rossmann store-sales data. Several directions remain open. Our contextual bets are linear in the supplied feature dictionary and might consequently fail to be sufficiently representative of miscalibrations in high dimensions. Therefore, extending the framework to nonlinear betting classes (e.g., kernelized or neural skeptics) while preserving validity and obtaining useful regret-topower guarantees is an important next step. Our power guarantees are necessarily feature-relative: if the auditor does not observe features that expose the violation, the corresponding audit may remain valid but powerless. The current framework also tests one quantile level at a time; aggregating evidence across quantiles, or comparing two forecasters head-to-head, are natural extensions of the same betting machinery. Finally, our experiments focus on count-valued demand data and one foundation model forecaster; broader benchmarks across continuous, multivariate, and non-retail settings, as well as elicitable functionals beyond quantiles, are left for future work.
References Anastasios N. Angelopoulos and Stephen Bates. Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning, 16(4):494–591, 2023. doi: 10.1561/2200000101. Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken, Andreas Auer, Boran Han, Pedro Mercado, Syama Sundar Rangapuram, Huibin Shen, Lorenzo Stella, Xiyuan Zhang, et al. Chronos-2: From univariate to universal forecasting. arXiv preprint arXiv:2510.15821, 2025. Sebastian Arnold, Alexander Henzi, and Johanna F. Ziegel. Sequentially valid tests for forecast calibration. The Annals of Applied Statistics, 17(3):1909–1935, 2023. doi: 10.1214/22-AOAS1697. Ying Cao and Zuo-Jun Max Shen. Quantile forecasting and data-driven inventory management under nonstationary demand. Operations Research Letters, 47(6):465–472, 2019. Philippe Casgrain, Martin Larsson, and Johanna F. Ziegel. Sequential testing for elicitable functionals via supermartingales. Bernoulli, 30(2):1347–1374, 2024. doi: 10.3150/23-BEJ1634. Peter Christoffersen and Denis Pelletier. Backtesting value-at-risk: A duration-based approach. Journal of Financial Econometrics, 2(1):84–108, 2004. 15
Peter F. Christoffersen. Evaluating interval forecasts. International Economic Review, 39(4):841–862, 1998. doi: 10.2307/2527341. Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. Wiley-Interscience, 2 edition, 2006. ISBN 978-0-471-24195-9. Will Cukierski. Rossmann store sales, 2015. URL https://www.kaggle.com/c/ rossmann-store-sales. Licensed under the Open Data Commons Open Database License (ODbL) v1.0. A. Philip Dawid and Vladimir G. Vovk. Prequential probability: Principles and properties. Bernoulli, 5(1):125–162, 1999. ISSN 13507265. URL http://www.jstor.org/stable/3318616. Francis X. Diebold, Todd A. Gunther, and Anthony S. Tay. Evaluating density forecasts with applications to financial risk management. International Economic Review, 39(4):863–883, 1998. doi: 10.2307/2527342. Colin Doms, Sarah C Kramer, and Jeffrey Shaman. Assessing the use of influenza forecasts and epidemiological modeling in public health decision making in the united states. Scientific reports, 8(1):12406, 2018. Robert F Engle and Simone Manganelli. Caviar: Conditional autoregressive value at risk by regression quantiles. Journal of business & economic statistics, 22(4):367–381, 2004. Tobias Fissler and Johanna F Ziegel. Higher order elicitability and Osband’s principle. Annals of Statistics, 44(4):1680–1707, 2016. Isaac Gibbs and Emmanuel J. Candès. Adaptive conformal inference under distribution shift. In Advances in Neural Information Processing Systems, volume 34, 2021. Tilmann Gneiting. Making and evaluating point forecasts. Journal of the American Statistical Association, 106(494):746–762, 2011. Tilmann Gneiting, Fadoua Balabdaoui, and Adrian E. Raftery. Probabilistic forecasts, calibration and sharpness. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 69(2): 243–268, 2007. doi: 10.1111/j.1467-9868.2007.00587.x. Peter Grünwald, Rianne de Heide, and Wouter Koolen. Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86(5):1091–1128, 11 2024. ISSN 1369-7412. doi: 10.1093/jrsssb/qkae011. URL https://doi.org/10.1093/jrsssb/qkae011. Elad Hazan, Amit Agarwal, and Satyen Kale. Logarithmic regret algorithms for online convex optimization. Machine Learning, 69(2):169–192, 2007. Yannick Hoga and Matei Demetrescu. Monitoring value-at-risk and expected shortfall forecasts. Management Science, 69(5):2954–2971, 2023. Steven R. Howard and Aaditya Ramdas. Sequential estimation of quantiles with applications to A/B testing and best-arm identification. Bernoulli, 28(3):1704–1728, 2022. doi: 10.3150/21-BEJ1388. Roger Koenker and Gilbert Bassett Jr. Regression quantiles. Econometrica: journal of the Econometric Society, pages 33–50, 1978.
16
S. Kullback and R. A. Leibler. On information and sufficiency. The Annals of Mathematical Statistics, 22(1):79–86, 1951. doi: 10.1214/aoms/1177729694. Chenghao Liu, Taha Aksu, Juncheng Liu, Xu Liu, Hanshu Yan, Quang Pham, Silvio Savarese, Doyen Sahoo, Caiming Xiong, and Junnan Li. Moirai 2.0: When less is more for time series forecasting. arXiv preprint arXiv:2511.11698, 2025. Paul Mineiro and Steven Howard. Time-uniform confidence bands for the CDF under nonstationarity. In Advances in Neural Information Processing Systems, volume 36, 2023. Jacob Montiel, Max Halford, Saulo Martiello Mastelini, Geoffrey Bolmier, Raphael Sourty, Robin Vaysse, Adil Zouitine, Heitor Murilo Gomes, Jesse Read, Talel Abdessalem, et al. River: machine learning for streaming data in python. 2021. Francesco Orabona and Dávid Pál. Coin betting and parameter-free online learning. Advances in Neural Information Processing Systems, 29, 2016. Aaditya Ramdas and Ruodu Wang. Hypothesis testing with e-values. Foundations and Trends® in Statistics, 1(1-2):1–390, 2025. ISSN 2978-4212. doi: 10.1561/3600000002. Aaditya Ramdas, Johannes Ruf, Martin Larsson, and Wouter M. Koolen. Testing exchangeability: Fork-convexity, supermartingales and e-processes. International Journal of Approximate Reasoning, 141:83–109, 2022. doi: 10.1016/j.ijar.2021.06.017. Aaditya Ramdas, Peter Grünwald, Vladimir Vovk, and Glenn Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38(4):576–601, 2023. doi: 10.1214/23-STS894. Yaniv Romano, Evan Patterson, and Emmanuel J. Candès. Conformalized quantile regression. In Advances in Neural Information Processing Systems, volume 32, 2019. Glenn Shafer. Testing by betting: A strategy for statistical and scientific communication. Journal of the Royal Statistical Society: Series A, 184(2):407–431, 2021. Glenn Shafer and Vladimir Vovk. Game-theoretic foundations for probability and finance. John Wiley & Sons, 2019. Glenn Shafer, Alexander Shen, Nikolai Vereshchagin, and Vladimir Vovk. Test martingales, bayes factors and p-values. Statistical Science, 26(1):84–101, 2011. doi: 10.1214/10-STS347. Shai Shalev-Shwartz. Online learning and online convex optimization. Foundations and Trends in Machine Learning, 4(2):107–194, 2012. doi: 10.1561/2200000018. Jean Ville. Étude Critique de la Notion de Collectif. Gauthier-Villars, Paris, 1939. Vladimir Vovk and Ruodu Wang. E-values: Calibration, combination, and applications. The Annals of Statistics, 49(3):1736–1754, 2021. doi: 10.1214/20-AOS2020. Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World. Springer, New York, 2005. Abraham Wald. Sequential tests of statistical hypotheses. The Annals of Mathematical Statistics, 16(2):117–186, 1945. doi: 10.1214/aoms/1177731118.
17
Qiuqi Wang, Ruodu Wang, and Johanna F. Ziegel. E-backtesting. Management Science, 2025. doi: 10.1287/mnsc.2023.01659. Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86(1):1–27, 2024. doi: 10.1093/jrsssb/qkad009. Qingchuan Yang, Simon Mahns, Sida Li, Anri Gu, Jibang Wu, and Haifeng Xu. Llm-as-a-prophet: Understanding predictive intelligence with prophet arena. arXiv preprint arXiv:2510.17638, 2025. Martin Zinkevich. Online convex programming and generalized infinitesimal gradient ascent. In Proceedings of the 20th international conference on machine learning (icml-03), pages 928–936, 2003.
18
A
Related work
Forecast evaluation and conditional calibration. Classical forecast evaluation already identifies the core failure mode motivating our work: forecasts can be calibrated on average while being predictably wrong after conditioning on information available at forecast time. For interval forecasts, Christoffersen (1998) showed that unconditional coverage alone is inadequate when violations cluster dynamically. For density and distributional forecasts, PIT diagnostics and calibration taxonomies clarify how marginal, exceedance, and probabilistic calibration differ (Diebold et al., 1998; Gneiting et al., 2007). In risk management, the analogous single-quantile problem is value-at-risk backtesting (Christoffersen and Pelletier, 2004; Engle and Manganelli, 2004). Our work studies the sharp conditional quantile version of this program. Rather than evaluating an entire predictive distribution, we audit a deployed black-box quantile forecast qbt through the hit Zt = 1{Yt ≤ qbt } − α. This is narrower than full distributional calibration, but it is exactly the property relevant to one-tail decisions such as inventory control and VaR-based risk management. The main novelty is that we index the calibration null by the auditor’s monitoring information, thereby separating marginal validity from feature-specific detectability. Safe anytime-valid inference and sequential quantile methods. Classical backtests are typically fixed-horizon, whereas deployed forecasting systems are monitored continuously. Safe anytime-valid inference addresses this problem using nonnegative test supermartingales, e-values, and e-processes, whose validity is preserved under optional stopping (Ville, 1939; Shafer et al., 2011; Shafer, 2021; Vovk and Wang, 2021; Ramdas et al., 2022, 2023; Grünwald et al., 2024; Ramdas and Wang, 2025). This viewpoint is natural in prequential forecast evaluation, where forecasts and outcomes arrive sequentially (Dawid and Vovk, 1999; Shafer and Vovk, 2019). Several recent works apply this methodology to quantiles, CDFs, or forecast calibration, but with different statistical targets. Howard and Ramdas (2022) construct time-uniform confidence sequences for population quantiles from an i.i.d. stream, with applications to A/B testing and best-arm identification. Their object is quantile estimation from samples; our object is testing whether a random, previously issued forecast qbt satisfies P (Yt ≤ qbt | Ht−1 ) = α. This extra layer is crucial: even if one can estimate a population quantile sequentially, auditing a black-box conditional quantile forecaster requires testing the conditional hit probability of the issued forecasts. Mineiro and Howard (2023) extend time- and value-uniform CDF inference to nonstationary streams by targeting the running average of conditional distributions. This is complementary to our work: they 19
estimate an entire changing CDF, whereas we test a forecast-specific conditional calibration null and study feature-aware power. Arnold et al. (2023) develop sequentially valid tests for calibration of probabilistic forecasts, closely related to PIT/probabilistic calibration. This is of separate interest, but it requires distributional forecasts and is not directly recoverable from quantile forecasts alone. Elicitable functionals, sequential testing, and e-backtesting.
Quantiles are elicitable
and identifiable functionals: the pinball loss is strictly consistent, and the identification function Vα (y, q) = 1{y ≤ q} − α vanishes at an α-quantile (Koenker and Bassett Jr, 1978; Gneiting, 2011; Fissler and Ziegel, 2016). The work closest to ours at the methodological level is Casgrain et al. (2024), who develop sequential tests for general elicitable and identifiable functionals via supermartingales and predictable mixing, and use OCO regret to obtain power guarantees. Our test can be viewed as a quantile-hit specialization of this general paradigm, but with a different null and a different power question. We audit random, covariate-dependent forecasts qbt and explicitly allow the skeptic’s monitoring information H to be a strict sub-filtration of the full pre-outcome information. This leads to a hierarchy of nulls and to feature-specific power statements: a coarser valid test can be powerless against violations visible only through richer features. In contrast, the standard elicitable-functional setup conditions on the full available filtration. Moreover, by specializing to the bounded two-point hit process Zt ∈ {−α, 1 − α}, our validity theory requires no moment, stationarity, parametric, or tail assumptions on Yt . The e-backtesting framework of Wang et al. (2025) is also closely related. They introduce backtest e-statistics and construct e-processes for risk measure forecasts, especially VaR and ES. Their main motivation is that VaR testing alone is insufficient for regulatory market-risk backtesting, and that ES forecasts require model-free anytime-valid tests. Our overlap with their VaR case is mathematical but not conceptual: the VaR e-statistic 1{Yt > qbt }/(1 − α) corresponds to one endpoint of our no-bankruptcy betting interval. The main distinctions are that e-backtesting uses the full market information filtration and focuses on one-sided risk-measure underestimation, whereas we index the null by the auditor’s information set and study how missing or available features determine power. Their optimality and power analysis is developed for settings such as i.i.d. losses or i.i.d. e-statistics, while our main power result is a finite-time regret-to-power guarantee for non-i.i.d., history-dependent streams with persistent feature-aligned predictable edge.
20
Conformal prediction, monitoring, and online learning.
Our goal is also distinct from con-
formal prediction. Conformal prediction and conformalized quantile regression construct prediction sets or intervals with coverage guarantees (Vovk et al., 2005; Romano et al., 2019; Angelopoulos and Bates, 2023), and adaptive conformal methods update these constructions under distribution shift to maintain long-run coverage frequencies (Gibbs and Candès, 2021). We instead audit a fixed deployed forecaster: we do not wrap, recalibrate, or modify its predictions. The question is whether its realized errors are conditionally exploitable by an auditor with a specified information set. Finally, our contextual betting procedures use online convex optimization to learn feature-aware bets. Classical OCO algorithms such as projected online gradient descent, follow-the-regularizedleader, online Newton step, and coin-betting provide regret guarantees for adapting decisions from sequential data (Zinkevich, 2003; Hazan et al., 2007; Shalev-Shwartz, 2012; Orabona and Pál, 2016). In our setting, regret has a statistical interpretation: it is the finite learning cost paid before the skeptic can exploit feature-aligned miscalibration. Thus the paper ties together classical conditional forecast evaluation, anytime-valid testing, elicitable-functional martingales, and contextual online learning into a feature-aware audit for black-box conditional quantile forecasters.
B
Methodological details
B.1
Game protocol
Definition B.1 (Protocol for conditional quantile calibration game). Skeptic starts with wealth M0 = 1. For each round t = 1, 2, . . ., the game proceeds as follows: 1. The predictable context Ct , assumed Ft−1 -measurable, is revealed and available to Forecaster and Skeptic. 2. Forecaster announces a real-valued forecast q̂t ∈ R, which is assumed to be Ft−1 -measurable. 3. Skeptic chooses a betting multiplier λt ∈ Λα := [−1/(1 − α), 1/α], assumed Ft−1 -measurable. 4. Reality reveals the outcome Yt , assumed Ft -measurable. 5. Skeptic’s wealth is updated by Mt = Mt−1 (1 + λt Zt ). The process (Mt ) has the betting interpretation that the skeptic starts with wealth one and repeatedly stakes a predictable fraction against the forecaster’s asserted hit probability. Under P0 , each legal one-step game is conditionally fair. 21
B.2
OCO algorithms
In this subsection, we record three standard OCO algorithms, whose variants are used in our experimental analysis. We also introduce the regret rates for the corresponding algorithms from standard sources. Throughout, the learner chooses θt ∈ K before observing Zt , uses the betting fraction λt = ⟨θt , ϕt ⟩, and then observes the convex loss gt (θ) := − log (1 + ⟨θ, ϕt ⟩Zt ) . On K = {θ : ∥θ∥2 ≤ 1/(2R)}, the denominator in ∇gt is at least 1/2, so ∥∇gt (θ)∥2 = −
Zt ϕt ≤ 2R. 1 + ⟨θ, ϕt ⟩Zt 2
Thus the losses are convex and uniformly Lipschitz on K. Moreover, exp(−gt (θ)) = 1 + ⟨θ, ϕt ⟩Zt is affine and positive on K, so gt is exp-concave on K. Online gradient descent. Projected online gradient descent updates in the direction of the negative observed gradient and projects back to K:
OGD θt+1 = ΠK θtOGD − ηt ∇gt (θtOGD ) ,
where ΠK is Euclidean projection. For convex Lipschitz losses over a bounded Euclidean domain, √ projected OGD achieves O( T ) static regret against the best fixed comparator in hindsight (Zinkevich, 2003, Theorem 1); see also Shalev-Shwartz (2012, Corollary 2.7). Follow-the-regularized-leader.
Euclidean FTRL chooses the next betting parameter by mini-
mizing past losses plus a stabilizing quadratic regularizer: FTRL θt+1 = arg min θ∈K
( t X
1 gi (θ) + ∥θ∥22 . 2η i=1 )
With η = Θ(T −1/2 ) when the horizon is known, or with a standard doubling trick when it is not, √ Euclidean FTRL has O( T ) regret for convex Lipschitz losses. This follows from the strongly convex regularizer analysis of Follow-the-Regularized-Leader (Shalev-Shwartz, 2012, Theorem 2.11 and Corollary 2.12).
22
Algorithm
Regret guarantee √ Projected OGD (Zinkevich, RTctx (θ) = O( T ) for 2003, Theorem 1); (Shalev- convex Lipschitz losses Shwartz, 2012, Corollary 2.7)
Contextual betting update θt+1 = ΠK (θt − ηt ∇gt (θt )) .
√ Euclidean FTRL (Shalev- RTctx (θ) = O( T ) for Shwartz, 2012, Theorem 2.11 convex Lipschitz losses and Corollary 2.12)
θt+1 = arg min θ∈K
Online Newton Step (Hazan RTctx (θ) = O(d log T ) for et al., 2007, Theorem 2) exp-concave losses
( t X
1 gi (θ) + ∥θ∥22 2η i=1
) .
At = At−1 + ∇t ∇⊤ t , −1 t θt+1 = ΠA K θt − ηAt ∇t .
Table B.1: Standard OCO algorithm updates for learning contextual betting directions. Here gt (θ) = − log(1 + ⟨θ, ϕt ⟩Zt ), ∇t = ∇gt (θt ), and the next-round bet is given by λt+1 = ⟨θt+1 , ϕt+1 ⟩. Online Newton Step.
Because the betting losses are exp-concave, one can exploit curvature
using Online Newton Step. With At = At−1 + ∇t ∇⊤ t ,
∇t := ∇gt (θtONS ), the update is
ONS ONS t θt+1 = ΠA − ηA−1 t ∇t , K θt ⊤ where ΠA K (y) := arg minθ∈K (θ − y) A(θ − y). For exp-concave losses with bounded gradients over a
bounded domain, ONS achieves logarithmic regret, specifically O(d log T ) in dimension d (Hazan et al., 2007, Theorem 2). Our proof instantiates this general result for the contextual betting loss and gives the explicit bound used in the main text. The implication of the regret rates presented in Table B.1 for our stopping time result in Theorem 3.2 is immediate: all three algorithms have sublinear regret in the adversarial OCO sense, and hence satisfy lim supT →∞ rTT = 0 < κ2 for every κ > 0. Therefore the certification time nctx γ (r) is finite for each of these learners. The difference is quantitative: OGD and Euclidean FTRL give √ T learning cost, whereas ONS gives logarithmic learning cost for the exp-concave betting losses.
23
C
Theoretical details
C.1
Missing definitions
Definition C.1 (Conditional e-variables). A nonnegative Ft -measurable random variable Et is a conditional e-variable for P0 at time t if, for every P ∈ P0 , EP [Et | Ft−1 ] ≤ 1 P -a.s. If (Et )t≥1 is a sequence of conditional e-variables and Mt :=
i=1 Ei with M0 = 1, then (Mt )t≥0 is a test
Qt
supermartingale, since EP [Mt | Ft−1 ] ≤ Mt−1 .
C.2
Supporting lemmas
Lemma C.2 (Conditional Hoeffding bound with support-length control). Let (mt )t≥1 be a martingale-difference sequence with respect to a filtration (Ft )t≥0 ; that is, each mt is Ft -measurable, integrable, and E[mt | Ft−1 ] = 0 a.s. Suppose there exist deterministic constants ct ≥ 0 such that, for each t, there are Ft−1 -measurable random variables at and bt satisfying at ≤ mt ≤ bt a.s. and bt − at ≤ ct a.s. Let VT :=
P
T X t=1
2 t=1 ct . Then, for every T ≥ 1 and x > 0,
PT !
mt ≤ −x
2x2 ≤ exp − VT
!
,
P
T X
!
mt ≥ x
t=1
2x2 ≤ 2 exp − VT
!
,
with the convention that the right-hand side is 0 when VT = 0. Proof of Lemma C.2. If VT = 0, then ct = 0 for every t ≤ T . Hence bt − at = 0 a.s. for every t ≤ T , so mt = at = bt a.s.; since E[mt | Ft−1 ] = 0, it follows that mt = 0 a.s. for every t ≤ T . The claimed inequalities are then trivial, so assume VT > 0. Fix t and η ∈ R. Conditional on Ft−1 , the random variable mt has conditional mean zero and is supported in [at , bt ], whose length is at most ct . Conditional Hoeffding’s lemma gives E[eηmt | Ft−1 ] ≤ exp(η 2 c2t /8) a.s. For η ∈ R, define t X
t η2 X Lt (η) := exp η mi − c2i , 8 i=1 i=1
!
L0 (η) := 1.
Then (Lt (η))Tt=0 is a nonnegative supermartingale, because η 2 c2t E[Lt (η) | Ft−1 ] = Lt−1 (η) exp − 8 Consequently E[LT (η)] ≤ 1, equivalently E exp(η
E[eηmt | Ft−1 ] ≤ Lt−1 (η).
2 t=1 mt ) ≤ exp(η VT /8).
PT
24
!
For the lower tail, take η > 0 and apply Markov’s inequality to exp(−η
P
T X
!
mt ≤ −x
−ηx
≤e
E exp −η
t=1
T X
!
mt
t=1
Optimizing over η > 0 gives η ⋆ = 4x/VT , and therefore P(
t=1 mt ):
PT
η 2 VT ≤ exp −ηx + 8
!
.
2 t=1 mt ≤ −x) ≤ exp(−2x /VT ).
PT
The same argument applied to −mt gives P(
2 t=1 mt ≥ x) ≤ exp(−2x /VT ).
PT
upper and lower tail bounds by the union bound yields P(|
C.3
Combining the
2 t=1 mt | ≥ x) ≤ 2 exp(−2x /VT ).
PT
Missing results
Proposition C.3 (Bounded martingale difference sequence property). If P ∈ P0 (H), then (Zt )t≥1 is a bounded martingale difference sequence (MDS) with respect to H: for every t ≥ 1, Zt is Ht -measurable, Zt ∈ {−α, 1 − α} P -a.s., and EP [Zt | Ht−1 ] = 0 P -a.s. Proof of Proposition C.3. Fix t ≥ 1. By construction, Zt = 1{Yt ≤ qbt } − α. The measurability of Zt follows from the Ht−1 -measurability of q̂t , the Ht -measurability of Yt , and the inclusion Ht−1 ⊆ Ht Since 1{Yt ≤ qbt } is binary, Zt ∈ {−α, 1 − α} and is bounded. Finally, for P ∈ P0 (H), EP [Zt | Ht−1 ] = P (Yt ≤ q̂t | Ht−1 ) − α = 0 P -a.s. for every t ≥ 1. Thus (Zt )t≥1 is a bounded martingale difference sequence with respect to H. Proposition C.4 (Oracle Kelly betting strategy). For each t, the extended-real maximizer of pt −α ψt over Λα is λ⋆t = α(1−α) . If pt ∈ (0, 1), this maximizer lies in ri(Λα ) = (−1/(1 − α), 1/α) and
both endpoints have value −∞. If pt = 0 or pt = 1, the same formula gives the corresponding no-bankruptcy endpoint. In all cases, ψt (λ⋆t ) = KL(Bern(pt ) ∥ Bern(α)). Proof of Proposition C.4. For a fixed stake λ, define the conditional expected log-growth ψt (λ) := E[log(1 + λZt ) | Ft−1 ] = pt log{1 + λ(1 − α)} + (1 − pt ) log(1 − αλ), with conventions 0 log 0 := 0 and a log 0 := −∞ for a > 0. Fix t and write p := pt . On ri(Λα ), ψt is differentiable with derivative p(1−α) ′ ψt′ (λ) = 1+λ(1−α) − (1−p)α 1−αλ . If p ∈ (0, 1), then ψt is strictly concave on ri(Λα ), and solving ψt (λ) = 0
gives λ = (p − α)/{α(1 − α)}, which lies in ri(Λα ). At either endpoint one possible wealth multiplier is zero; since both outcomes have positive probability, the expected log-growth there is −∞. If p = 0, then ψt (λ) = log(1 − αλ) is decreasing and is maximized at −1/(1 − α). If p = 1, then ψt (λ) = log{1 + λ(1 − α)} is increasing and is maximized at 1/α. Both endpoint cases agree with (C.4). Substituting (C.4) gives 1 + λ⋆t (1 − α) = p/α and 1 − αλ⋆t = (1 − p)/(1 − α), hence ψt (λ⋆t ) = p log(p/α) + (1 − p) log{(1 − p)/(1 − α)} = KL(Bern(p) ∥ Bern(α)). 25
Lemma C.5 (Quadratic lower bound on full-feature log-growth). For every law Q, every θ ∈ K, and every t ≥ 1, h
i
ϕ EQ log(1 + ⟨θ, ϕt ⟩Zt ) Ht−1 ≥ bQ t (θ)
Q-a.s.
ϕ Moreover, if bQ t (θ) > 0 on an event of positive probability, then mt ̸= 0 on an event of positive
probability. Hence any law admitting a uniformly positive full-feature quadratic edge violates P0ϕ . Proof. Set λt = ⟨θ, ϕt ⟩. Since |λt Zt | ≤ 1/2, the elementary inequality log(1 + x) ≥ x − x2 for ϕ x ∈ [−1/2, 1/2] gives log(1 + λt Zt ) ≥ λt Zt − λ2t Zt2 . Taking conditional expectation given Ht−1 , and ϕ using that λt is Ht−1 -measurable, yields
h
i
ϕ ϕ ϕ EQ log(1 + λt Zt ) Ht−1 ≥ λt EQ [Zt | Ht−1 ] − λ2t EQ [Zt2 | Ht−1 ]
≥ λt mϕt − λ2t , ϕ Q Q 2 because Zt2 ≤ 1. This is exactly bQ t (θ). If mt = 0, then bt (θ) = −⟨θ, ϕt ⟩ ≤ 0. Therefore bt (θ) > 0
implies mϕt ̸= 0. A uniformly positive edge therefore implies failure of the full-feature calibration null. Proposition C.6 (Linear full-feature alternatives are full-feature alternatives). For every κ > 0 and ϕ t0 ∈ N, we have that linear predictable-edge alternative is a subclass of Qϕ , i.e., Qϕ,lin κ,t0 (Φ) ⊆ P \ P0 . Q ⋆ ⋆ Proof of Proposition C.6. Let Q ∈ Qϕ,lin κ,t0 (Φ). Then there exists θ ∈ K such that bt (θ ) ≥ κ > 0 Qϕ ϕ ϕ ⋆ a.s. for every t ≥ t0 . By Lemma C.5, bQ t (θ ) > 0 implies mt ̸= 0. Therefore EQ [Zt | Ht−1 ] = mt ̸= 0
with positive probability for every t ≥ t0 . Hence Q ∈ / P0ϕ . Definition C.7 (Cumulative full-feature edge alternative). Fix a predictable feature sequence Φ = (ϕt )t≥1 and a deterministic real sequence D = (DT )T ≥1 . For θ ∈ K, define the predictable cumulative full-feature edge BTQ (θ) :=
Q Q ϕ ϕ 2 t=1 bt (θ), where bt (θ) := ⟨θ, ϕt ⟩mt − ⟨θ, ϕt ⟩ , and mt :=
PT
ϕ EQ [Zt | Ht−1 ]. The cumulative full-feature edge alternative with lower envelope D is
n
Qϕcum (D; Φ) := Q : ∃θ⋆ ∈ K such that BTQ (θ⋆ ) ≥ DT
o
Q-a.s. for every T ≥ 1 .
Theorem C.8 (Full-feature contextual regret-to-power under cumulative edge). Assume that the ϕ online learner outputs Ht−1 -measurable vectors θt ∈ K and satisfies the pathwise regret bound
RTctx (θ) ≤ rT for every θ ∈ K and every T ≥ 1, where (rT )T ≥1 is deterministic. Define τγctx := inf{T ≥ 1 : MTctx ≥ 1/γ}. Fix γ ∈ (0, 1) and a deterministic lower envelope D = (DT )T ≥1 . Suppose 26
Q ∈ Qϕcum (D; Φ). Define sT (D) := DT − rT − log(1/γ). Then, for every T ≥ 1 such that sT (D) > 0, we have Q(τγctx > T ) ≤ exp
2sT (D)2 − T
!
.
Proof of Theorem C.8. Set λ⋆t := ⟨θ⋆ , ϕt ⟩. By regret against θ⋆ , log MTctx = −
T X
gt (θt ) ≥ −
T X
t=1
gt (θ⋆ ) − rT =
t=1
T X
log(1 + λ⋆t Zt ) − rT .
t=1
Because |λ⋆t Zt | ≤ 1/2, log(1 + λ⋆t Zt ) ≥ λ⋆t Zt − (λ⋆t )2 , and therefore log MTctx ≥
⋆ ⋆ 2 t=1 {λt Zt − (λt ) } −
PT
rT . ϕ Define the martingale difference ∆St := λ⋆t Zt − EQ [λ⋆t Zt | Ht−1 ] and set ST :=
t=1 ∆St . The
PT
ϕ sequence (∆St )t≥1 is a martingale-difference sequence with respect to Hϕ , since EQ [∆St | Ht−1 ]= ϕ ϕ ϕ ϕ EQ [λ⋆t Zt − EQ [λ⋆t Zt | Ht−1 ] | Ht−1 ] = 0. Since λ⋆t is Ht−1 -measurable, EQ [λ⋆t Zt | Ht−1 ] = λ⋆t mϕt .
Hence
T X
{λ⋆t Zt − (λ⋆t )2 } = ST +
t=1
Thus log MTctx ≥ ST +
Q ⋆ t=1 bt (θ ) − rT
PT
T X
T X Q
t=1
t=1
{λ⋆t mϕt − (λ⋆t )2 } = ST +
bt (θ⋆ ).
≥ ST + DT − rT . On the event {τγctx > T }, one has
log MTctx < log(1/γ) due to monotonocity of the logarithm, and so ST < log(1/γ) + rT − DT = −sT (D). Therefore {τγctx > T } ⊆ {ST < −sT (D)}. ϕ It remains to control the lower tail of ST . Conditional on Ht−1 , the variable λ⋆t Zt takes values
in {−αλ⋆t , (1 − α)λ⋆t }, and therefore its conditional support has length |(1 − α)λ⋆t − (−αλ⋆t )| = |λ⋆t |. ϕ Centering by EQ [λ⋆t Zt | Ht−1 ] only translates this conditional support, so the centered increment
∆St also has conditional support length |λ⋆t |. Since θ⋆ ∈ K and ∥ϕt ∥2 ≤ R, |λ⋆t | = |⟨θ⋆ , ϕt ⟩| ≤ 1 ∥θ⋆ ∥2 ∥ϕt ∥2 ≤ 2R R = 12 ≤ 1.
Applying Hoeffding–Azuma’s inequality for martingales with conditional range lengths (Lemma
C.2 with ct = 1 for every t) gives for every x > 0, Q(ST ≤ −x) ≤ exp − Taking x = sT (D) > 0 proves the result.
2 P2x
T 12 t=1
2
= exp − 2x T
.
Remark C.9 (Lower envelopes as non-i.i.d. drift conditions). The lower envelope DT in Definition C.7 is a non-i.i.d. analogue of a positive drift condition. To make the comparison precise, recall the simple i.i.d. likelihood-ratio setting: let X1 , X2 , . . . be i.i.d. under an alternative law Q1 , and let P0 be a simple null. If ℓ(X) := log(dQ1 /dP0 )(X), then under Q1 , EQ1 [ℓ(X1 )] = KL(Q1 ∥P0 ), where KL denotes the Kullback–Leibler information (Kullback and Leibler, 1951). Hence EQ1 [
27
t=1 ℓ(Xt )] =
PT
T KL(Q1 ∥P0 ). Thus, in the i.i.d. likelihood-ratio problem, cumulative evidence has a linear expected drift with per-observation rate KL(Q1 ∥P0 ). This is the classical signal rate underlying sequential likelihood-ratio testing (Wald, 1945) and the i.i.d. Chernoff–Stein Lemma (Cover and Thomas, 2006, Theorem 11.8.3). In the present problem, there need not be a repeated observation law, a stationary ϕ 2 distribution, or independent increments. The one-step predictable edge bQ t (θ) = ⟨θ, ϕt ⟩mt − ⟨θ, ϕt ⟩ ϕ is Ht−1 -measurable and may vary with time, covariates, and the realized past. Consequently, there
is generally no single population drift parameter analogous to KL(Q1 ∥P0 ). The deterministic lower envelope BTQ (θ⋆ ) :=
Q ⋆ t=1 bt (θ ) ≥ DT
PT
Q-a.s. for every T ≥ 1 records, instead, the cumulative
amount of predictable feature-aligned miscalibration available to a fixed comparator θ⋆ along the possibly non-i.i.d. stream. Remark C.10 (Connection to linear-full feature alternative in Theorem 3.2). The class Qϕ,lin κ,t0 (Φ) that we studied in Definition 3.1 Theorem 3.2, is a simple uniform-edge subclass of the more general cumulative-edge alternative studied here. From the proof of Theorem 3.2, we can observe that ϕ κ,t0 ; Φ), with D κ,t0 := κT − 2(t − 1). In this case, the envelope D κ,t0 is linear, Qϕ,lin 0 κ,t0 (Φ) ⊆ Qcum (D T T
but the data stream may still be non-i.i.d. It asserts only linear growth of cumulative predictable edge. Such growth may arise from i.i.d. repetitions, but it may also arise from nonstationary or history-dependent streams, including changepoints and time-varying feature distributions.
C.4
Missing proofs
Proof of Proposition 2.2. Fix P ∈ P0 (G). Then, for every t, EP [Zt | Gt−1 ] = 0 P -a.s. Since Ht−1 ⊆ Gt−1 , the tower property yields EP [Zt | Ht−1 ] = EP [EP [Zt | Gt−1 ] | Ht−1 ] = 0 P -a.s. Hence P ∈ P0 (H). Applying this implication along the filtration chain Hmarg ⊆ Hw ⊆ Hϕ gives the displayed hierarchy. Proof of Theorem 2.5. Fix P ∈ P0 (H). Nonnegativity follows from λt ∈ Λα and Zt ∈ {−α, 1 − α}. The process is adapted to the monitored information because Mt−1 and λt are Ht−1 -measurable and Zt is Ht -measurable. For each finite t, Mt is integrable by induction, since each multiplier is bounded. Using H-predictability and the definition of P0 (H), EP [Mt | Ht−1 ] = EP [Mt−1 (1 + λt Zt ) | Ht−1 ] = Mt−1 {1 + λt EP [Zt | Ht−1 ]} = Mt−1 .
28
Here, the second equality is because Mt−1 is Ht−1 -measurable. The final equality due to the bounded martingale difference sequence (MDS) property of (Zt )t≥1 with respect to H, (Proposition C.3). Thus (Mt ) is a nonnegative P -martingale with initial value M0 = 1. Since P was arbitrary in P0 (H), the process is a test martingale for the composite null. Proof of Corollary 2.6. For every P ∈ P0 (H), Theorem 2.5 shows that (Mt ) is a nonnegative P martingale with M0 = 1, and hence a nonnegative supermartingale. Ville’s inequality for nonnegative supermartingales gives
!
P sup Mt ≥ c
≤
t≥0
EP [M0 ] 1 = . c c
Taking c = 1/γ and using {τγ < ∞} = {supt≥1 Mt ≥ 1/γ} ⊆ {supt≥0 Mt ≥ 1/γ} proves the claim. Proof of Corollary 2.7. Fix any P ∈ P0 (G). Since Ht−1 ⊆ Gt−1 , any Ht−1 -measurable bet λt is also Gt−1 -measurable. In other words, any H-predictable betting process is automatically G-predictable. Moreover, by definition of P0 (G), EP [Zt | Gt−1 ] = 0
P -a.s.
The no-bankruptcy constraint gives 1 + λt Zt ≥ 0 because Zt ∈ {−α, 1 − α} and λt ∈ Λα . Hence Mt ≥ 0. Because Mt−1 and λt are both Gt−1 -measurable, we may pull them out when conditioning on Gt−1 . Therefore, EP [Mt | Gt−1 ] = EP [Mt−1 (1 + λt Zt ) | Gt−1 ] = Mt−1 (1 + λt EP [Zt | Gt−1 ]) = Mt−1 . Thus (Mt )t≥0 is a nonnegative P -martingale with respect to G for every P ∈ P0 (G). Consequently marg it is a test martingale, and hence an e-process, for the stronger null P0 (G). Finally, since Ht−1 ⊆ w ⊆ Hϕ , the preceding argument applies first with (H, G) = (Hmarg , Hw ), then with (Hmarg , Hϕ ), Ht−1 t−1
and finally with (Hw , Hϕ ). This proves the stated transfer of validity. Proof of Proposition 2.8. Since Q ∈ / P0 (G), the stronger G-conditional calibration null fails. Since Q ∈ P0 (H), Theorem 2.5 implies that every H-predictable no-bankruptcy wealth process is a nonnegative Q-martingale with respect to H. Ville’s inequality then gives Q(τγ < ∞) ≤ γ. It
29
remains to prove the expected log-wealth bound. For each t, we have EQ [1 + λt Zt | Ht−1 ] = 1 + λt EQ [Zt | Ht−1 ] = 1. Since log is concave, Jensen’s inequality gives EQ [log(1 + λt Zt ) | Ht−1 ] ≤ log EQ [1 + λt Zt | Ht−1 ] = 0. Summing from t = 1 to T gives the following EQ [log MT ] =
T X
EQ [log(1 + λt Zt )] ≤ 0.
t=1
Thus, a coarser skeptic has no positive systematic log-growth under Q, even though Q violates calibration at the richer information level. Q ⋆ ⋆ Proof of Theorem 3.2. Since Q ∈ Qϕ,lin κ,t0 (Φ), there exists θ ∈ K such that bt (θ ) ≥ κ Q-a.s. for
every t ≥ t0 . For t < t0 , write λ⋆t = ⟨θ⋆ , ϕt ⟩. Since |λ⋆t | ≤ 1/2 and mϕt ∈ [−α, 1 − α] ⊆ [−1, 1], 1 1 3 ⋆ ⋆ ϕ ⋆ 2 ⋆ ⋆ 2 bQ = − ≥ −1. t (θ ) = λt mt − (λt ) ≥ −|λt | − (λt ) ≥ − − 2 4 4 Therefore, for every T ≥ t0 ,
Q ⋆ t=1 bt (θ ) ≥ −(t0 − 1) + κ(T − t0 + 1) ≥ κT − 2(t0 − 1) =: DT , where
PT
the last inequality uses κ ≤ 1. If T ≥ nctx γ (r), then by definition log(1/γ) + 2(t0 − 1) + rT ≤ κT /2, and hence sT (D) = DT − rT − log(1/γ) ≥ κT /2. Theorem C.8 gives 2(κT /2)2 Q(τγctx > T ) ≤ exp − T
!
κ2 T = exp − 2
!
.
Since the events {τγctx > T } decrease to {τγctx = ∞}, the exponential bound implies Q(τγctx < ∞) = 1. ctx Finally, let n = nctx γ (r) and write τ = τγ . By the tail-sum formula for nonnegative integer-valued
random variables, EQ [τ ] =
P∞
T =0 Q(τ > T ). For T < n, we use the trivial bound Q(τ > T ) ≤ 1,
2
while for T ≥ n the preceding exponential tail bound gives Q(τ > T ) ≤ exp − κ 2T . Hence ∞ X
κ2 T EQ [τ ] ≤ n + exp − 2 T =n
!
=n+
∞ X
−κ2 /2
e
T =n
T
2
e−κ n/2 1 =n+ . 2 /2 ≤ n + −κ 1−e 1 − e−κ2 /2 2
Since κ ∈ (0, 1], x = κ2 /2 ∈ (0, 1/2], and the elementary inequality 1 − e−x ≥ x/2 gives 1 − e−κ /2 ≥ 30
4 κ2 ctx ctx 4 . Therefore, we obtain the desired result EQ [τγ ] ≤ nγ (r) + κ2 .
D
Experimental details
D.1
Simulation protocol
Synthetic Negative-Binomial panel. The validity experiment and the synthetic-data half of the real-data experiment (both in Section 4) share a panel of N =1,000 count series over T =1,460 days. For each series and day t, demand follows Dt ∼ NegBin(nt , pt ),
µt = exp x⊤ t β ,
pt =
nt , nt + µ t
with eight-dimensional covariate vector (w)
(w)
(y)
(y)
xt = 1, sint , cost , promot , logpricet , sint , cost , gt .
The components are: (w)
• Weekly seasonality sint (y)
sin(2πt/365), cost
(w)
= sin(2πt/7), cost
(y)
= cos(2πt/7) and yearly seasonality sint
=
= cos(2πt/365).
• A 2-state Markov promotion flag promot ∈ {0, 1} with Pr(promot =1 | promot−1 =0) = 0.06 and Pr(promot =1 | promot−1 =1) = 0.70. • A random-walk log-price logpricet = logpricet−1 + 0.02 ξt plus an iid jitter 0.10 ηt , with ξt , ηt ∼ N (0, 1). • An AR(1) latent factor gt = 0.95 gt−1 + 0.10 ζt , ζt ∼ N (0, 1). The coefficient vector is β = (6, 0.20, −0.08, 0.35, −0.45, 0.08, −0.08, 0.15), and the NB dispersion is nt = 10 on promotion days and nt = 15 otherwise. Per-series random seeds are drawn deterministically from a SHA-256 hash of the tag ("cov", T, series_id), so the panel is fully reproducible. Oracle quantiles. The oracle in the validity experiment of Section 4 issues the analytic Negative−1 Binomial α-quantile q̂t = FNB(n (α) at Q=99 levels α ∈ {0.01, 0.02, . . . , 0.99} per series. Since t ,pt )
31
the conditional law is exactly NegBin(nt , pt ), the oracle prediction matches the true conditional quantile and any rejection by the betting test counts as a Type-I error by design. OCO Skeptic implementation.
The AdaGrad, FTRL (with ℓ1 proximal regularisation), and
SGD updates used by the betting Skeptic are implemented via the corresponding optimizers in the River streaming-machine-learning package (Montiel et al., 2021). The ONS Skeptic is implemented as a custom subclass of River’s base optimizer following the Online Newton Step of Hazan et al. (2007); we maintain the inverse Hessian incrementally via a rank-1 Sherman–Morrison update. River’s own Newton optimizer is not used.
D.2
Additional simulation results
Type-I validity for SGD and FTRL Skeptics.
For completeness we report the two OCO
optimisers omitted from Figure 4. The setup is identical: N =1,000 oracle series of length T =1,460, nominal level αtest =0.05, rejection threshold 1/α=20 from Ville’s inequality. Both optimisers track the 5% Ville floor across the full range of α, matching the AdaGrad and ONS panels in Figure 4.
Figure 6: Type-I rejection rate vs. quantile α for the SGD and FTRL Skeptics (N =1,000 oracle series, T =1,460, αtest =0.05). Grey band: one-sided 99% MC tolerance. Companion to Figure 4.
32
Per-feature rejection across context windows (Chronos-2). We complement Figure 5 (Chronos-2 with a 365-day sliding window in the main body) with the same combined Rossmann | Synthetic panels for the remaining context windows. The top-1 hyperparameter per pipeline is auto-selected by mean rejection on the full reference skeptic; emphasised single-feature skeptics are the top-2 per panel (computed dynamically per cell), with the remaining single-feature configurations drawn as faded grey landscape. See Figure 5 for the full explanation of the panel layout. Rossmann · chronos2 · sliding_30
Synthetic · chronos2 · sliding_30
1,115 stores · ftrl (lr0.1-l10.1-shrink0.5)
1,000 SKUs · ftrl (lr0.1-l10.1-shrink0.5)
rejection rate
100%
75%
50%
25%
5%
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
α (quantile level) marginal
promo
is_saturday
sin_w
full
Figure 7: Per-feature rejection rate vs. quantile α for Chronos-2 with a 30-day sliding window (FTRL betting on both pipelines). Companion to Figure 5.
Rossmann · chronos2 · sliding_90
Synthetic · chronos2 · sliding_90
1,115 stores · adagrad (lr0.1-shrink0.5)
1,000 SKUs · ftrl (lr0.1-l10.1-shrink0.5)
rejection rate
100%
75%
50%
25%
5%
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
α (quantile level) marginal
promo
is_saturday
log_price
full
Figure 8: Per-feature rejection rate vs. quantile α for Chronos-2 with a 90-day sliding window (Rossmann: AdaGrad; Synthetic: FTRL). Companion to Figure 5.
33
Rossmann · chronos2 · sliding_180
Synthetic · chronos2 · sliding_180
1,115 stores · ftrl (lr0.1-l10.1-shrink0.5)
1,000 SKUs · ftrl (lr0.1-l10.1-shrink1.0)
rejection rate
100%
75%
50%
25%
5%
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
α (quantile level) marginal
promo
is_saturday
sin_y
full
Figure 9: Per-feature rejection rate vs. quantile α for Chronos-2 with a 180-day sliding window (FTRL betting on both pipelines). Companion to Figure 5. Per-feature rejection: Moirai-2 forecaster.
The same panel for the second foundation
forecaster (Moirai-2) at the matching 365-day sliding window. Moirai-2 attains substantially higher rejection rates across all features, indicating heavier conditional miscalibration; nevertheless the per-feature ordering remains informative for diagnosis. Rossmann · moirai2 · sliding_365
Synthetic · moirai2 · sliding_365
1,115 stores · adagrad (lr0.1-shrink0.5)
1,000 SKUs · ftrl (lr0.01-l10.001-shrink1.0)
rejection rate
100%
75%
50%
25%
5%
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
0.01
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
0.99
α (quantile level) marginal
promo
target_lag14
z_prev
sin_w
full
Figure 10: Per-feature rejection rate vs. quantile α for Moirai-2 with a 365-day sliding window (Rossmann: AdaGrad; Synthetic: FTRL). Companion to Figure 5.
34
Hyperparameter sensitivity at the main-paper cell. We map the rejection-rate landscape across the full hyperparameter grid for each OCO optimiser at the main-paper cell (Chronos-2 with a 365-day sliding window) at the median quantile α=0.5, where rejection-rate signal-to-noise is best for ranking hyperparameters. Rows are the (η, λ1 , shrink) HP triples on disk; columns are skeptic configurations (full reference plus each single feature, restricted to the configurations covered consistently across all three optimisers). Rows are sorted by descending row-mean rejection rate, so the strongest HPs sit at the top of every panel.
lr=0.1
l1=—
shrink=0.5
0.44
0.57
0.03
0.23
0.09
0.07
0.09
0.05
0.11
0.09
0.10
0.05
0.08
0.07
0.58
0.09
0.03
1.0
lr=0.1
l1=—
shrink=1.0
0.23
0.57
0.04
0.23
0.09
0.07
0.09
0.05
0.09
0.08
0.09
0.05
0.09
0.06
0.57
0.09
0.03
0.8
lr=1
l1=—
shrink=1.0
0.15
0.28
0.08
0.10
0.05
0.06
0.12
0.11
0.12
0.13
0.13
0.12
0.13
0.11
0.13
0.12
0.14
0.6
lr=1
l1=—
shrink=0.5
0.09
0.34
0.06
0.13
0.06
0.06
0.07
0.07
0.10
0.09
0.08
0.06
0.10
0.09
0.19
0.07
0.07
0.4
lr=0.01
l1=—
shrink=0.5
0.27
0.17
0.00
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.23
0.00
0.00
0.2
lr=0.01
l1=—
shrink=1.0
0.27
0.17
0.00
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.23
0.00
0.00
0.0
l ful
o
om pr
l_h oo
d oli
h sc is_
ay
da ur
y
t sa
is_
sin
w_
do
os
c w_
do
7
lag
rs_
me
sto cu
lag
s_
r me
sto cu
14
lag
s_
r me
sto cu
21
28
lag
rs_
me
sto cu
ag
t_l
e rg
ta
7
ag
t_l
e rg
ta
14
ag
t_l
e rg
ta
21
ag
t_l
e rg
ta
28
r
da
len
ca
ss
nly
_o
ev pr
rejection rate (across SKUs)
Rossmann [model=chronos2, strategy=sliding_365, algo=adagrad] — rejection rate × skeptic (α=0.5; rows sorted by row mean across all skeptics)
le re tu
fea
z_
Figure 11: Rossmann HP sweep: AdaGrad, Chronos-2 with a 365-day sliding window, α=0.5.
lr=0.1
l1=—
shrink=0.5
0.14
0.69
0.15
0.26
0.12
0.06
0.18
0.11
0.17
0.15
0.19
0.13
0.15
0.13
0.20
0.16
0.13
1.0
lr=0.1
l1=—
shrink=1.0
0.07
0.45
0.11
0.23
0.12
0.06
0.13
0.09
0.12
0.10
0.14
0.13
0.13
0.12
0.08
0.16
0.13
0.8
lr=1
l1=—
shrink=0.5
0.13
0.42
0.06
0.19
0.05
0.05
0.10
0.09
0.15
0.12
0.11
0.09
0.14
0.11
0.26
0.06
0.06
0.6
lr=0.01
l1=—
shrink=0.5
0.18
0.68
0.10
0.18
0.05
0.04
0.04
0.02
0.04
0.05
0.05
0.02
0.04
0.04
0.19
0.05
0.02
0.4
lr=1
l1=—
shrink=1.0
0.12
0.15
0.16
0.16
0.01
0.07
0.11
0.08
0.11
0.10
0.13
0.08
0.13
0.06
0.12
0.07
0.06
0.2
lr=0.01
l1=—
shrink=1.0
0.08
0.44
0.06
0.20
0.05
0.04
0.04
0.02
0.04
0.05
0.05
0.02
0.04
0.04
0.08
0.05
0.02
0.0
l ful
o
om pr
h ol_
ho sc is_
d oli
ay
da ur
t sa
is_
y
sin
w_
do
os
c w_
do
7
lag
rs_
me
sto cu
lag
s_
r me
sto cu
14
lag
s_
r me
sto cu
21
28
lag
rs_
me
sto cu
ag
t_l
e rg
ta
7
ag
t_l
e rg
ta
14
ag
t_l
e rg
ta
21
ag
t_l
e rg
ta
28
r
da
len
ca
ss
nly
_o
ev pr
z_
rejection rate (across SKUs)
Rossmann [model=chronos2, strategy=sliding_365, algo=sgd] — rejection rate × skeptic (α=0.5; rows sorted by row mean across all skeptics)
le re tu
fea
Figure 12: Rossmann HP sweep: SGD, Chronos-2 with a 365-day sliding window, α=0.5.
35
Rossmann [model=chronos2, strategy=sliding_365, algo=ftrl] — rejection rate × skeptic (α=0.5; rows sorted by row mean across all skeptics) l1=0.1
shrink=0.5
0.46
0.59
0.03
0.22
0.07
0.06
0.07
0.04
0.08
0.08
0.08
0.03
0.07
0.06
0.58
0.08
0.03
lr=0.1
l1=0.01
shrink=0.5
0.44
0.59
0.03
0.22
0.07
0.06
0.07
0.04
0.08
0.08
0.08
0.03
0.07
0.06
0.58
0.08
0.03
lr=0.1
l1=0.001
shrink=0.5
0.44
0.59
0.03
0.22
0.07
0.06
0.07
0.04
0.08
0.08
0.08
0.03
0.07
0.06
0.58
0.08
0.03
0.60
0.03
0.21
0.07
0.06
0.06
0.04
0.07
0.07
0.06
0.03
0.06
0.05
0.58
0.07
0.03
lr=0.1
l1=1
shrink=0.5
0.50
lr=0.1
l1=0.1
shrink=1.0
0.40
0.59
0.03
0.22
0.07
0.06
0.07
0.04
0.07
0.08
0.08
0.03
0.06
0.06
0.58
0.08
0.03
lr=0.1
l1=1
shrink=1.0
0.47
0.60
0.03
0.21
0.07
0.06
0.06
0.04
0.06
0.07
0.06
0.03
0.06
0.05
0.58
0.07
0.03
lr=0.1
l1=0.01
shrink=1.0
0.39
0.59
0.03
0.22
0.07
0.06
0.07
0.04
0.07
0.08
0.08
0.03
0.07
0.06
0.58
0.08
0.03
lr=0.1
l1=0.001
shrink=1.0
0.39
0.59
0.03
0.22
0.07
0.06
0.07
0.04
0.07
0.08
0.08
0.03
0.07
0.06
0.58
0.08
0.03
lr=1
l1=1
shrink=0.5
0.12
0.63
0.07
0.20
0.07
0.05
0.10
0.06
0.12
0.08
0.11
0.05
0.10
0.06
0.33
0.10
0.09
lr=1
l1=0.01
shrink=0.5
0.11
0.64
0.07
0.18
0.04
0.05
0.10
0.05
0.13
0.09
0.10
0.04
0.10
0.07
0.31
0.08
0.08
lr=1
l1=0.001
shrink=0.5
0.11
0.64
0.07
0.18
0.04
0.05
0.10
0.05
0.13
0.09
0.10
0.04
0.10
0.07
0.30
0.09
0.08
lr=1
l1=0.1
shrink=0.5
0.11
0.64
0.07
0.18
0.04
0.05
0.09
0.05
0.12
0.09
0.11
0.04
0.10
0.07
0.32
0.08
0.08
lr=1
l1=1
shrink=1.0
0.06
0.58
0.07
0.14
0.06
0.05
0.08
0.04
0.10
0.06
0.08
0.05
0.11
0.04
0.18
0.10
0.09
lr=1
l1=0.01
shrink=1.0
0.07
0.53
0.05
0.14
0.03
0.03
0.08
0.03
0.12
0.05
0.08
0.04
0.13
0.04
0.10
0.09
0.09
lr=1
l1=0.001
shrink=1.0
0.07
0.53
0.04
0.15
0.03
0.03
0.08
0.03
0.12
0.05
0.08
0.04
0.13
0.04
0.10
0.09
0.09
lr=1
l1=0.1
shrink=1.0
0.07
0.54
0.04
0.14
0.03
0.03
0.08
0.03
0.12
0.05
0.08
0.04
0.13
0.04
0.11
0.08
0.09
lr=0.01
l1=0.001
shrink=0.5
0.17
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.13
0.00
0.00
lr=0.01
l1=0.001
shrink=1.0
0.17
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.13
0.00
0.00
lr=0.01
l1=0.01
shrink=0.5
0.17
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.13
0.00
0.00
lr=0.01
l1=0.01
shrink=1.0
0.17
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.13
0.00
0.00
lr=0.01
l1=0.1
shrink=0.5
0.17
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.13
0.00
0.00
lr=0.01
l1=0.1
shrink=1.0
0.17
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.13
0.00
0.00
lr=0.01
l1=1
shrink=0.5
0.15
0.02
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.11
0.00
0.00
lr=0.01
l1=1
shrink=1.0
0.15
0.02
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.11
0.00
0.00
lr=0.001
l1=0.001
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.001
shrink=1.0
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.01
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.01
shrink=1.0
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.1
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.1
shrink=1.0
0.00
l
ful
0.00
0.00
da oli
h
ol_
ho
sc
is_
0.00
y
o
om pr
0.00
y
is_
s
da ur at
0.00
os
in
do
s w_
do
me
sto
cu
0.00
0.00
lag rs_
me
sto
cu
0.00
4
7
c w_
1 lag
rs_
me
sto
cu
0.00
1
2 lag
rs_
me
sto
cu
0.00
7 ag t_l
8
rs_
2 lag
e rg
ta
0.00
t_
e rg
ta
14 lag
0.00
0.00
1
t
e rg
ta
0.00
ar
8
g2 _la
t
e rg
ta
g2 _la
ca
d len
0.00
z_
ev pr
nly _o
1.0
0.8
0.6
0.4
rejection rate (across SKUs)
lr=0.1
0.2
0.0
0.00
s
t fea
s ele ur
Figure 13: Rossmann HP sweep: FTRL, Chronos-2 with a 365-day sliding window, α=0.5. The hyperparameter highlighted in the main paper sits in the upper rows of this panel.
36
Synthetic [model=chronos2, strategy=sliding_365, algo=adagrad] — rejection rate × skeptic (α=0.5; rows sorted by row mean across all skeptics) l1=—
shrink=1.0
0.30
0.27
0.05
0.08
0.05
0.06
0.36
0.41
0.11
0.06
0.09
lr=0.1
l1=—
shrink=0.5
0.27
0.27
0.05
0.08
0.05
0.06
0.35
0.41
0.11
0.06
0.09
lr=0.01
l1=—
shrink=0.5
0.28
0.01
0.00
0.00
0.00
0.00
0.01
0.12
0.01
0.00
0.00
lr=0.01
l1=—
shrink=1.0
0.28
0.01
0.00
0.00
0.00
0.00
0.01
0.12
0.01
0.00
0.00
lr=1
l1=—
shrink=0.5
0.03
0.04
0.03
0.04
0.03
0.04
0.07
0.03
0.03
0.03
0.03
lr=1
l1=—
shrink=1.0
0.03
0.03
0.03
0.03
0.02
0.03
0.03
0.03
0.03
0.03
0.04
lr=0.001
l1=—
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=—
shrink=1.0
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
y
g
1.0
0.8
0.6
0.4
0.2
l ful ) int
(jo
re xtu ) mi -hoc t s po
pt
ce er
int
ev pr
z_
_w
sin
w
s_
co
o
om pr
log
e
ric
_p
_y
sin
s_
co
rejection rate (across SKUs)
lr=0.1
0.0
(
Figure 14: Synthetic HP sweep: AdaGrad, Chronos-2 with a 365-day sliding window, α=0.5.
Synthetic [model=chronos2, strategy=sliding_365, algo=sgd] — rejection rate × skeptic (α=0.5; rows sorted by row mean across all skeptics) l1=—
shrink=0.5
0.12
0.25
0.03
0.04
0.02
0.03
0.12
0.49
0.18
0.06
0.06
lr=0.01
l1=—
shrink=1.0
0.09
0.25
0.03
0.04
0.02
0.03
0.09
0.49
0.18
0.06
0.06
lr=0.1
l1=—
shrink=0.5
0.03
0.11
0.09
0.07
0.05
0.07
0.10
0.16
0.04
0.03
0.06
lr=0.1
l1=—
shrink=1.0
0.03
0.08
0.09
0.07
0.05
0.06
0.05
0.12
0.03
0.02
0.05
lr=1
l1=—
shrink=0.5
0.03
0.03
0.03
0.03
0.04
0.04
0.04
0.05
0.03
0.03
0.03
lr=1
l1=—
shrink=1.0
0.03
0.05
0.03
0.03
0.03
0.04
0.04
0.03
0.03
0.03
0.04
lr=0.001
l1=—
shrink=0.5
0.11
0.01
0.00
0.00
0.00
0.00
0.07
0.06
0.00
0.00
0.00
lr=0.001
l1=—
shrink=1.0
0.07
0.01
0.00
0.00
0.00
0.00
0.03
0.06
0.00
0.00
0.00
y
g
1.0
0.8
0.6
0.4
0.2
l ful ) int
(jo
re xtu ) mi -hoc t s po
pt
ce er
int
ev pr
z_
_w
sin
w
s_
co
o
om pr
log
e
ric
_p
_y
sin
s_
co
rejection rate (across SKUs)
lr=0.01
0.0
(
Figure 15: Synthetic HP sweep: SGD, Chronos-2 with a 365-day sliding window, α=0.5.
37
Synthetic [model=chronos2, strategy=sliding_365, algo=ftrl] — rejection rate × skeptic (α=0.5; rows sorted by row mean across all skeptics) lr=0.1
l1=0.001
shrink=1.0
0.43
0.32
0.05
0.07
0.04
0.05
0.39
0.49
0.16
0.07
0.08
lr=0.1
l1=0.1
shrink=0.5
0.43
0.33
0.05
0.07
0.04
0.05
0.39
0.49
0.16
0.07
0.08
lr=0.1
l1=0.1
shrink=1.0
0.43
0.32
0.05
0.07
0.04
0.05
0.39
0.49
0.16
0.07
0.08
lr=0.1
l1=0.001
shrink=0.5
0.42
0.33
0.05
0.07
0.04
0.05
0.39
0.49
0.16
0.07
0.08
lr=0.1
l1=0.01
shrink=1.0
0.43
0.32
0.05
0.07
0.04
0.05
0.39
0.49
0.16
0.07
0.08
lr=0.1
l1=0.01
shrink=0.5
0.42
0.33
0.05
0.07
0.04
0.05
0.39
0.49
0.16
0.07
0.08
lr=1
l1=0.1
shrink=0.5
0.04
0.17
0.06
0.06
0.05
0.06
0.25
0.12
0.02
0.03
0.06
lr=1
l1=0.01
shrink=0.5
0.04
0.17
0.06
0.06
0.05
0.05
0.25
0.12
0.03
0.02
0.06
lr=1
l1=0.001
shrink=0.5
0.03
0.16
0.06
0.05
0.05
0.05
0.25
0.12
0.03
0.02
0.05
lr=1
l1=0.1
shrink=1.0
0.04
0.08
0.07
0.07
0.04
0.05
0.09
0.10
0.03
0.03
0.04
lr=1
l1=0.01
shrink=1.0
0.04
0.07
0.07
0.07
0.04
0.05
0.09
0.10
0.03
0.02
0.04
lr=1
l1=0.001
shrink=1.0
0.04
0.07
0.07
0.07
0.04
0.05
0.09
0.10
0.03
0.02
0.04
lr=0.01
l1=0.001
shrink=0.5
0.24
0.01
0.00
0.00
0.00
0.00
0.01
0.08
0.00
0.00
0.00
lr=0.01
l1=0.001
shrink=1.0
0.24
0.01
0.00
0.00
0.00
0.00
0.01
0.08
0.00
0.00
0.00
lr=0.01
l1=0.01
shrink=0.5
0.24
0.01
0.00
0.00
0.00
0.00
0.01
0.08
0.00
0.00
0.00
lr=0.01
l1=0.01
shrink=1.0
0.24
0.01
0.00
0.00
0.00
0.00
0.01
0.08
0.00
0.00
0.00
lr=0.01
l1=0.1
shrink=0.5
0.24
0.01
0.00
0.00
0.00
0.00
0.01
0.08
0.00
0.00
0.00
lr=0.01
l1=0.1
shrink=1.0
0.24
0.01
0.00
0.00
0.00
0.00
0.01
0.08
0.00
0.00
0.00
lr=0.001
l1=0.001
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.001
shrink=1.0
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.01
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.01
shrink=1.0
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.1
shrink=0.5
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
lr=0.001
l1=0.1
shrink=1.0
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
y
g
1.0
0.6
0.4
l ful ) int
(jo
re xtu ) mi -hoc st o (p
pt
ce er
int
ev pr
z_
_w
sin
w
s_ co
o om pr
log
e ric
_p
_y
sin
s_ co
rejection rate (across SKUs)
0.8
0.2
0.0
Figure 16: Synthetic HP sweep: FTRL, Chronos-2 with a 365-day sliding window, α=0.5. The hyperparameter highlighted in the main paper sits in the upper rows of this panel.
38
D.3
Computational considerations
Forecaster quantile predictions (Chronos-2, Moirai-2) were generated on a single workstation with one NVIDIA GeForce RTX 4090 GPU; predictions are cached as Parquet files and reused across all betting experiments. Betting tests, validity runs, and Monte Carlo heatmaps are CPU-only and run via a process pool of ∼8 workers on a consumer multi-core CPU. End-to-end runtime per HP bundle is on the order of minutes; the full HP sweep across both pipelines and all (model, strategy, optimiser) cells totals several hundred CPU-hours. Preliminary and exploratory sweeps not reported in the paper consumed a similar order of magnitude of compute.
AI Usage We used an LLM (GPT v5.5) as an assistive tool for: (i) improving exposition via language editing; and (ii) aiding understanding of theoretical concepts. All LLM-generated suggestions were independently verified by the co-author team.
39