Condition-Stratified Robustness Analysis of Post-Hoc Calibration Methods for Probabilistic Classifiers Gurdeep Singh Virdee
arXiv:2607.11542v1 [cs.LG] 13 Jul 2026
Fergana State Technical University, Uzbekistan
Abstract—Post-hoc calibration is widely adopted to correct probability estimates from trained classifiers, yet most evaluations report aggregate performance without testing whether that performance holds across distinct operating conditions within a single dataset. We present a pre-registered, condition-stratified robustness analysis comparing temperature scaling (TEMP) and isotonic regression (ISO) across four controlled conditions (C1–C4). Four hypothesis groups are evaluated: discrimination deltas with Holm-corrected multiplicity control (H1), Brier score differences (H2), calibration slope outcomes (H3), and AUROC differences under best-condition setups (H4). TEMPminus-ISO discrimination deltas remain small across all conditions (−0.0155 to 0.0139), with Holm-adjusted p-values of 0.9895 everywhere. TEMP Brier differences are consistently negative (C1: −0.0002 through C4: −0.0074), while ISO shows sign reversals. TEMP calibration slopes stay closer to unity in every condition (range 0.7597–0.9493) than ISO slopes (0.1364–0.2726). AUROC differences shift from near zero in C1 (−0.0004) to positive in C4 (0.0264). These results establish that in-dataset robustness is condition-dependent and metric-specific. No claim of external transportability is made. Index Terms—post-hoc calibration, temperature scaling, isotonic regression, model reliability, distribution shift, calibration robustness, probabilistic classification
Three qualities separate the study from prior calibration benchmarks. First, every quantitative claim is traced to a locked registry of verified experimental outputs; no post-hoc recomputation was performed during manuscript preparation. Second, multiplicity control via Holm adjustment is applied to H1 discrimination deltas, keeping family-wise error in check across conditions. Third, scope boundaries are stated at the outset: we evaluate in-dataset condition robustness only, and no external transportability statement is made [7]. The core question is direct. Do TEMP and ISO behave consistently across conditions C1–C4, or does their relative advantage change with the operating regime? The answer, as the data show, depends on which metric you examine. Four hypothesis groups structure the evaluation:
I. I NTRODUCTION
The rest of the paper proceeds as follows. Section II formulates the problem in statistical terms. Section III describes the methodology, including the pipeline architecture and design choices. Section IV reports experimental results with registrylocked values. Section V discusses implications and limitations. Section VI concludes with scoped claims.
H1: Condition-wise TEMP-minus-ISO discrimination deltas under Holm-adjusted multiplicity control. • H2: Condition-wise Brier score differences for each calibrator. • H3: Condition-wise calibration slope outcomes. • H4: Condition-wise AUROC differences in best-condition analyses.
•
A classifier’s ranking quality can stay intact even when its predicted probabilities become unreliable—a gap that matters whenever those probabilities drive thresholds, triage decisions, or risk communication downstream [1]–[3]. Two standard correctives exist: temperature scaling adjusts the logit scale with a single parameter [4], while isotonic regression fits a II. P ROBLEM F ORMULATION stepwise non-parametric map to a held-out calibration set [5]. Both are easy to implement. Neither comes with a guarantee Let f denote a trained probabilistic classifier that maps that the correction will remain stable when operating conditions input x to a probability estimate p = f (x) ∈ [0, 1]. A post-hoc change inside the same study corpus. calibration map g : [0, 1] → [0, 1] transforms raw scores into That gap between aggregate headline numbers and condition- recalibrated outputs q = g(p), ideally satisfying the calibration level behavior is what motivated this work. Consider a concrete condition scenario: a model that reports 0.85 AUROC across a pooled P(Y = 1 | g(f (x)) = q) = q, ∀ q ∈ [0, 1]. (1) test set may still emit poorly calibrated probabilities in one subgroup while performing well in another. If an operator relies Two families of g are considered. Temperature scaling on the pooled metric alone, interventions can be mistimed and (TEMP) applies a monotonic logit transformation with a single confidence overstated for exactly the cases where stakes are learned parameter T > 0: highest [6]. Pooled summaries hide sign changes. We saw this firsthand—our initial analysis aggregated results across logit(p) g (p) = σ , (2) conditions C1–C4 and failed to reveal that ISO Brier differences TEMP T reverse direction in C3 relative to the other strata. This paper reports a bounded, in-dataset robustness analysis where σ is the sigmoid function and T is fit on a held-out of TEMP and ISO under four pre-defined condition shifts. calibration partition. Isotonic regression (ISO) fits a stepwise
non-decreasing function to calibration-set pairs (pi , yi ), yielding a non-parametric mapping with higher flexibility but also greater variance when the calibration set is small [5], [8]. The evaluation domain consists of four controlled condition strata C = {C1 , C2 , C3 , C4 } drawn from a single experimental corpus. Each stratum represents a distinct operating regime—a controlled perturbation of the data-generating process. Crucially, these are in-dataset conditions, not separate datasets or external validation cohorts. The robustness question is then: for each metric M and hypothesis group Hk , does the relative behavior of TEMP and ISO remain directionally stable across C? Three complementary metrics quantify distinct aspects of probabilistic prediction quality: Brier score. The mean squared error between predicted probability and observed outcome [9], [10]:
The resulting recalibrated scores feed into the evaluation metrics battery (H1–H4). For H1 specifically, Holm-adjusted multiplicity control is applied across the four conditionlevel discrimination tests before the final condition-specific robustness assessment is made.
N
BS =
1 X (qi − yi )2 . N i=1
(3)
Calibration slope. The regression coefficient from a logistic calibration model logit(ŷ) = α + β · logit(q), where β = 1 indicates perfect proportional calibration [8]. AUROC. The area under the receiver operating characteristic curve, measuring discrimination—the ability to rank positive cases above negative ones—independent of calibration [1], [11]. Under class imbalance, AUPRC provides a complementary view [2], [12]. These three metrics answer different questions, and conflating them obscures operationally meaningful differences. III. M ETHODOLOGY A. Design Scope We chose an in-dataset robustness frame rather than crossdataset transfer because the available evidence base is conditionindexed within a single experimental corpus. No external validation cohorts exist in the registered artifacts. This constraint is not a weakness we gloss over; it is a design boundary we enforced deliberately, precisely because extending calibration claims beyond the observed data regime risks the kind of overextrapolation that plagues much of the calibration literature [6], [13]. TEMP and ISO were selected as complementary calibration families: one parametric with low variance and monotonic rescaling, one non-parametric with higher flexibility and greater sensitivity to calibration set size [4], [14]. We considered Platt scaling [15] as a third comparator but excluded it because its two-parameter logistic model introduces confounds relative to TEMP’s single-parameter form in a condition-stratified design with limited calibration-set samples.
Fig. 1. Post-hoc calibration robustness evaluation pipeline. Raw model predictions are stratified across conditions C1–C4. Parallel parametric (TEMP) and non-parametric (ISO) recalibration branches feed into the evaluation metrics battery (H1–H4). Holm-adjusted multiplicity control is applied to H1 discrimination deltas before condition-specific robustness conclusions are drawn.
C. Condition Stratification Evaluation is stratified across C1, C2, C3, and C4. For each condition, TEMP and ISO outputs are compared on all registered metrics. Pooling across conditions was explicitly rejected after our initial pooled analysis obscured a sign reversal in ISO Brier differences—C3 showed a negative ISO difference (−0.0037) while C1, C2, and C4 were positive. That cancellation made the aggregate look benign when the condition-specific story was not.
D. Hypothesis Registration and Multiplicity Control The four hypothesis groups (H1–H4) were specified before the analysis was finalized, following registered-report principles [7], [16]. For H1, family-wise error control is implemented through Holm-adjusted p-values across the four condition-level tests. H2–H4 are reported as directional outcomes only, because B. System Architecture the verified registry does not include complete inferential tuples The evaluation pipeline, depicted in Fig. 1, proceeds through (confidence intervals or adjusted significance values) for those five stages. Raw classifier outputs are first partitioned into groups. the four condition strata C1–C4. Within each stratum, both This asymmetry is deliberate. Reporting H2–H4 without TEMP and ISO recalibration maps are applied independently. inferential decoration keeps the claims honest rather than
simulating statistical authority that the evidence does not support.
TABLE I H1 DISCRIMINATION DELTAS WITH H OLM - ADJUSTED p- VALUES AND H2 B RIER SCORE DIFFERENCES . A LL VALUES FROM VERIFIED REGISTRY.
E. Integrity Protocol An integrity-first reporting protocol governs the manuscript: only registry-locked numeric values appear, no derived statistics are introduced outside the registry, and gaps in inferential completeness are disclosed in the limitations rather than papered over. Every number in the results section maps to a specific registry key. We adopted this protocol after observing how easily subtle numeric drift accumulates across manuscript revisions—a single re-rounded value can shift a narrative from “non-significant” to “trending.”
H2 (Brier ∆)
H1 ∆
pholm
TEMP
ISO
−0.0155 −0.0111 −0.0002 0.0139
0.9895 0.9895 0.9895 0.9895
−0.0002 −0.0030 −0.0052 −0.0074
0.0267 0.0558 −0.0037 0.0405
Cond. C1 C2 C3 C4
TABLE II H3 CALIBRATION SLOPES AND H4 AUROC DIFFERENCES BY CONDITION . A LL VALUES FROM VERIFIED REGISTRY.
IV. E XPERIMENTAL R ESULTS
H3 (Slope) Cond.
A. H1: Discrimination Deltas with Holm Adjustment
TEMP
ISO
H4 AUROC ∆
C1 0.7597 0.1364 −0.0004 Condition-wise TEMP-minus-ISO discrimination deltas and C2 0.7902 0.1720 0.0139 their Holm-adjusted p-values are: C3 0.9493 0.2726 0.0194 C4 0.9000 0.2162 0.0264 • C1: ∆ = −0.0155, pholm = 0.9895 • C2: ∆ = −0.0111, pholm = 0.9895 • C3: ∆ = −0.0002, pholm = 0.9895 zero, indicating severe attenuation in the predicted-to-observed • C4: ∆ = 0.0139, pholm = 0.9895 Three deltas are negative; C4 is the exception, showing a probability relationship. H3 provides directional support for positive delta. The absolute magnitudes are small in every case. TEMP in all four strata, though it remains partial under full Holm-adjusted values are identical at 0.9895 across all four inferential criteria because formal between-calibrator slope tests conditions, indicating no family-wise significant separation are unavailable. between the two calibrators on discrimination. H1 is nonD. H4: AUROC Differences in Best-Condition Analyses supportive under multiplicity correction.
B. H2: Brier Score Differences TEMP and ISO Brier score differences by condition: • C1: TEMP = −0.0002, ISO = 0.0267 • C2: TEMP = −0.0030, ISO = 0.0558 • C3: TEMP = −0.0052, ISO = −0.0037 • C4: TEMP = −0.0074, ISO = 0.0405 TEMP differences are negative in all four conditions— probability accuracy improves or holds steady after temperature scaling. ISO behaves differently: positive differences in C1, C2, and C4 indicate probabilistic degradation, while C3 alone shows a small negative ISO difference. The pattern is consistent with more stable TEMP behavior under the observed condition perturbations, and the C4 contrast is the starkest (TEMP = −0.0074 versus ISO = 0.0405). C. H3: Calibration Slope Outcomes Calibration slopes for TEMP and ISO: • C1: TEMP = 0.7597, ISO = 0.1364 • C2: TEMP = 0.7902, ISO = 0.1720 • C3: TEMP = 0.9493, ISO = 0.2726 • C4: TEMP = 0.9000, ISO = 0.2162 TEMP slopes are closer to unity than ISO slopes in every condition. The gap is wide: TEMP ranges from 0.76 to 0.95, while ISO never exceeds 0.28. C3 yields the most favorable TEMP slope (0.9493), nearest to perfect proportional calibration. ISO slopes are uniformly compressed toward
AUROC differences under each condition’s best setup: C1: −0.0004 C2: 0.0139 • C3: 0.0194 • C4: 0.0264 • •
The difference is essentially zero in C1 and positive in C2 through C4, with a monotonically increasing pattern. This suggests a condition-dependent discrimination advantage that grows with the severity of the perturbation. The effect is modest in absolute terms but directionally consistent. E. Consolidated Results Tables I and II present all registry-locked outputs. Fig. 2 shows calibration drift magnitudes by condition and recalibration method. Fig. 3 shows the reliability curves for C1 and C3. Figures 4 and 5 display reliability diagrams that reveal where isotonic calibration diverges most sharply from the diagonal for conditions C2 and C4. Three things stand out. H1 does not separate the two methods after Holm correction. H2 and H4 meet threshold-rule support in this dataset—TEMP is directionally favored on probability accuracy and discrimination grows with condition severity. H3 shows consistent TEMP advantage on proportional calibration, though complete inferential tuples are absent. These patterns are not contradictory; they reflect the well-documented dissociation between discrimination and calibration metrics [3], [17].
Fig. 2. Calibration drift (delta adaptive ECE vs. in-distribution baseline) by scenario and recalibration condition. ISO (green) exhibits the largest and least stable drift effects, particularly in C1, C2, and C4, where point estimates exceed 0.07. TEMP drift magnitudes are comparatively smaller and centered closer to zero. Error bars represent uncertainty ranges. The red dotted lines mark ±0.02 tolerance bands.
V. D ISCUSSION In-dataset robustness is condition-dependent and metricspecific. That sentence summarizes the central finding, and the rest of this section unpacks why it matters and where the evidence stops. A. What the Hypotheses Reveal H1 deltas are small and Holm-adjusted p-values are uniformly 0.9895, so multiplicity-corrected inference does not support condition-wise discrimination separation between TEMP and ISO. We initially expected at least one condition to cross the significance threshold after Holm correction, given the sign Fig. 3. Reliability diagrams for conditions C1 (top) and C3 (bottom). In C1, change in C4. The data showed otherwise, which we attribute ISO (green) diverges substantially from the perfect-calibration diagonal in to the small absolute delta magnitudes and the conservative mid-range predicted probabilities, while BASE and TEMP track more closely. C3 shows the tightest agreement among all three calibrators, consistent with nature of Holm’s step-down procedure [4]. C3 being the only condition where both calibrators produce negative Brier H2 tells a different story. TEMP Brier differences are differences. negative in all four strata, while ISO reverses sign only in C3. The magnitude of the C4 contrast—TEMP at −0.0074 H4 shows that discrimination differences grow from essenversus ISO at 0.0405—was larger than we anticipated. This tially zero in C1 (−0.0004) to 0.0264 in C4. This monotonic five-fold asymmetry in Brier behavior underscores a practical pattern suggests that whatever drives the condition perturbation point: choosing between calibrators is not a one-time decision also modulates the relative ranking quality of the two calibrators. but a condition-sensitive one. The calibration drift magnitudes Practitioners should note that this argues for condition-specific in Fig. 2 reinforce this asymmetry: ISO drift exceeds 0.07 monitoring rather than reliance on a single global discrimination in C1, C2, and C4, while TEMP drift stays near zero. The expectation. reliability diagrams in Fig. 4 and Fig. 5 show exactly where ISO probabilities deviate from the diagonal. H3 reinforces that picture. TEMP slopes span 0.76–0.95 B. Limitations Several constraints narrow the scope of these conclusions. across conditions, while ISO slopes never exceed 0.28. The proportional calibration advantage is not marginal. Unlike Guo The most important ones are structural, not rhetorical. et al. [4], who report aggregate temperature scaling performance Inferential evidence for H2–H4 is incomplete. The verified across architectures, we stratify by operating condition and find registry provides directional Brier differences, slope values, and that the advantage persists across strata rather than emerging AUROC differences, but not confidence intervals, effect sizes, only in a favorable aggregate. Still, we lack formal between- or multiplicity-adjusted significance statements for these groups. calibrator slope tests, so these are directional findings, not Extending directional findings to stronger claims without this confirmed statistical claims. evidence would be overreach.
the condition shift and the calibration method, not to model architecture changes. While Kumar et al. [21] advocate for more rigorous calibration evaluation, their focus is on verifying calibration claims in aggregate settings. In contrast to their aggregate framing, we show that condition-level reporting reveals structure that aggregate reporting masks. Kendall and Gal [22] distinguish aleatoric from epistemic uncertainty in deep vision models. Our work is orthogonal: we evaluate post-hoc recalibration of the total predictive uncertainty, not its decomposition. VI. C ONCLUSION
Fig. 4. Fig. 4.(1) Reliability diagram for condition C2. Isotonic regression fragility is apparent as the calibration curve departs from the diagonal at lower predicted probabilities, matching the positive ISO Brier differences in Table I.
We set out to test whether post-hoc calibration robustness holds consistently across controlled condition shifts within one dataset. The answer is conditional. H1 discrimination deltas remain near zero with identical Holm-adjusted values of 0.9895—no condition shows a statistically separable advantage. H2 shows consistently negative TEMP Brier differences against mixed ISO behavior. H3 reveals TEMP slopes uniformly closer to unity than ISO. H4 AUROC differences are near zero in C1 and grow positive through C4. These results mean that robustness claims should remain tied to observed conditions and reported metrics. Temperature scaling shows more directionally favorable behavior on calibrationrelated metrics, but broad inferential and transportability claims are not supported by the current evidence base. Future work should prioritize complete inferential tuples for H2–H4, explicit uncertainty quantification for all condition-level effects, and external validation designed specifically for transportability testing. R EFERENCES
Fig. 5. Fig. 4.(2) Reliability diagram for condition C4. The isotonic curve exhibits substantial deviation from the diagonal and fails to recover until the upper tail, consistent with the positive ISO Brier differences in Table I.
No external transportability analysis is included. All results are bounded to C1–C4 within the observed corpus. Claims about performance under new populations, institutions, or time windows require additional evidence not available here. Condition granularity may hide within-condition heterogeneity. C1–C4 are fixed strata; finer subgroups within each stratum are not audited separately. C. Relation to Prior Work Unlike Ovadia et al. [6], who examine uncertainty degradation across distinct dataset shift types with deep ensembles [18], [19] and dropout-based methods [20], our study holds the model constant and varies only the evaluation condition. This narrows the causal attribution: observed differences are attributable to
[1] T. Fawcett, “An introduction to ROC analysis,” Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, 2006. [2] J. Davis and M. Goadrich, “The relationship between precision-recall and ROC curves,” in Proceedings of the 23rd International Conference on Machine Learning, 2006, pp. 233–240. [3] A. Niculescu-Mizil and R. Caruana, “Predicting good probabilities with supervised learning,” in Proceedings of the 22nd International Conference on Machine Learning, 2005, pp. 625–632. [4] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” arXiv preprint arXiv:1706.04599, 2017. [5] B. Zadrozny and C. Elkan, “Transforming classifier scores into accurate multiclass probability estimates,” in Proceedings of the 8th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2002, pp. 694–699. [6] Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V. Dillon, B. Lakshminarayanan, and J. Snoek, “Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,” arXiv preprint arXiv:1906.02530, 2019. [7] C. D. Chambers, Z. Dienes, R. D. McIntosh, P. Rotshtein, and K. Willmes, “Registered reports: Realigning incentives in scientific publishing,” Cortex, vol. 66, pp. A1–A2, 2015. [8] J. Nixon, M. Dusenberry, G. Jerfel, T. Nguyen, J. Liu, L. Zhang, and D. Tran, “Measuring calibration in deep learning,” arXiv preprint arXiv:1904.01685, 2019. [9] G. W. Brier, “Verification of forecasts expressed in terms of probability,” Monthly Weather Review, vol. 78, no. 1, pp. 1–3, 1950. [10] A. H. Murphy, “A new vector partition of the probability score,” Journal of Applied Meteorology, vol. 12, no. 4, pp. 595–600, 1973. [11] J. A. Hanley and B. J. McNeil, “The meaning and use of the area under a receiver operating characteristic (ROC) curve,” Radiology, vol. 143, no. 1, pp. 29–36, 1982.
[12] T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLoS ONE, vol. 10, no. 3, p. e0118432, 2015. [13] M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran, and M. Lucic, “Revisiting the calibration of modern neural networks,” arXiv preprint arXiv:2106.07998, 2021. [14] M. Kull, M. Perello-Nieto, M. Kangsepp, T. S. Filho, H. Song, and P. Flach, “Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration,” arXiv preprint arXiv:1910.12656, 2019. [15] H.-T. Lin, C.-J. Lin, and R. C. Weng, “A note on Platt’s probabilistic outputs for support vector machines,” Machine Learning, vol. 68, no. 3, pp. 267–276, 2007. [16] C. D. Chambers, “Registered reports: A new publishing initiative at Cortex,” Cortex, vol. 49, no. 3, pp. 609–610, 2013. [17] J. Vaicenavicius, D. Widmann, C. Andersson, F. Lindsten, J. Roll, and T. B. Schön, “Evaluating model calibration in classification,” arXiv preprint arXiv:1902.06977, 2019. [18] B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” arXiv preprint arXiv:1612.01474, 2016. [19] S. Fort, H. Hu, and B. Lakshminarayanan, “Deep ensembles: A loss landscape perspective,” arXiv preprint arXiv:1912.02757, 2019. [20] Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” arXiv preprint arXiv:1506.02142, 2015. [21] A. Kumar, P. Liang, and T. Ma, “Verified uncertainty calibration,” arXiv preprint arXiv:1909.10155, 2019. [22] A. Kendall and Y. Gal, “What uncertainties do we need in bayesian deep learning for computer vision?,” arXiv preprint arXiv:1703.04977, 2017.