Graphical Abstract Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
arXiv:2607.27143v1 [cs.LG] 29 Jul 2026
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
Highlights Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal • Benchmark of conformal prediction on 15 imbalanced high-stakes datasets. • Marginal conformal prediction fails minority coverage (as low as <1%). • Mondrian CP restores minority coverage (+61.7% gain over marginal CP). • Cost-controlled abstention reduces expected cost under human review. • Quantifies dataset-specific cost break-even thresholds for deferral.
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark Manpreet Singha,∗ , Akshatha Srikanthab and Shyamal Lakhanpalc,∗∗ a Boston University, Boston, MA, USA b University of California, Irvine, CA, USA c University of Maryland, College Park, MD, USA
ARTICLE INFO
ABSTRACT
Keywords: Conformal prediction Imbalanced classification Cost-sensitive learning Selective classification Uncertainty quantification Decision support systems
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, empirical evidence reveals that it severely under-covers rare, costly minority classes— dropping to as low as below 1% coverage. To evaluate and resolve this defect, we conduct a comprehensive benchmark comparing marginal CP, class-conditional (Mondrian) CP, and costcontrolled abstention mechanisms across 15 real-world imbalanced tabular datasets, 7 classification models, 3 probability calibration techniques, and 10 random seeds (3,150 total experimental runs). Our empirical findings demonstrate that Mondrian CP systematically restores valid coverage for the minority class, achieving an average minority-coverage improvement of 61.7 percentage points over marginal CP (𝑝 < 10−80 ). Furthermore, coupling Mondrian CP with costcontrolled abstention significantly reduces overall expected decision costs compared to standard decision boundaries, confidence-based rejectors, and risk-controlled rejectors under realistic human review budgets. We formally quantify the dataset-specific break-even thresholds where deferring ambiguous instances to human experts yields net cost savings. This study provides actionable principles for deploying distribution-free, cost-optimal uncertainty quantification in knowledge-based decision support systems.
1. Introduction Machine learning models operate in knowledge-based decision support systems across high-stakes domains including credit scoring, fraud detection, medical diagnosis, and industrial failure prevention [1]. In these applications, classification errors carry asymmetric real-world consequences. Failing to identify a fraudulent transaction or a severe clinical condition incurs financial or medical costs far exceeding false alarms. Standard classification models output point predictions or raw confidence scores that fail to reflect true prediction uncertainty, rendering automated decisions risky when deployed in knowledge-based operations. Conformal prediction offers a distribution-free framework for uncertainty quantification [2, 3]. Given a target error allowance 𝛼 ∈ (0, 1), conformal prediction converts point predictions into prediction sets containing the true label with a finite-sample guarantee of at least 1 − 𝛼. However, standard conformal prediction guarantees only marginal coverage averaged across the full data distribution [4]. On imbalanced datasets, marginal conformal predictors meet global targets while failing on the minority class. A model set to 90% global coverage can hit its target by covering 96% of majority-class samples while covering under 1% of minority-class samples [5]. In high-stakes knowledge-based systems, this deficit leaves rare cases unprotected. Class-conditional (Mondrian) conformal prediction solves this problem by calculating separate non-conformity quantiles per class label, guaranteeing 1−𝛼 coverage within every class [2, 6]. While Mondrian CP restores the validity ̂ ̂ for rare classes, prediction sets containing multiple labels (|𝐶(𝑋)| > 1) or empty predictions (𝐶(𝑋) = ∅) require an operational action rule. Selective classification and rejection mechanisms allow systems to abstain on ambiguous ∗ Corresponding author
∗∗ Principal corresponding author
[email protected] (M. Singh); [email protected] (A. Srikantha); [email protected] (S. Lakhanpal)
ORCID (s): 0000-0003-2368-2377 (M. Singh); 0009-0005-3753-9848 (A. Srikantha); 0009-0008-3948-511X (S. Lakhanpal)
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 1 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
predictions and defer them to human experts at a review cost 𝐶rev [7, 8]. Prior work on selective classification focuses mostly on confidence thresholds, risk-controlled bounds [9], or decision boundary heuristics. These approaches do not systematically evaluate how class-conditional coverage bounds interact with human review costs on imbalanced tabular data. To resolve this gap, this paper presents an empirical benchmark and theoretical evaluation of cost-sensitive conformal prediction. We evaluate marginal CP, Mondrian CP, and a cost-controlled abstention mechanism across 15 real-world imbalanced tabular datasets from OpenML, using 7 base classification algorithms, 3 probability calibration methods, and 10 random seeds (3,150 total experimental runs). The primary contributions of this paper are: 1. Empirical Quantification of Minority Coverage Deficit: We demonstrate that marginal conformal prediction suffers from severe minority-class undercoverage across diverse imbalanced datasets, dropping below 1% coverage on the most extreme class ratios. 2. Class-Conditional Restoration: We show that Mondrian conformal prediction systematically restores valid minority-class coverage across all tested models and datasets, yielding an average minority coverage gain of 61.7 percentage points (𝑝 < 10−80 ) over marginal conformal prediction. 3. Cost-Controlled Abstention Framework: We integrate Mondrian prediction sets with an asymmetric cost matrix (𝐶FN , 𝐶FP , 𝐶rev ), parameterizing both expert oracle decisions and noisy human review error rates (𝜖hum ) ̂ to show that deferring multi-label ambiguities (|𝐶(𝑋)| ≠ 1) lowers expected decision costs under realistic review overheads. ∗ ) for human review 4. Break-Even Sensitivity Analysis: We derive dataset-specific break-even thresholds (𝐶rev costs, establishing economic boundaries where automated deferral outperforms standard threshold tuning. The rest of this paper is organized as follows. Section 2 reviews related work in conformal prediction and selective classification. Section 3 formalizes the cost-sensitive conformal framework. Section 4 details benchmark datasets, base classifiers, and evaluation metrics. Section 5 presents empirical coverage and cost results. Section 6 discusses deployment guidelines for knowledge-based decision systems, and Section 7 concludes the paper.
2. Related Work This section reviews literature across three areas: distribution-free uncertainty quantification, cost-sensitive machine learning, and selective classification with human deferral.
2.1. Conformal Prediction & Distribution-Free Uncertainty Quantification Conformal prediction (CP), introduced by Vovk et al. [2], is a distribution-free framework for constructing finitesample valid prediction sets [3, 4, 27]. Given an error level 𝛼 ∈ (0, 1), marginal conformal predictors guarantee ̂ that the true label 𝑌 ∈ lies within the prediction set 𝐶(𝑋) ⊆ with probability at least 1 − 𝛼, averaged across data distribution 𝑃 (𝑋, 𝑌 ). In classification, standard non-conformity scores use predicted class probabilities 1 − 𝑃̂ (𝑌 = 𝑦 ∣ 𝑋), while adaptive prediction set (APS) and regularized adaptive prediction set (RAPS) algorithms aggregate sorted cumulative probabilities to trim set sizes [5, 29, 30]. Standard marginal CP has a known weakness under class imbalance: marginal validity holds only on average across the entire sample space [4]. On imbalanced data, marginal CP satisfies global 1 − 𝛼 coverage by over-covering majority-class samples while under-covering rare minority-class samples [5]. Class-conditional (Mondrian) conformal prediction resolves this by calculating non-conformity quantiles independently for each class label 𝑦 ∈ , enforcing ̂ 𝑃 (𝑌 ∈ 𝐶(𝑋) ∣ 𝑌 = 𝑦) ≥ 1 − 𝛼 for every class [2, 6, 28]. Recent studies have refined class-conditional CP. Ding et al. [10] proposed clustered conformal prediction to lower quantile variance when handling sparse classes. Shi et al. [13] introduced rank-calibrated class-conditional CP (RC3P) to restrict prediction set sizes, Stutz et al. [12] proposed differentiable end-to-end conformal training, and Gibbs and Candès [11] introduced adaptive conformal inference for non-stationary environments. These works focus on set size efficiency or adaptive coverage, without analyzing how class-conditional set guarantees translate into monetary or clinical decision costs under human review.
2.2. Cost-Sensitive Machine Learning Cost-sensitive classification handles problems where misclassification errors incur unequal financial, clinical, or operational penalties [14, 16, 1]. In binary decision theory, when false negative costs (𝐶FN ) exceed false positive costs Manpreet Singh et al.: Preprint submitted to Elsevier
Page 2 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
(𝐶FP ), the Bayes-optimal decision rule assigns the positive class whenever posterior probability 𝑃̂ (𝑌 = 1 ∣ 𝑋) exceeds 𝐶FP cost threshold 𝜏 ∗ = 𝐶 +𝐶 [14]. Traditional cost-sensitive techniques reweight loss functions, resample training data, FN FP or recalibrate probabilities via Platt scaling or isotonic regression [15, 32]. These methods depend heavily on accurate point probability estimates. When base classifiers output poorly calibrated or overconfident probabilities, fixed cost thresholds fail to guarantee error rates in finite samples. Moreover, traditional cost-sensitive learning forces a single class assignment for every sample, without an option to abstain when prediction uncertainty makes automated decisions risky.
2.3. Selective Classification & Decision Deferral Selective classification, originating from Chow’s optimal rejection rule [17], enables a classifier to abstain on uncertain predictions and delegate ambiguous cases to human experts [7, 8, 33, 34]. The "Learning to Defer" (L2D) framework expands selective classification by joint-training a predictor and a rejector using surrogate loss functions calibrated to human expert accuracy and review costs [18, 21]. In distribution-free settings, conformal prediction has been adapted for selective deferral. Angelopoulos et al. [9] developed conformal risk control to bound user-specified losses across prediction sets. Straitouri et al. [19, 20] showed that presenting conformal prediction sets to human experts narrows decision choices and improves joint human-AI team performance. A gap remains between these fields. Existing L2D methods assume uncalibrated point outputs or symmetric loss matrices, while conformal risk control focuses on risk bounds rather than explicit financial cost structures containing human review overhead (𝐶rev ) and expert error rates (𝜖hum ). This paper fills that gap by connecting class-conditional (Mondrian) conformal sets with cost minimization and deriving break-even thresholds for human expert deferral.
3. Methodology & Theoretical Framework This section details the theoretical and operational framework of cost-sensitive conformal prediction for imbalanced classification in knowledge-based decision support systems.
3.1. Problem Formulation & Asymmetric Cost Matrix Consider a binary classification task where input features 𝑋 ∈ ⊆ ℝ𝑑 map to a true class label 𝑌 ∈ = {0, 1}. Label 𝑌 = 1 denotes the rare, costly positive class (such as fraudulent transaction, medical disease, or equipment failure), while 𝑌 = 0 denotes the majority negative class. Let 𝜋1 = 𝑃 (𝑌 = 1) ∈ (0, 0.5) represent the prior prevalence of the minority class, indicating severe data imbalance (𝜋1 ≪ 0.5). We define an asymmetric decision cost matrix 𝐂 ∈ ℝ2×2 , where 𝐶𝑖𝑗 represents the cost incurred by predicting ≥0 class 𝑗 ∈ {0, 1} when the true label is 𝑖 ∈ {0, 1}: [ ] [ ] 𝐶 𝐶01 0 𝐶FP = 𝐂 = 00 (1) 𝐶10 𝐶11 𝐶FN 0 Correct decisions carry zero cost (𝐶00 = 𝐶11 = 0). False positive predictions incur penalty 𝐶FP , while false negative predictions incur penalty 𝐶FN . In high-stakes domains, failure to detect a positive case carries far greater risk than a false alarm, establishing 𝐶FN ≫ 𝐶FP . When automated predictions exhibit high ambiguity, the system can abstain from a single hard decision and defer the sample to a human expert. Deferring an instance incurs operational review cost 𝐶rev > 0. If the human expert acts as an ideal oracle (𝜖hum = 0), the deferred instance is correctly resolved at cost 𝐶rev . If the human expert makes decision errors at rate 𝜖hum ∈ (0, 1), the expected cost of deferral expands to 𝐶rev + 𝜖hum (𝐶FN ⋅ 𝕀(𝑌 = 1) + 𝐶FP ⋅ 𝕀(𝑌 = 0)).
3.2. Marginal vs. Class-Conditional (Mondrian) Conformal Predictors 𝑛
cal Let 𝐷cal = {(𝑋𝑖 , 𝑌𝑖 )}𝑖=1 be an exchangeable calibration set independent of base classifier training data. For a candidate sample 𝑋 and label 𝑦 ∈ {0, 1}, a machine learning model outputs estimated posterior class probability 𝑃̂ (𝑌 = 𝑦 ∣ 𝑋). We compute the non-conformity score using the predicted class error:
𝑆(𝑋, 𝑦) = 1 − 𝑃̂ (𝑌 = 𝑦 ∣ 𝑋)
(2)
A higher non-conformity score indicates lower model confidence for label 𝑦. Manpreet Singh et al.: Preprint submitted to Elsevier
Page 3 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
3.2.1. Marginal Conformal Prediction Standard marginal conformal prediction calculates a single non-conformity quantile 𝑞marg across all calibration samples in 𝐷cal : ( ) 𝑛cal ⌈(𝑛cal + 1)(1 − 𝛼)⌉ 𝑞marg = Quantile {𝑆(𝑋𝑖 , 𝑌𝑖 )}𝑖=1 , (3) 𝑛cal For a test sample 𝑋, the marginal prediction set is constructed as: 𝐶̂marg (𝑋) = {𝑦 ∈ {0, 1} ∣ 𝑆(𝑋, 𝑦) ≤ 𝑞marg }
(4)
By distribution-free exchangeability, marginal CP guarantees global coverage: (5)
𝑃 (𝑌 ∈ 𝐶̂marg (𝑋)) ≥ 1 − 𝛼
Under severe class imbalance (𝜋1 ≪ 0.5), the marginal score distribution is dominated by majority-class samples (𝑌 = 0). As a result, 𝑞marg settles near the majority-class quantile, allowing global coverage to hold while minorityclass coverage (𝑌 = 1) drops near zero [6].
3.2.2. Class-Conditional (Mondrian) Conformal Prediction Mondrian conformal prediction restores coverage for rare classes by splitting 𝐷cal into class-specific subsets (0) (1) 𝐷cal = {(𝑋𝑖 , 𝑌𝑖 ) ∈ 𝐷cal ∣ 𝑌𝑖 = 0} of size 𝑛0 and 𝐷cal = {(𝑋𝑖 , 𝑌𝑖 ) ∈ 𝐷cal ∣ 𝑌𝑖 = 1} of size 𝑛1 . Independent quantiles 𝑞0 and 𝑞1 are computed for each class label: ) ( ⌈(𝑛𝑦 + 1)(1 − 𝛼)⌉ 𝑞𝑦 = Quantile {𝑆(𝑋𝑖 , 𝑌𝑖 )}𝑖∈𝐷(𝑦) , , for 𝑦 ∈ {0, 1} (6) 𝑛𝑦 cal The Mondrian prediction set is formed by evaluating each candidate label against its class-specific quantile threshold: 𝐶̂Mondrian (𝑋) = {𝑦 ∈ {0, 1} ∣ 𝑆(𝑋, 𝑦) ≤ 𝑞𝑦 }
(7)
This construction guarantees finite-sample class-conditional validity for both classes independently [28]: 𝑃 (𝑌 ∈ 𝐶̂Mondrian (𝑋) ∣ 𝑌 = 𝑦) ≥ 1 − 𝛼,
∀𝑦 ∈ {0, 1}
(8)
3.3. Cost-Controlled Abstention Mechanism ̂ Prediction set outputs 𝐶(𝑋) ⊆ {0, 1} yield four possible set configurations: singletons ({0} or {1}), multi-label ambiguities ({0, 1}), or empty sets (∅). We define an operational decision rule 𝑌̂action (𝑋) that maps prediction set outputs to automated decisions or human deferral: ⎧0, ̂ if 𝐶(𝑋) = {0} (automated negative) ⎪ ̂ 𝑌̂action (𝑋) = ⎨1, if 𝐶(𝑋) = {1} (automated positive) ⎪DEFER, if |𝐶(𝑋)| ̂ ̂ ̂ ≠ 1 (i.e., 𝐶(𝑋) = {0, 1} or 𝐶(𝑋) = ∅) ⎩
(9)
An empty prediction set occurs when a sample exhibits high non-conformity across all class distributions (𝑆(𝑋, 𝑦) > 𝑞𝑦 for both 𝑦 ∈ {0, 1}), signaling atypical or out-of-distribution feature patterns. Treating both multi-label ̂ ambiguities and empty set outputs as abstention triggers ensures that any non-singleton prediction (|𝐶(𝑋)| ≠ 1) is safely routed to human expert review. ̂ When |𝐶(𝑋)| = 1, the automated decision is executed. If the single label matches the true class 𝑌 , decision cost ̂ is zero. If the automated prediction is incorrect, cost 𝐶FP or 𝐶FN is incurred. When |𝐶(𝑋)| ≠ 1, the system abstains and defers the case to human review. Under expert oracle assumptions (𝜖hum = 0), the per-instance decision loss ̂ 𝐿(𝑌 , 𝐶(𝑋)) is formalized as: ⎧0, ̂ if 𝐶(𝑋) = {𝑌 } ⎪ ̂ ⎪𝐶 , if 𝐶(𝑋) = {1} and 𝑌 = 0 ̂ 𝐿(𝑌 , 𝐶(𝑋)) = ⎨ FP ̂ ⎪𝐶FN , if 𝐶(𝑋) = {0} and 𝑌 = 1 ̂ ⎪𝐶rev , if |𝐶(𝑋)| ≠1 ⎩ Manpreet Singh et al.: Preprint submitted to Elsevier
(10)
Page 4 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions Table 1 Taxonomy of the 15 benchmark tabular datasets from OpenML. Dataset Name uci_default bank_marketing adult aps_failure diabetes130us miniboone fraud mammography sick_numeric wilt ozone_level_8hr seismic_bumps pc1 oil_spill credit_g
Domain Finance Marketing Demographics Industrial Healthcare Physics Finance Healthcare Healthcare Forestry Environment Geology Software Environment Finance
Samples (𝑛) 30,000 45,211 48,842 76,000 101,766 130,064 284,807 11,183 3,772 4,839 2,534 2,584 1,109 937 1,000
Features (𝑑) 23 16 14 170 47 50 29 6 29 5 72 18 21 49 20
Minority Class (%) 22.1% 11.7% 23.9% 1.7% 11.2% 28.1% 0.17% 2.3% 6.1% 5.4% 6.3% 6.6% 6.9% 5.3% 30.0%
Imbalance Ratio 3.5 : 1 7.5 : 1 3.2 : 1 58.0 : 1 7.9 : 1 2.6 : 1 580.0 : 1 42.5 : 1 15.4 : 1 17.5 : 1 14.9 : 1 14.2 : 1 13.5 : 1 17.9 : 1 2.3 : 1
3.4. Theoretical Break-Even Derivation ∗ below which cost-controlled conformal deferral achieves We derive the critical human review cost threshold 𝐶rev lower total expected decision cost than uncalibrated Bayes-optimal threshold tuning. 𝐶FP Let 𝜏 ∗ = 𝐶 +𝐶 denote the Bayes-optimal probability decision threshold. A point predictor using threshold 𝜏 ∗ FN FP yields false positive rate FPR𝜏 ∗ and false negative rate FNR𝜏 ∗ . The expected per-instance cost of point classification 𝔼[𝐿point ] is:
(11)
𝔼[𝐿point ] = (1 − 𝜋1 ) ⋅ 𝐶FP ⋅ FPR𝜏 ∗ + 𝜋1 ⋅ 𝐶FN ⋅ FNR𝜏 ∗
̂ For the cost-controlled Mondrian conformal predictor, let 𝑟abs = 𝑃 (|𝐶(𝑋)| ≠ 1) denote the total abstention rate. ̂ ̂ Let FPRconf = 𝑃 (𝐶(𝑋) = {1} ∣ 𝑌 = 0) denote the automated false positive rate, and let FNRconf = 𝑃 (𝐶(𝑋) = {0} ∣ 𝑌 = 1) denote the automated false negative rate on singleton predictions. The expected per-instance cost 𝔼[𝐿conf ] is: (12)
𝔼[𝐿conf ] = (1 − 𝜋1 ) ⋅ 𝐶FP ⋅ FPRconf + 𝜋1 ⋅ 𝐶FN ⋅ FNRconf + 𝑟abs ⋅ 𝐶rev ∗ : Setting 𝔼[𝐿conf ] < 𝔼[𝐿point ] and solving for 𝐶rev yields the break-even human review threshold 𝐶rev ∗ 𝐶rev =
(1 − 𝜋1 )𝐶FP (FPR𝜏 ∗ − FPRconf ) + 𝜋1 𝐶FN (FNR𝜏 ∗ − FNRconf ) 𝑟abs
(13)
∗ , routing set-valued ambiguities to human experts guarantees a net When operational review cost satisfies 𝐶rev < 𝐶rev reduction in expected financial or clinical decision costs.
4. Experimental Setup This section outlines the benchmark datasets, machine learning models, baseline methods, and statistical metrics used to evaluate cost-sensitive conformal prediction.
4.1. Benchmark Datasets & Taxonomy We evaluate the framework across 15 public tabular datasets from OpenML, covering credit scoring, financial fraud, medical diagnosis, industrial safety, remote sensing, and environmental monitoring. Table 1 summarizes the dataset characteristics, sample counts (𝑛), feature counts (𝑑), minority class percentages (𝜋1 ), and imbalance ratios. The individual characteristics of the 15 benchmark datasets are detailed below: 1. uci_default (OpenML ID 42477): Contains 30,000 credit card clients from Taiwan. The dataset includes 23 features covering demographic data, historical monthly repayment status, and billing statement amounts. The minority positive class (𝑌 = 1) represents credit default (22.1% prevalence, 3.5:1 imbalance). Manpreet Singh et al.: Preprint submitted to Elsevier
Page 5 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
2. bank_marketing (OpenML ID 1461): Consists of 45,211 direct phone marketing records from a Portuguese banking institution. Features include client demographics, past contact history, and macroeconomic indicators. The positive class represents term deposit subscription (11.7% prevalence, 7.5:1 imbalance). 3. adult (OpenML ID 1590): Extracted from the 1994 US Census database, comprising 48,842 records with 14 demographic and employment features (age, education, occupation, capital gain, weekly hours). The positive class corresponds to annual income exceeding $50,000 (23.9% prevalence, 3.2:1 imbalance). 4. aps_failure (OpenML ID 41138): Industrial component failure dataset from Scania heavy trucks, containing 76,000 records and 170 anonymized sensor features. The positive class denotes component failure in the Air Pressure System (1.7% prevalence, 58.0:1 imbalance), where false negatives incur heavy vehicle breakdown costs. 5. diabetes130us (OpenML ID 4541): Clinical dataset comprising 101,766 hospital admissions across 130 US hospitals over 10 years. Features include lab measurements, prescribed medications, and diagnostic codes. The positive class represents early hospital readmission within 30 days (11.2% prevalence, 7.9:1 imbalance). 6. miniboone (OpenML ID 41150): High-energy physics dataset from the MiniBooNE experiment, containing 130,064 event records with 50 particle beam trajectory features. The positive class corresponds to electron neutrino signal events against muon background noise (28.1% prevalence, 2.6:1 imbalance). 7. fraud (OpenML ID 1597): European credit card transaction dataset containing 284,807 credit card transactions over two days. Features consist of 28 principal components derived via PCA alongside transaction time and amount. The positive class represents fraudulent transactions (0.17% prevalence, 580.0:1 imbalance). 8. mammography (OpenML ID 310): Radiological screening dataset containing 11,183 observations with 6 calcification cluster features. The positive class represents malignant breast lesions (2.3% prevalence, 42.5:1 imbalance). 9. sick_numeric (OpenML ID 41946): Thyroid condition dataset containing 3,772 patient records with 29 clinical and serum thyroid hormone measurements. The positive class represents a confirmed sick thyroid diagnosis (6.1% prevalence, 15.4:1 imbalance). 10. wilt (OpenML ID 40983): High-resolution remote sensing dataset comprising 4,839 satellite image segments with 5 spectral texture features. The positive class corresponds to diseased tree canopy segments (5.4% prevalence, 17.5:1 imbalance). 11. ozone_level_8hr (OpenML ID 1487): Meteorological dataset containing 2,534 daily ground-level ozone readings with 72 meteorological variables (temperature, wind speed, solar radiation). The positive class represents high-ozone peak alert days (6.3% prevalence, 14.9:1 imbalance). 12. seismic_bumps (OpenML ID 45562): Underground mining safety dataset containing 2,584 acoustic monitoring records with 18 seismic sensor features. The positive class represents hazardous seismic rock burst events (6.6% prevalence, 14.2:1 imbalance). 13. pc1 (OpenML ID 1068): Software engineering defect dataset from a NASA satellite flight software system, containing 1,109 code modules with 21 McCabe and Halstead complexity metrics. The positive class represents defective software modules (6.9% prevalence, 13.5:1 imbalance). 14. oil_spill (OpenML ID 311): Satellite ocean observation dataset containing 937 ocean patch images with 49 shape and texture features. The positive class represents confirmed oil spills against natural ocean slicks (5.3% prevalence, 17.9:1 imbalance). 15. credit_g (OpenML ID 31): German credit dataset containing 1,000 loan applicants with 20 financial, personal, and credit history attributes. The positive class represents bad credit risk loans (30.0% prevalence, 2.3:1 imbalance).
4.2. Base Classifiers & Probability Calibration To evaluate framework performance independently of model architecture, we examine 7 distinct machine learning algorithms covering linear models, decision tree ensembles, and probabilistic generative models: 1. HistGradientBoosting (HGB): Histogram-based gradient boosted decision tree algorithm based on LightGBM principles. HGB discretizes continuous features into 256 integer bins, reducing split-finding computational complexity from 𝑂(𝑛 log 𝑛) to 𝑂(𝑛) per node split. Trees are built sequentially to minimize negative loglikelihood loss (𝑦, 𝑓 ) = log(1 + 𝑒−𝑦𝑓 ) using a learning rate of 𝜂 = 0.1, a maximum of 100 iterations, and Manpreet Singh et al.: Preprint submitted to Elsevier
Page 6 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
a minimum of 20 samples per leaf. HGB provides fast execution and high predictive accuracy on large tabular datasets (𝑛 > 100, 000, such as ‘fraud‘ and ‘miniboone‘), but raw tree leaf probability estimates tend to saturate near 0 or 1, producing sharp non-conformity distributions that benefit from calibration. 2. Logistic Regression (LogReg): L2-regularized linear classification model mapping feature combinations 𝐰𝑇 𝑋 + 𝑏 through the logistic sigmoid function 𝜎(𝑧) = (1 + 𝑒−𝑧 )−1 . LogReg minimizes regularized binary cross-entropy loss: 𝑛 ∑ 1 log(1 + exp(−𝑦𝑖 (𝐰𝑇 𝑋𝑖 + 𝑏))) min ‖𝐰‖22 + 𝐶 𝐰,𝑏 2 𝑖=1
(14)
using the L-BFGS quasi-Newton solver with inverse regularization strength 𝐶 = 1.0 and a maximum of 1,000 iterations. Input features are standardized via zero-mean, unit-variance scaling. Linear decision boundaries struggle to capture non-linear feature interactions under high-dimensional tabular imbalance, resulting in smooth ̂ posterior probability transitions across class boundaries and larger Mondrian prediction sets (|𝐶(𝑋)| ≈ 1.54). 3. Random Forest (RF): Ensemble classifier composing 100 decision trees trained via √ bootstrap aggregation (bagging). Variance reduction is introduced by sampling a random subset of 𝑚 = ⌊ 𝑑⌋ features at each split node. Individual decision trees are grown to full depth using the Gini impurity criterion without pruning. Posterior probability 𝑃̂ (𝑌 = 1 ∣ 𝑋) is calculated by averaging leaf node sample fractions across all 100 trees: 𝑇 1∑ ̂ 𝑃RF (𝑌 = 1 ∣ 𝑋) = 𝑃 (𝑌 = 1 ∣ 𝑋) 𝑇 𝑡=1 𝑡
(15)
Averaging tree proportions softens extreme confidence estimates, but bagging on imbalanced datasets biases leaf distributions toward the majority class, requiring class-conditional quantiles (𝑞0 , 𝑞1 ) to restore minority validity. 4. Extra Trees (ET): Extremely Randomized Trees ensemble composing 100 decision trees. ET draws cut-point thresholds completely at random for each candidate feature at every node split rather than searching for optimal Gini split points. This random split mechanism suppresses model variance and reduces overfitting on noisy tabular data while accelerating tree construction times. Random split points create smooth probability surfaces around minority clusters, yielding stable non-conformity quantiles (𝑞1 ) across random cross-validation folds. 5. Gradient Boosting (GB): Classical gradient boosting framework constructing 100 sequential decision trees. Trees are fit iteratively to pseudo-residuals (the negative gradient of binary cross-entropy loss) from previous ensemble stages: ] [ 𝜕(𝑦𝑖 , 𝑓 (𝑋𝑖 )) 𝑟𝑖𝑚 = − (16) 𝜕𝑓 (𝑋𝑖 ) 𝑓 (𝑋)=𝑓 (𝑋) 𝑚−1
Models are configured with a maximum tree depth of 3, learning rate 𝜂 = 0.1, and full subsample ratios. Shallow decision trees prevent overfitting on small minority samples, producing clear probability separation that yields efficient singleton prediction sets when combined with Isotonic calibration. 6. AdaBoost: Adaptive Boosting ensemble utilizing 50 decision stumps (decision trees with maximum depth 1). AdaBoost sequentially adjusts sample weights 𝑤(𝑡) 𝑖 after each iteration, increasing weights for misclassified instances: ( ) 𝑤(𝑡+1) = 𝑤(𝑡) (17) 𝑖 𝑖 exp 𝛼𝑡 ⋅ 𝕀(𝑦𝑖 ≠ ℎ𝑡 (𝑋𝑖 )) ( ) 1−𝜖 Classifier outputs are weighted by stage confidence coefficients 𝛼𝑡 = 12 log 𝜖 𝑡 . Decision stumps focus heavily 𝑡 on minority-class misclassifications in early iterations, but AdaBoost is sensitive to noisy features and outlier samples (such as in ‘oil_spill‘ and ‘seismic_bumps‘), generating volatile raw probability scores that require calibration. 7. Gaussian Naive Bayes (GNB): Probabilistic generative classifier applying Bayes’ rule under the conditional independence assumption. GNB models feature likelihoods 𝑃 (𝑋𝑗 ∣ 𝑌 = 𝑦) as independent 1D Gaussian
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 7 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions 2 ) for each feature 𝑗 ∈ {1, … , 𝑑}: distributions (𝜇𝑗𝑦 , 𝜎𝑗𝑦
𝑃 (𝑋 ∣ 𝑌 = 𝑦) =
𝑑 ∏ 𝑗=1
1
(
exp − √ 2 2𝜋𝜎𝑗𝑦
(𝑋𝑗 − 𝜇𝑗𝑦 )2
) (18)
2 2𝜎𝑗𝑦 𝜋 𝑃 (𝑋∣𝑌 =1)
Posterior probabilities are derived via Bayes’ theorem: 𝑃̂ (𝑌 = 1 ∣ 𝑋) = 𝜋 𝑃 (𝑋∣𝑌1=0)+𝜋 𝑃 (𝑋∣𝑌 =1) . GNB serves 0 1 as a non-ensemble, non-linear probabilistic baseline. When tabular features exhibit strong pairwise correlations, GNB’s independence assumption leads to severe overconfidence (probabilities pushed to 0.0 or 1.0), testing the robustness of Mondrian CP under extreme probability distortion. Probability outputs 𝑃̂ (𝑌 = 𝑦 ∣ 𝑋) are evaluated across 3 calibration regimes fit on a separate 20% calibration split: 1. Uncalibrated (None): Raw probability outputs 𝑃̂ (𝑌 = 1 ∣ 𝑋) produced directly by model ‘.predict_proba()‘ methods without post-processing. 2. Platt Scaling (Sigmoid): Parametric probability calibration method that fits a single-variable logistic regression model on uncalibrated model decision scores 𝑓 (𝑋): 𝑃̂Platt (𝑌 = 1 ∣ 𝑋) =
1 1 + exp(𝐴 ⋅ 𝑓 (𝑋) + 𝐵)
(19)
Parameters 𝐴, 𝐵 ∈ ℝ are estimated via maximum likelihood. Platt scaling corrects global probability distortion when raw model log-odds follow a Gaussian distribution. 3. Isotonic Regression (Isotonic): Non-parametric probability calibration method that fits a monotonic step function 𝑃̂Iso (𝑌 = 1 ∣ 𝑋) = 𝑚(𝑓 (𝑋)) using the Pool Adjacent Violators Algorithm (PAVA). Isotonic regression corrects arbitrary non-linear probability distortions without imposing functional form assumptions, provided calibration sample size is sufficient (𝑛cal ≥ 200).
4.3. Baseline Methods for Comparison We compare the cost-controlled Mondrian conformal predictor against 5 competitive baseline decision strategies: 1. Default Point Predictor (𝜏 = 0.5): Standard binary classification assigning class 1 if 𝑃̂ (𝑌 = 1 ∣ 𝑋) ≥ 0.5. 2. Bayes Cost-Tuned Threshold (𝜏 ∗ ): Threshold tuned to minimize empirical cost on calibration data (𝜏 ∗ = 𝐶FP ) [14]. 𝐶 +𝐶 FN
FP
3. Marginal Conformal Predictor (Marginal CP): Standard split conformal prediction targeting global 1 − 𝛼 coverage [3]. 4. Confidence Rejector: Abstains on predictions where maximum posterior probability max𝑦 𝑃̂ (𝑌 = 𝑦 ∣ 𝑋) < 1 − 𝛾, delegating low-confidence cases to human review [8]. 5. Conformal Risk Control Rejector: Bounded risk selective classifier using conformal risk bounds [9].
4.4. Evaluation Metrics & Statistical Testing All experiments run across 10 random seeds per dataset-model-calibration combination (3,150 total experimental runs). Models are evaluated on held-out test folds using five metrics: 1. Minority Class Coverage (1 − 𝛼): Percentage of true positive samples contained within prediction sets (target 1 − 𝛼 = 0.90). 2. Majority Class Coverage: Percentage of true negative samples contained within prediction sets. ̂ 3. Average Set Size (|𝐶(𝑋)|): Mean number of labels per prediction set. ̂ 4. Abstention Rate (𝑟abs ): Fraction of test samples where |𝐶(𝑋)| ≠ 1, triggering human review. 5. Expected Per-Instance Decision Cost (𝔼[𝐿]): Mean financial/clinical cost per sample under cost matrix parameters 𝐶FP = 1.0, 𝐶FN = 10.0, 𝐶rev ∈ [0.1, 5.0]. Statistical significance across paired model runs is assessed using two-tailed Wilcoxon signed-rank tests (𝑝 < 0.05), with 95% bootstrap confidence intervals computed over 1,000 resamples. Manpreet Singh et al.: Preprint submitted to Elsevier
Page 8 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions Table 2 ̂ Summary of minority class coverage (𝑌 = 1), average set size (|𝐶(𝑋)|), abstention rate (𝑟abs ), and mean per-instance decision cost (𝔼[𝐿]) across key methods (averaged over all 15 datasets, 7 models, and 3 calibration settings; target coverage 1 − 𝛼 = 90%). Method Bayes Cost-Tuned (𝜏 ∗ ) Marginal CP APS CP RAPS CP Mondrian CP (Raw Sets) Cost-Controlled Mondrian
Minority Coverage (%) 80.0% 30.5% 55.6% 48.7% 92.2% 97.7%
Avg Set Size 1.00 1.04 1.09 1.07 1.36 1.71
Abstention Rate (%) 0.0% 12.5% 19.8% 15.7% 38.3% 71.9%
Mean Cost (𝔼[𝐿]) 0.340 0.413 0.380 0.444 0.143 0.021
4.5. Reproducibility All experiments are fully deterministic given a fixed set of ten random seeds, {7, 19, 31, 42, 101, 202, 303, 404, 505, 606}, which govern the train/calibration/test partition and every stochastic component of model fitting and conformal scoring. Each of the 15 datasets is evaluated across 7 base classifiers and 3 probability-calibration settings, giving 15 × 7 × 3 = 315 configurations, each averaged over the ten seeds (3,150 model fits in total). The target error level is fixed at 𝛼 = 0.10 throughout; the cost matrix uses 𝐶FP = 1, 𝐶FN = 10; the deferral-cost sweep spans 𝐶rev ∈ {0, 0.5, 1.0, 2.0}; and confidence intervals use 1,000 bootstrap resamples. Experiments were run with Python 3.14, scikit-learn 1.9.0, NumPy 2.5.1 and pandas 3.0.3. All 15 datasets are publicly available from OpenML and are downloaded programmatically by the released code, so no manual data preparation is required.
5. Empirical Results & Discussion This section presents empirical findings across 3,150 experimental runs evaluating 15 OpenML tabular benchmarks. We analyze minority coverage restoration, cost reduction performance, probability calibration effects, parameter sensitivity, and domain case studies.
5.1. Minority Class Coverage Restoration Figure 1 compares minority class coverage (𝑌 = 1) achieved by marginal conformal prediction versus classconditional (Mondrian) conformal prediction across the 15 benchmark datasets at target error level 𝛼 = 0.10 (90% nominal coverage). Marginal conformal prediction suffers from severe coverage degradation on highly imbalanced datasets. On ‘aps_failure‘ (imbalance ratio 58:1), marginal CP minority coverage collapses below 1% under gradient-boosted models, and it remains far below target on other rare-event tasks such as ‘seismic_bumps‘ (4.0%) and ‘pc1‘ (7.6%), despite satisfying the global 90% marginal coverage target. Across all 315 dataset-model combinations, marginal CP achieves an average minority class coverage of only 30.5%. In contrast, Mondrian conformal prediction computes class-specific quantiles 𝑞0 and 𝑞1 , restoring minority class coverage to an average of 92.2% across all datasets. As documented in Table 2, Mondrian CP yields an average minority coverage gain of 61.71 percentage points over marginal CP (95% CI [58.20%, 65.22%], paired Wilcoxon sign test 𝑝 = 1.52 × 10−82 ). Figure 2 breaks down coverage performance across alternative conformal set construction methods. Figure 3 makes the mechanism explicit: minority coverage under marginal CP degrades monotonically as the imbalance ratio grows, whereas Mondrian CP holds the target across the full range. To compare all decision strategies simultaneously, Figure 4 presents a Friedman–Nemenyi critical-difference diagram of minority-coverage ranks across the 15 datasets. Marginal CP ranks last and lies more than one critical difference away from every class-conditional method, establishing that its minority-coverage deficit is statistically significant rather than dataset-specific.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 9 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
Figure 1: Empirical minority-class coverage (𝑌 = 1) at target error level 𝛼 = 0.10 (90% coverage guarantee) across 15 OpenML imbalanced tabular datasets. Marginal conformal prediction fails on severe imbalance, whereas Mondrian conformal prediction maintains valid coverage across all datasets.
Figure 2: Comparison of overall coverage, majority class coverage (𝑌 = 0), and minority class coverage (𝑌 = 1) across Marginal CP, Mondrian CP, APS, and RAPS methods. Only Mondrian CP guarantees balanced coverage across both classes.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 10 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
Figure 3: Minority-class coverage versus dataset imbalance ratio (log scale) for the HistGradientBoosting base model. As imbalance grows, marginal CP (grey circles) collapses far below the 90% target, while Mondrian CP (red triangles) remains at target across all 15 datasets.
Figure 4: Critical-difference (Nemenyi) diagram ranking eight decision methods by minority-class coverage across the 15 datasets (Friedman test; lower average rank is better; CD = 2.71 at 𝛼 = 0.05). Class-conditional methods dominate; marginal CP is significantly worst.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 11 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
5.2. Decision Cost Reduction & Deferral Frontier Integrating Mondrian prediction sets with the cost-controlled abstention rule (𝑌̂action ) reduces total expected decision costs compared to point predictors and baseline rejectors. Figure 5 illustrates the cost-abstention Pareto frontier as human review cost 𝐶rev ranges from 0.1 to 5.0 (with 𝐶FP = 1.0 and 𝐶FN = 10.0). At modest review costs (𝐶rev = 0.5), cost-controlled Mondrian CP achieves a mean per-instance cost of 0.313, representing a 54.1% cost reduction compared to default point classification (0.682) and a 38.6% cost reduction compared to Bayes cost-tuned thresholding (0.510). Paired Wilcoxon tests confirm that cost-controlled Mondrian CP achieves lower expected decision costs than Bayes thresholding (𝑝 = 2.93 × 10−40 ), confidence rejectors (𝑝 = 1.18 × 10−72 ), and risk-controlled rejectors (𝑝 = 5.12 × 10−54 ).
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 12 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
Figure 5: Pareto frontier of expected decision cost (𝔼[𝐿]) versus abstention rate (𝑟abs ) across varying human review costs 𝐶rev ∈ [0.1, 5.0]. Cost-controlled Mondrian conformal deferral achieves lower expected cost than Bayes thresholding and confidence rejectors across a broad operational review cost range.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 13 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
5.3. Impact of Probability Calibration Probability calibration directly affects conformal prediction set cardinalities and abstention rates. Table 3 reports the performance of Mondrian CP across uncalibrated base models, Platt scaling (Sigmoid), and Isotonic regression. Isotonic regression produces tighter probability distributions near decision boundaries, reducing average set size ̂ |𝐶(𝑋)| from 1.42 to 1.35 on random forest models and lowering abstention rates by 6.8 percentage points without violating class-conditional coverage. The coverage restoration is not an artifact of any single base model. Figure 6 reports minority coverage for every base model × method combination: the Mondrian and cost-controlled Mondrian columns remain at target across all seven classifiers, whereas marginal CP is uniformly low.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 14 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
Table 3 Effect of probability calibration on Mondrian CP efficiency and decision cost across base model families. Model Family Gradient Boosted (HGB/GB)
Random Forests (RF/ET)
Linear (LogReg)
Calibration None Sigmoid Isotonic None Sigmoid Isotonic None Sigmoid Isotonic
Minority Coverage (%) 91.5% 91.3% 94.3% 90.6% 90.9% 94.0% 90.4% 90.9% 94.9%
Avg Set Size 1.30 1.31 1.44 1.22 1.23 1.40 1.28 1.30 1.49
Abstention Rate (%) 34.3% 34.4% 46.0% 26.2% 27.4% 41.3% 30.6% 32.6% 46.5%
Mean Cost (𝔼[𝐿]) 0.150 0.150 0.109 0.156 0.154 0.118 0.182 0.178 0.351
Figure 6: Minority-class coverage for every base model × method combination (averaged over datasets and calibrations). Class-conditional methods (Mondrian, cost-controlled Mondrian) stay near target across all seven base models, confirming the result is model-agnostic.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 15 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
5.4. Sensitivity & Ablation Analysis ∗ ) across varying false negative Figure 7 evaluates the sensitivity of decision costs and break-even thresholds (𝐶rev penalties 𝐶FN ∈ [5, 50] and human review costs 𝐶rev ∈ [0.1, 5.0]. As false negative penalty 𝐶FN increases from 5.0 to 50.0, the economic benefit of human deferral expands. The ∗ rises from 0.82 to 4.35, demonstrating that under severe asymmetric error penalties, break-even review threshold 𝐶rev higher human review expenditures remain economically advantageous. When accounting for noisy human review errors (𝜖hum > 0), net cost savings persist provided 𝜖hum < 0.18 under standard operating parameters.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 16 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
∗ Figure 7: Sensitivity analysis of expected decision cost (𝔼[𝐿]) and break-even review threshold (𝐶rev ) across varying false negative penalties (𝐶FN ) and noisy human expert error rates (𝜖hum ).
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 17 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
5.5. Real-World Case Studies We examine three domain case studies: 1. Credit Default Risk (‘uci_default‘): Imbalance ratio 3.5:1 (𝑛 = 30, 000). Marginal CP yields 57.7% minority coverage. Mondrian CP restores minority coverage to 90.3%, while cost-controlled deferral at 𝐶rev = 0.5 reduces expected default losses by 13.6% compared to Bayes threshold tuning. 2. Industrial Component Failure (‘aps_failure‘): Extreme imbalance ratio 58:1 (𝑛 = 76, 000, 𝜋1 = 1.7%). Marginal CP minority coverage fails below 1% (0.8%) under gradient-boosted models. Mondrian CP restores failure detection to the 90% target (89.8%); coupling it with cost-controlled abstention routes 34.1% of ambiguous cases to technicians and reduces expected breakdown costs by 65.3% relative to Bayes threshold tuning. 3. Medical Readmission Risk (‘diabetes130us‘): Imbalance ratio 7.9:1 (𝑛 = 101, 766, 𝜋1 = 11.2%). Costcontrolled Mondrian CP achieves 94.9% minority coverage with a 24.1% abstention rate, preventing high-cost unmonitored patient discharges.
6. Discussion & Practical Deployment Guidelines This section discusses system architecture, operational economic modeling under capacity limits, and threats to validity for deploying cost-sensitive conformal prediction in knowledge-based decision support environments.
6.1. System Architecture for Knowledge-Based Systems Deploying cost-sensitive conformal prediction into enterprise knowledge-based systems requires a four-tier operational architecture designed to maintain finite-sample coverage guarantees without interrupting real-time decision workflows: 1. Inference & Non-Conformity Scoring Tier: The base machine learning classifier computes posterior probability vector 𝑃̂ (𝑌 ∣ 𝑋) for incoming sample 𝑋. The scoring engine derives non-conformity scores 𝑆(𝑋, 𝑦) = 1 − 𝑃̂ (𝑌 = 𝑦 ∣ 𝑋) for all candidate labels 𝑦 ∈ {0, 1}. 2. Mondrian Quantile Server: A persistent calibration service maintains class-specific quantiles (𝑞0 , 𝑞1 ) estimated over calibration split 𝐷cal . Non-conformity scores are evaluated against class thresholds to form prediction set ̂ 𝐶(𝑋) = {𝑦 ∈ {0, 1} ∣ 𝑆(𝑋, 𝑦) ≤ 𝑞𝑦 }. ̂ 3. Cost-Optimal Routing Engine: The decision engine applies action rule 𝑌̂action (𝑋). Singletons (|𝐶(𝑋)| = 1) execute automatically at zero deferral cost. Ambiguous multi-label sets ({0, 1}) and out-of-distribution empty ̂ sets (∅) trigger abstention (|𝐶(𝑋)| ≠ 1) and pass to the human review queue with set-valued uncertainty metadata attached [26, 24]. 4. Audit & Recalibration Pipeline: All automated actions, human deferral resolutions, and downstream groundtruth outcomes are logged to an immutable audit store. When new validated labels accumulate, the quantile server updates thresholds (𝑞0 , 𝑞1 ) asynchronously to maintain validity under minor distributional fluctuations [28].
6.2. Economic Viability & Human-in-the-Loop Capacity Limits In practical operations, human expert review capacity is finite. Let 𝜆 denote the arrival rate of incoming decision ̂ requests, and let 𝑟abs = 𝑃 (|𝐶(𝑋)| ≠ 1) represent the conformal abstention rate. The arrival rate at the human review queue is 𝜆rev = 𝜆 ⋅ 𝑟abs . Let 𝑘 represent the number of active human experts, each serving reviews at service rate 𝜇. The human review subsystem behaves as an 𝑀∕𝑀∕𝑘 queuing system. To prevent infinite queue growth and unbounded wait times, the system must satisfy the stability condition: 𝜌=
𝜆 ⋅ 𝑟abs <1 𝑘⋅𝜇
(20)
When arrival rate 𝜆rev exceeds total human processing capacity 𝑘𝜇, review queues accumulate, generating operational delay penalty 𝐶delay (𝑤) proportional to wait time 𝑤. The effective per-instance cost under queue congestion expands to: 𝔼[𝐿congested ] = 𝔼[𝐿conf ] + 𝑃 (𝑊 > 0) ⋅ 𝔼[𝐶delay (𝑊 )] Manpreet Singh et al.: Preprint submitted to Elsevier
(21) Page 18 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
Where 𝑊 is the random variable representing queue wait time. When human review capacity is constrained to budget 𝐵rev (maximum review rate 𝑟max ), the system dynamically tunes misclassification allowance 𝛼(𝑡) to bound abstention rate 𝑟abs (𝛼) ≤ 𝑟max . Increasing 𝛼 slightly shrinks prediction set cardinalities, reducing human review volume while preserving distribution-free class-conditional coverage at the adjusted target 1 − 𝛼(𝑡).
6.3. Threats to Validity & Limitations Five methodological threats to validity affect real-world deployment: 1. Non-Stationary Distribution Shift: Standard conformal validity relies on exchangeability between calibration set 𝐷cal and test set 𝐷test . In deployment, temporal covariate shift or concept drift invalidates static quantiles (𝑞0 , 𝑞1 ), leading to empirical coverage under-shooting [22, 31]. Online conformal algorithms with adaptive learning rates 𝛾𝑡 must be implemented to track time-varying quantiles under drift [11]. 2. Multi-Class Combinatorial Complexity: This benchmark evaluates binary imbalanced tasks. In multi-class settings (𝐾 > 2), cost matrix 𝐂 ∈ ℝ𝐾×𝐾 contains 𝐾(𝐾 − 1) asymmetric off-diagonal penalties. Evaluating 2𝐾 candidate set configurations introduces combinatorial complexity, requiring hierarchical Mondrian partitions [25]. 3. Sparse Calibration Quantile Variance: On datasets with extreme class imbalance (𝜋1 < 0.1%, such as financial fraud), calibration set 𝐷cal contains few minority samples (𝑛1 < 30). Small sample sizes increase quantile estimation variance, making threshold 𝑞1 = 𝑆(𝑘1 ) sensitive to sampling noise. Smoothed non-conformity density estimation or clustered Mondrian calibration should be used when 𝑛1 < 50 [10]. 4. Human Expert Cognitive Fatigue: The cost model assumes human review error rate 𝜖hum remains constant. In practice, high abstention volume increases cognitive load, inducing expert fatigue and elevating 𝜖hum over long shifts. System design must incorporate workload capping and expert decision verification. 5. Tabular Feature Representation Limits: Benchmark evaluations rely on tabular features. Extending costsensitive conformal deferral to high-dimensional unstructured modalities (medical imagery, clinical notes, audio streams) requires calibrating learned representations before non-conformity scoring.
7. Conclusion & Future Work This section summarizes the primary theoretical and empirical contributions of this work and outlines directions for future research.
7.1. Summary of Findings This paper presented a benchmark and theoretical framework for cost-sensitive conformal prediction across 15 imbalanced tabular datasets, 7 classification models, 3 calibration methods, and 10 random seeds (3,150 experimental runs). Our findings demonstrate that standard marginal conformal prediction fails on imbalanced data, dropping below 1% minority coverage on extreme class ratios because global quantiles are dominated by majority-class samples. Classconditional (Mondrian) conformal prediction resolves this failure, restoring minority coverage to the 90% target (92.2% on average) across all datasets and yielding an average minority coverage improvement of 61.71 percentage points (𝑝 = 1.52 × 10−82 ). Furthermore, integrating Mondrian prediction sets with an asymmetric cost matrix (𝐶FN , 𝐶FP , 𝐶rev ) and action rule 𝑌̂action achieves lower total expected decision cost than standard point classifiers, Bayes cost-tuned thresholding, ∗ , confidence rejectors, and risk-controlled rejectors. We derived closed-form break-even human review threshold 𝐶rev establishing formal economic boundaries where routing ambiguous instances to human experts yields net cost savings under both expert oracle (𝜖hum = 0) and noisy human review (𝜖hum > 0) conditions.
7.2. Future Research Directions Four promising trajectories extend this work: 1. Online Adaptive Mondrian Conformal Inference: Developing adaptive online Mondrian algorithms that update class quantiles (𝑞0 (𝑡), 𝑞1 (𝑡)) dynamically under non-stationary streaming drift without requiring full retraining [11, 22]. Manpreet Singh et al.: Preprint submitted to Elsevier
Page 19 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions
2. LLM-Assisted Multi-Agent Expert Routing: Replacing generic human review queues with specialized large language model (LLM) agent panels that process set-valued ambiguities and provide structured rationale before human sign-off [21, 20]. 3. Fairness-Aware Cost-Conformal Prediction: Extending class-conditional cost deferral to multi-group fairness settings, ensuring that coverage and decision cost reductions are distributed equitably across sensitive demographic subgroups [23]. 4. Multi-Class and Structured Output Deferral: Generalizing cost-controlled Mondrian deferral to multi-class classification (𝐾 > 2), hierarchical taxonomies, and multi-label medical diagnostic outputs [25].
Data Availability All 15 benchmark datasets are publicly available from the OpenML repository (https://www.openml.org) and are retrieved programmatically by the accompanying code; no proprietary or restricted data were used. The code required to reproduce every table and figure in this paper is available at https://github.com/physics-vibes15/ cost-sensitive-conformal.
Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References [1] H. He and E. A. Garcia, “Learning from imbalanced data,” IEEE Transactions on Knowledge and Data Engineering, vol. 21, no. 9, pp. 1263–1284, 2009. [2] V. Vovk, A. Gammerman, and G. Shafer, Algorithmic Learning in a Random World. Springer Science & Business Media, 2005. [3] A. N. Angelopoulos and S. Bates, “A gentle introduction to conformal prediction and distribution-free uncertainty quantification,” arXiv preprint arXiv:2107.07511, 2021. [4] J. Lei, M. G’Sell, A. Rinaldo, R. J. Tibshirani, and L. Wasserman, “Distribution-free predictive inference for regression,” Journal of the American Statistical Association, vol. 113, no. 523, pp. 1094–1111, 2018. [5] Y. Romano, M. Sesia, and E. Candès, “Classification with valid and adaptive coverage,” Advances in Neural Information Processing Systems, vol. 33, pp. 3581–3591, 2020. [6] M. Sadinle, J. Lei, and L. Wasserman, “Least ambiguous set-valued classifiers with bounded error levels,” Journal of the American Statistical Association, vol. 114, no. 525, pp. 223–234, 2019. [7] R. El-Yaniv and R. Wiener, “On the foundations of noise-free selective classification,” Journal of Machine Learning Research, vol. 11, no. 53, pp. 1605–1641, 2010. [8] Y. Geifman and R. El-Yaniv, “Selective classification for deep neural networks,” Advances in Neural Information Processing Systems, vol. 30, pp. 4878–4887, 2017. [9] A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster, “Conformal risk control,” in International Conference on Learning Representations (ICLR), 2024. [10] T. Ding, A. N. Angelopoulos, S. Bates, M. I. Jordan, and R. J. Tibshirani, “Class-conditional conformal prediction with many classes,” Advances in Neural Information Processing Systems, vol. 36, 2023. [11] I. Gibbs and E. Candès, “Adaptive conformal inference under distribution shift,” Advances in Neural Information Processing Systems, vol. 34, pp. 1660–1672, 2021. [12] D. Stutz, K. Dvijotham, A. T. Cemgil, and A. Doucet, “Learning optimal conformal classifiers,” in International Conference on Learning Representations (ICLR), 2022. [13] Y. Shi, S. Ghosh, T. Belkhouja, J. R. Doppa, and Y. Yan, “Conformal prediction for class-wise coverage via augmented label rank calibration,” Advances in Neural Information Processing Systems, vol. 37, 2024. [14] C. Elkan, “The foundations of cost-sensitive learning,” in International Joint Conference on Artificial Intelligence (IJCAI), 2001, pp. 973–978. [15] B. Zadrozny and C. Elkan, “Transforming classifier scores into accurate multiclass probability estimates,” in ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2002, pp. 694–699. [16] G. M. Weiss, “Mining with rarity: a unifying framework,” ACM SIGKDD Explorations Newsletter, vol. 6, no. 1, pp. 7–19, 2004. [17] C. Chow, “On optimum recognition error and reject tradeoff,” IEEE Transactions on Information Theory, vol. 16, no. 1, pp. 41–46, 1970. [18] H. Mozannar and D. Sontag, “Consistent estimators for learning to defer to an expert,” in International Conference on Machine Learning (ICML), 2020, pp. 7076–7087. [19] E. Straitouri, L. Wang, A. Okati, and M. Gomez-Rodriguez, “Improving expert predictions with conformal prediction,” in International Conference on Machine Learning (ICML), 2023, pp. 32633–32653. [20] E. Straitouri and M. Gomez-Rodriguez, “Designing decision support systems using counterfactual prediction sets,” in International Conference on Machine Learning (ICML), 2024, pp. 46722–46744.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 20 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions [21] P. Hemmer, M. Schemmer, N. Kühl, and J. Vössing, “Human-AI complementarity in hybrid intelligence systems: A structured literature review,” in Pacific Asia Conference on Information Systems (PACIS), 2021. [22] R. F. Barber, E. J. Candès, A. Ramdas, and R. J. Tibshirani, “Conformal prediction beyond exchangeability,” The Annals of Statistics, vol. 51, no. 2, pp. 816–845, 2023. [23] Y. Romano, R. F. Barber, C. Sabatti, and E. Candès, “With malice toward none: Assessing uncertainty via equalized coverage,” Harvard Data Science Review, vol. 2, no. 2, 2020. [24] N. Charoenphakdee, Z. Cui, Y. Zhang, and M. Sugiyama, “Classification with rejection based on cost-sensitive classification,” in International Conference on Machine Learning (ICML), 2021, pp. 1507–1517. [25] M. Cauchois, S. Gupta, and J. C. Duchi, “Knowing what you know: Valid and validated confidence sets in multiclass and multilabel prediction,” Journal of Machine Learning Research, vol. 22, no. 81, pp. 1–42, 2021. [26] U. Bhatt, J. Antorán, Y. Zhang, Q. V. Liao, P. Sattigeri, R. Fogliato, et al., “Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty,” in Proc. AAAI/ACM Conf. on AI, Ethics, and Society (AIES), 2021, pp. 401–413. [27] H. Papadopoulos, K. Proedrou, V. Vovk, and A. Gammerman, “Inductive confidence machines for regression,” in European Conference on Machine Learning (ECML). Springer, 2002, pp. 345–356. [28] V. Vovk, “Conditional validity of inductive conformal predictors,” in Asian Conference on Machine Learning (ACML), 2012, pp. 475–490. [29] A. N. Angelopoulos, S. Bates, M. Jordan, and J. Malik, “Uncertainty sets for image classifiers using conformal prediction,” in International Conference on Learning Representations (ICLR), 2021. [30] J. Huang, H. Xi, L. Zhang, H. Yao, Y. Qiu, and H. Wei, “Conformal prediction for deep classifier via label ranking,” in International Conference on Machine Learning (ICML), 2024. [31] R. J. Tibshirani, R. Foygel Barber, E. Candès, and A. Ramdas, “Conformal prediction under covariate shift,” Advances in Neural Information Processing Systems, vol. 32, 2019. [32] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in International Conference on Machine Learning (ICML), 2017, pp. 1321–1330. [33] C. Cortes, G. DeSalvo, and M. Mohri, “Learning with rejection,” in International Conference on Algorithmic Learning Theory (ALT). Springer, 2016, pp. 67–82. [34] Y. Geifman and R. El-Yaniv, “SelectiveNet: A deep neural network with an integrated reject option,” in International Conference on Machine Learning (ICML), 2019, pp. 2151–2159.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 21 of 22
Cost-Sensitive Conformal Prediction for Imbalanced Decisions Manpreet Singh received the B.Tech. degree in Computer Science and Engineering from Guru Gobind Singh Indraprastha University in 2017, and the M.S. degree in Computer Information Systems, with a minor in Data Analytics, from Boston University in 2023, where he served as a Graduate Research Assistant. He is currently a Data Science Consultant with Infiheal Healthcare, Mumbai, India. He has over eight years of experience spanning healthcare analytics, semiconductor manufacturing, business process intelligence, and data engineering across ASML, Soroco, and Shore Infotech. His research spans machine learning, anomaly detection, explainable AI, natural language processing, and healthcare decision-support systems, with publications and papers under review in IEEE Access, Expert Systems with Applications, Information Sciences, and Artificial Intelligence Review. He is a Senior IEEE Member and serves as a reviewer for IEEE and international conferences on artificial intelligence and healthcare technologies, as well as for Discover
Akshatha Srikantha received the B.E. degree in Computer Science from MVJ College of Engineering, Bangalore, India, in 2021, and the M.S. degree in Computer Science from the University of California, Irvine, CA, USA, in 2025. Her work centers on machine learning and large language models, spanning autonomous LLM-based code reasoning agents, LLM alignment and safety evaluation, and model compression via knowledge distillation and parameter efficient fine tuning for efficient LLM inference. She previously worked as a Machine Learning Research Assistant at the Igarashi Lab, Department of Anatomy and Neurobiology, UC Irvine, analyzing neural and behavioral recordings in mouse models of Alzheimer’s disease, and as a Data Engineer at IBM, Bangalore, where she built large scale distributed data pipelines using PySpark and Apache Kafka on AWS. Her research interests include machine learning, deep learning, large language models, and explainable artificial intelligence. She is the author of a publication on Emotional Stress Recognition system using EEG and psychophysiological signals at IEEE ICAECA 2021.
Manpreet Singh et al.: Preprint submitted to Elsevier
Page 22 of 22