Single-Query Black-Box Calibration Auditing via Logit Bias
arXiv:2609.05125v1 [cs.LG] 4 Sep 2026
Roman Plaud Institut Polytechnique de Paris Onepoint, France
Antoine Saillenfest Onepoint, France
Thomas Bonald Institut Polytechnique de Paris
Matthieu Labeau Institut Polytechnique de Paris
Willem Waegeman Ghent University
Abstract Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.
1
Introduction
Standard calibration metrics for assessing the reliability of classifiers, such as the Expected Calibration Error (ECE) [12, 6], inherently require access to the model’s continuous output probabilities. When Large Language Models (LLMs) are deployed as zero-shot or few-shot classifiers on benchmarks like MMLU [7] or binary Question-Answering tasks [2], classification is typically performed by extracting the raw probabilities of specific target tokens (e.g., "True" versus "False" or "A" versus "B") [9]. However, for commercial, black-box models accessed via APIs, these continuous token probabilities are frequently restricted. API providers often hide output logits (logprobs) to protect proprietary architectures and prevent model distillation or imitation attacks [1]. This opacity severely limits independent auditing and practically prohibits calibration evaluation. To bypass this limitation, we exploit the logit_bias parameter, a feature exposed by some major APIs (such as OpenAI [13]), originally intended to give users semantic control over generation, such as suppressing specific vocabulary or enforcing formatting. We demonstrate that this parameter can be mathematically manipulated to evaluate exact probability thresholds. Querying these probability thresholds allows us to compute the calibration error of an opaque model over a full dataset without ever extracting a continuous logit. Contributions. In this paper, we introduce a novel empirical calibration estimator for black-box APIs equipped with a logit_bias parameter. Our method strictly requires only one API query per sample. We rigorously decompose the estimator’s bias and variance, proving its asymptotic consistency to the True Calibration Error (TCE). Our approach therefore provides an efficient, mathematically grounded framework for calibration auditing of black-box LLMs.
Preprint.
2
Related Work
Evaluating LLM calibration traditionally requires white-box access to exact output distributions. When native logprobs are hidden, researchers typically fall back on proxy methods that suffer from severe cost or accuracy limitations. Our single-query estimator directly resolves these deficiencies. Proxy Baselines: Verbalization and Sampling. When exact probabilities are inaccessible, standard approaches estimate confidence through generated text or sampling behavior. The simplest zero-shot baseline prompts the LLM to explicitly verbalize its certainty [11], though this is notoriously sensitive to framing and prone to sycophancy [18]. While more sophisticated techniques—such as reasoning decomposition or multi-prompt aggregation—can improve verbalized calibration [19], they remain subjective. Alternatively, self-consistency sampling estimates confidence via the empirical frequency of the majority class across multiple generations [17, 10]. However, generating multiple responses incurs a prohibitive query cost for large-scale datasets. Exact Extraction via Logit Bias. To bypass the unreliability of generative proxies, recent work introduces Iterative Logit Extraction. [1] demonstrated that continuous logits can be extracted from black-box models by iteratively manipulating the logit_bias parameter and observing output changes. While successful, this model-stealing approach relies on a query-intensive binary search to recover arbitrary precision. Our work avoids multi-query continuous extraction; by recognizing that binned calibration only requires evaluating discrete bounds, we mathematically map ECE thresholds to predefined logit biases, reducing black-box calibration evaluation to a single API call per sample.
3
Background: Traditional ECE Computation
To contextualize the necessity of our black-box approach, we briefly review standard calibration metrics to highlight their fundamental reliance on exact continuous probabilities. For a binary dataset S = {(Xi , Yi )}N i=1 , let f (X) ∈ (0, 1) denote the model’s exact continuous predicted probability for the positive class (e.g., the “True” token) and Y ∈ {0, 1} the true label. The standard Expected Calibration Error (ECE) [12, 6] approximates the True Calibration Error, TCE = E[|E[Y |f (X)] − f (X)|], by partitioning these predictions into M discrete bins. Letting Bm denote the set of sample indices falling into the m-th bin, the empirical ECE is defined as: [ bin = ECE
M X |Bm | m=1
N
where the empirical accuracy is acc(Bm ) = P conf(Bm ) = |B1m | i∈Bm f (Xi ).
|acc(Bm ) − conf(Bm )| 1 |Bm |
P
i∈Bm Yi
(1)
and the average confidence is
While smooth alternatives like Kernel Density Estimation [14] exist, all metrics share the same limitation: they require knowing the continuous probability f (Xi ) for every sample. Because APIs [ bin is impossible in practice. To prove that our proposed hide these predictions, computing an ECE single-query estimator overcomes this opacity, our empirical evaluation (Section 5) will simulate this restricted environment. This allows us to extract the hidden probabilities to compute an exact [ bin , which serves as the oracle against which our method and other baselines are benchmarked. ECE
4
The Single-Query Black-Box Estimator
[ blind , an estimator that computes calibration error using exactly one In this section, we introduce ECE API query per sample. Our approach relies on a simple idea: by injecting a targeted logit_bias to the binary output tokens, we can force a black-box API to evaluate threshold indicators of the form 1(f (Xi ) ≥ ti ). We then partition the dataset and use these indicators to construct a consistent estimator of the True Calibration Error (TCE) without ever extracting the underlying probabilities. Single-Query Threshold Evaluation. Let z1 and z0 denote the model’s raw logits corresponding to the exact positive and negative target tokens (e.g., the specific token IDs for “True” and “False”). By 2
restricting the decision to these two outcomes1 and ignoring the rest of the vocabulary, the softmax probability mathematically simplifies to a sigmoid over their difference: f (X) = σ(z1 − z0 ). Testing whether a sample’s confidence f (Xi ) exceeds a threshold ti ∈ (0, 1) is equivalent to bounding this logit gap: ti . (2) f (Xi ) ≥ ti ⇐⇒ z1 − z0 ≥ ln 1 − ti To evaluate this without white-box access, we define a threshold shift b = − ln(ti /(1 − ti )) and apply it to the positive token via the API’s logit_bias parameter. Additionally, to prevent the model from outputting synonymous but invalid tokens like “Yes” or “ Correct”, we apply a massive constant bias C (e.g., C = 50) to both the positive and negative targets to effectively suppress all other token probabilities to zero. By setting the API temperature to 0.0, the outputted token evaluates the inequality z1 +C +b ≥ z0 +C. Since the constant C cancels out, this condition simplifies to z1 − z0 ≥ −b. Consequently, if the API outputs the positive token, our indicator 1(f (Xi ) ≥ ti ) evaluates to 1; otherwise, it is 0. This recovers the exact threshold in a single query. Partition Scheme and Estimator. Standard binning evaluates whether f (Xi ) falls within a bin Bm = [tm , tm+1 ), which translates to the difference of two indicators: 1(f (Xi ) ≥ tm ) − 1(f (Xi ) ≥ tm+1 ). Because our budget allows only one query per sample, we cannot check both the upper and lower bounds for a single input. We resolve this by randomly partitioning the dataset S into M strictly disjoint subsets S1 , . . . , SM , N each containing exactly Nm = ⌊ M ⌋ independent samples. We replace the unknown continuous probability f (Xi ) with the constant bin midpoint cm = tm +t2m+1 , and estimate the local empirical gap, which directly approximates the m-th binning term |BNm | (acc(Bm ) − conf(Bm )) from Equation 1, by evaluating the upper and lower threshold indicators on adjacent subsets: X X 1 1 ˆ LC ∆ (cm − Yi )1(f (Xi ) ≥ tm ) − (cm − Yj )1(f (Xj ) ≥ tm+1 ) m = Nm Nm (Xi ,Yi )∈Sm
(Xj ,Yj )∈Sm+1
[ blind = The final estimator aggregates these local gaps over all M bins: ECE
PM
ˆ LC m=1 |∆m |
Theoretical Guarantees. We evaluate the consistency of our estimator against the True Calibration Error (TCE), defined as TCE = E[|E[Y |f (X)] − f (X)|]. Theorem 1 (Estimator Consistency). Let the true calibration function p → E[Y |f (X) = p] be L-Lipschitz. For a dataset of size N partitioned into M disjoint subsets, the Mean Squared Error [ blind with respect to the True Calibration Error (TCE) is strictly bounded by: (MSE) of ECE 3 2 M 1 [ E ECEblind − TCE ≤O + 2 (3) N M Consequently, if the number of bins scales with the dataset size as M ∝ N α for any scaling exponent 0 < α < 13 , the MSE strictly vanishes as N → ∞: 2 [ blind − TCE lim E ECE =0 (4) N →∞
[ blind is a provably consistent estimator of the TCE. The proof This theorem guarantees that ECE (detailed in Appendix B) establishes this by decomposing the MSE into a variance component and three distinct bias terms, all of which vanish under the required N α scaling law. Proof is inspired [ bin estimator. from [4] who performed similar bounding and decomposition for the standard ECE Optimal Bin Scaling By minimizing the theoretical upper bound of the MSE with respect to M , we show that the optimal bin count should scale as M ∝ N 1/5 . (See Appendix C for the derivation.) 1 In practice, this includes semantically equivalent token variants (e.g., " True", "TRUE"). For clarity of exposition, we derive the mechanism here for a single pair of tokens. The full derivation, proving that this thresholding mechanism holds perfectly across sets of multiple token variants, is provided in Appendix D.
3
5
Empirical Evaluation
[ blind on BoolQ [2] and a binarized MMLU [7], created by Experimental Setup. We evaluate ECE splitting each 4-way question into four independent Yes/No questions. To establish an exact whitebox ground truth from continuous probabilities, we simulate opaque APIs using four open-weight models: Qwen-2.5-7B-Instruct [15], Llama-3.1-8B-Instruct [5], Mistral-7B-Instruct-v0.3 [8], and Gemma-2-9B-IT [16]. Baselines. We compare our estimator against Verbalized Confidence (K = 1), Monte Carlo Sampling (K ∈ {1..8}, T = 1.0), and Iterative Logit Extraction [1] (K ∈ {1..8}) (Baselines implementation are detailed in Appendix D). To prevent vocabulary bleeding and ensure a fair comparison, all logit-based baselines apply the same C = 50 bias (Section 4). [ bin . We evaluate the Mean Absolute Error between each proxy estimator and the ground truth ECE
Figure 1: Cost-error Pareto frontier. Average MAE against the white-box oracle across four [ blind establishes models and two datasets. ECE the optimal trade-off, outperforming sampling and matching K = 4 Iterative Logit Extraction.
Figure 2: ECE Contribution Curve. Llama-3.18B evaluated on MMLU. It plots the marginal ˆ LC density-weighted contribution (∆ m ) of each threshold interval to the total error. Positive values indicate overconfidence.
[ blind outCost-Error Pareto Frontier. As shown in Figure 1, with a single-query budget, ECE performs Verbalized Confidence, reducing the average MAE from > 0.07 to near 0.01. Sampling is inefficient for calibration; even with K = 8 queries per sample, it plateaus at an MAE of 0.02. Iterative logit extraction [1] provides accurate continuous probabilities, and our 1-query estimator matches its performance at K = 5. While iterative logit extraction slightly surpasses our estimator [ blind defines a competitive Pareto at K ≥ 6, it requires at least 6x the query cost. Therefore, ECE frontier for cost-efficient calibration auditing. [ blind provides a visual Diagnostic Interpretability. Beyond producing a single ECE score, ECE diagnostic without requiring probabilities. Figure 2 shows the Blind Calibration Curve for Llama-3.18B on MMLU (extended curves are in Appendix F). Instead of plotting accuracy against confidence, ˆ LC this curve plots the signed local gap (∆ m ) per threshold interval. This visualizes the net densityweighted miscalibration: positive values indicate overconfidence, and negative values indicate underconfidence. For example, Figure 2 shows Llama-3.1-8B is underconfident at lower probabilities and overconfident in higher regions. In standard reliability diagrams, large visual gaps might represent [ blind small fractions of the dataset. In contrast, the absolute sum of our bars equals the final ECE score. This helps practitioners isolate which probability regions actually degrade model reliability.
6
Conclusion
Evaluating the calibration of opaque LLMs is challenging when API providers restrict continuous [ blind , an estimator that leverages the logit_bias probabilities. To solve this, we introduced ECE parameter to evaluate exact probability thresholds using strictly one query per sample. This eliminates 4
the prohibitive costs of multi-query extraction or sampling and the inaccuracies of verbalizing methods. We proved the estimator’s asymptotic consistency with the True Calibration Error and empirically demonstrated that it matches the accuracy of expensive baselines at a fraction of the cost. Ultimately, [ blind provides the research community with a mathematically rigorous, highly economical ECE framework for the independent auditing of commercial foundation models.
5
References [1] Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, and Florian Tramèr. Stealing part of a production language model. In Forty-first International Conference on Machine Learning, 2024. [2] Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 3694–3700, 2019. [3] Bradley Efron and Charles Stein. The jackknife estimate of variance. The Annals of Statistics, pages 586–596, 1981. [4] Futoshi Futami and Masahiro Fujisawa. Information-theoretic generalization analysis for expected calibration error. In Advances in Neural Information Processing Systems, volume 37, pages 84246–84297, 2024. [5] Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vítor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, 6
Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma. The llama 3 herd of models, 2024. [6] Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330. PMLR, 06–11 Aug 2017. 7
[7] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. [8] Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b, 2023. [9] Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac H Dodds, Nova DasSarma, Eli Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022. [10] Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664, 2023. [11] Stephanie Lin, Jacob Hilton, and Owain Evans. Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334, 2022. [12] Mahdi Pakdaman Naeini, Gregory F Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, pages 2901–2907, 2015. [13] OpenAI. Using logit bias to alter token probability with the OpenAI API, 2026. Accessed: 2026-08-21. [14] Teodora Popordanoska, Raphael Sayer, and Matthew B. Blaschko. A consistent and differentiable lp canonical calibration error estimator. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 7933–7946, 2022. [15] Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report, 2025. [16] Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman, Shantanu Thakoor, Jean-Bastien Grill, Behnam Neyshabur, Olivier Bachem, Alanna Walton, Aliaksei Severyn, Alicia Parrish, Aliya Ahmad, Allen Hutchison, Alvin Abdagic, Amanda Carl, Amy Shen, Andy Brock, Andy Coenen, Anthony Laforge, Antonia Paterson, Ben Bastian, Bilal Piot, Bo Wu, Brandon Royal, Charlie Chen, Chintu Kumar, Chris Perry, Chris Welty, Christopher A. Choquette-Choo, Danila Sinopalnikov, David Weinberger, Dimple Vijaykumar, Dominika Rogozińska, Dustin Herbison, Elisa Bandy, Emma Wang, Eric Noland, Erica Moreira, Evan Senter, Evgenii Eltyshev, Francesco Visin, Gabriel Rasskin, Gary Wei, Glenn Cameron, Gus Martins, Hadi Hashemi, Hanna Klimczak-Plucińska, Harleen Batra, Harsh Dhand, Ivan Nardini, Jacinda Mein, Jack Zhou, James Svensson, Jeff Stanway, Jetha Chan, Jin Peng Zhou, Joana Carrasqueira, Joana Iljazi, Jocelyn Becker, Joe Fernandez, Joost van Amersfoort, Josh Gordon, Josh Lipschultz, Josh Newlan, Ju yeong Ji, Kareem Mohamed, Kartikeya Badola, Kat Black, Katie Millican, Keelin McDonell, Kelvin Nguyen, Kiranbir Sodhia, Kish Greene, Lars Lowe Sjoesund, Lauren Usui, Laurent Sifre, Lena Heuermann, Leticia Lago, Lilly McNealus, Livio Baldini Soares, Logan Kilpatrick, Lucas Dixon, Luciano Martins, Machel Reid, Manvinder Singh, Mark Iverson, Martin Görner, Mat Velloso, Mateo Wirth, Matt Davidow, Matt Miller, Matthew Rahtz, Matthew Watson, Meg Risdal, Mehran Kazemi, Michael Moynihan, Ming Zhang, Minsuk Kahng, Minwoo Park, Mofi Rahman, Mohit Khatwani, Natalie Dao, Nenshad Bardoliwalla, Nesh Devanathan, Neta Dumai, Nilay Chauhan, Oscar Wahltinez, Pankil Botarda, Parker Barnes, Paul Barham, Paul Michel, Pengchong Jin, 8
Petko Georgiev, Phil Culliton, Pradeep Kuppala, Ramona Comanescu, Ramona Merhej, Reena Jana, Reza Ardeshir Rokni, Rishabh Agarwal, Ryan Mullins, Samaneh Saadat, Sara Mc Carthy, Sarah Cogan, Sarah Perrin, Sébastien M. R. Arnold, Sebastian Krause, Shengyang Dai, Shruti Garg, Shruti Sheth, Sue Ronstrom, Susan Chan, Timothy Jordan, Ting Yu, Tom Eccles, Tom Hennigan, Tomas Kocisky, Tulsee Doshi, Vihan Jain, Vikas Yadav, Vilobh Meshram, Vishal Dharmadhikari, Warren Barkley, Wei Wei, Wenming Ye, Woohyun Han, Woosuk Kwon, Xiang Xu, Zhe Shen, Zhitao Gong, Zichuan Wei, Victor Cotruta, Phoebe Kirk, Anand Rao, Minh Giang, Ludovic Peran, Tris Warkentin, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, D. Sculley, Jeanine Banks, Anca Dragan, Slav Petrov, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Sebastian Borgeaud, Noah Fiedel, Armand Joulin, Kathleen Kenealy, Robert Dadashi, and Alek Andreev. Gemma 2: Improving open language models at a practical size, 2024. [17] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022. [18] Miao Xiong, Zhiyuan Hu, Xinyang Lu, Ruibo Li, Jie Fu, Zhengxiao Liu, et al. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. The Twelfth International Conference on Learning Representations, 2023. [19] Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, Tongshuang Wu, and Jianshu Chen. Fact-and-reflection (FaR) improves confidence calibration of large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 8702–8718, Bangkok, Thailand, August 2024. Association for Computational Linguistics.
9
A
Limitations and Future Work
[ blind provides a provably consistent and highly query-efficient framework for black-box While ECE calibration auditing, our methodology is subject to several theoretical and practical limitations: • Dependence on API Infrastructure: The primary limitation of our methodology is its reliance on commercial API providers exposing a logit_bias parameter. As [1] demonstrate, this parameter enables model-stealing attacks via iterative log-probability extraction. Consequently, providers are restricting its use. OpenAI supports it for standard models but disables it for newer reasoning architectures (e.g., the o1 series). Anthropic recently removed this parameter from their API, and Google’s Gemini API ignores it. If the industry fully deprecates this feature, exact black-box calibration auditing via thresholding will become impossible. • Restriction to Binary Classification Tasks: Our technique relies on a binary choice where the probability simplifies to the difference between two token logits. In multi-class tasks involving several tokens (e.g., “A”, “B”, “C”, “D”), the probability of any single token depends on all the others. Consequently, evaluating a probability threshold for even the top label is not straightforward using a single logit_bias value. Furthermore, evaluating multi-class calibration requires choosing among several ECE definitions (such as top-label or marginal ECE). Developing a methodology to map these multi-class metrics to single API queries remains an open problem. • Optimal Allocation of Fixed Budgets: Our estimator is designed estimation calibration under a budget of exactly one query per sample. If an auditor has a larger budget (e.g., K = 3 or K = 5), our formulation does not dictate how to optimally allocate these additional queries. Future work must determine whether a larger budget is better spent evaluating multiple thresholds per sample or querying independent subsets. • Token Probability vs. Confidence: Our method measures the calibration of the model’s next-token predictive distribution over specific target words. However, as is common in zero-shot LLM evaluation, the raw probability mass a model assigns to the token “True” may not fully capture its confidence in the underlying factual claim [9, 10]. • Hidden Probability Mass and Clamping Effects: We apply a large constant bias (C = 50) to force the model to select between specified target tokens. If the model favors an unpredicted but valid synonym (e.g., “Correct” or “Yes”), that probability mass is suppressed. Note that this clamping effect applies to all evaluated logit-based estimators, including the white-box oracle. • Simulated API Environments: To compute the exact white-box ground truth required for our empirical evaluation, we simulated black-box constraints using open-weights models rather than querying live commercial endpoints. While mathematically equivalent to a live API exposing a logit_bias parameter, real-world production APIs often include undocumented prompt formatting, safety filters, or dynamic model routing that could potentially interfere with precise logit manipulations. • The “Oracle” is not a True Oracle: In our empirical evaluation, we treat the whitebox continuous binned ECE (ECEoracle ) as the ground truth. However, this oracle is an empirical estimate of the True Calibration Error (TCE) computed over a finite dataset, meaning it remains subject to standard finite-sample and discrete binning biases. • Requirement for Large Datasets: To optimally balance statistical subset noise against discretization bias, our theoretical framework requires the number of bins to scale as M ∝ N 1/5 . Consequently, the method depends on a sufficiently large dataset to achieve a reasonable resolution. For example, evaluating the standard BoolQ validation split (N = 3, 270 examples) strictly limits the optimal bin count to M = 5. This inevitably restricts the granularity of the calibration evaluation, making the estimator impractical for small benchmarks.
10
B
Detailed Proofs for Estimator Bounds and Consistency
In this section, we provide the complete mathematical proofs for the bounds on the bias and variance [ blind , establishing its consistency. of the single-query blind estimator ECE Let S = (Xi , Yi )N i=1 represent a dataset of N independent observations. The continuous probability space [0, 1] is partitioned into M equal-width bins Bm = [tm , tm+1 ), with respective constant midpoints cm = tm +t2m+1 . The dataset is randomly fractured into M strictly disjoint subsets N ⌋ or Nm + 1 samples. S1 , . . . , SM , each containing exactly Nm = ⌊ M For a specific sample Z = (X, Y ), threshold t, and midpoint cm , we define the threshold observation function: W (Z, t, cm ) = (cm − Y )1(f (X) ≥ t) (5) The local empirical gap is defined as: 1 X 1 ˆ LC ∆ W (Zi , tm , cm ) − m = Nm Nm i∈Sm
B.1
X
W (Zj , tm+1 , cm )
(6)
j∈Sm+1
Expected Value and Variance of the Local Empirical Gap
T rap ˆ LC ˆ LC Lemma 1 (Expected Value of ∆ = E[(cm − Y )1(f (X) ∈ Bm )]. Then E[∆ m ). Let ∆m m ]= ∆Tmrap .
Proof. By the linearity of expectation and i.i.d assumption: ˆ LC E[∆ m ] = E[(cm − Y )1(f (X) ≥ tm )] − E[(cm − Y )1(f (X) ≥ tm+1 )] Factoring out (cm − Y ), the difference of the two indicator functions evaluates to 1 strictly when f (X) falls between tm and tm+1 . Thus: T rap ˆ LC E[∆ m ] = E[(cm −Y )(1(f (X) ≥ tm )−1(f (X) ≥ tm+1 ))] = E[(cm −Y )1(f (X) ∈ Bm )] = ∆m
ˆ LC Lemma 2 (Variance of ∆ m ). The variance of the local empirical gap is bounded such that 2M LC ˆm ) ≤ V ar(∆ . N
Proof. Because Sm and Sm+1 are completely disjoint, their sample means are strictly independent. The variance of their difference is the sum of their variances: V ar(W (Z, tm , cm )) V ar(W (Z, tm+1 , cm )) ˆ LC V ar(∆ + m )= Nm Nm Given that cm ∈ [0, 1], Y ∈ {0, 1}, and the indicator is in {0, 1}, the random variable W is strictly bounded in the interval [−1, 1]. The maximum possible variance for a variable bounded in [−1, 1] is exactly 1. Substituting Nm = N/M , we obtain: ˆ LC V ar(∆ m )≤
1 2M 1 + = N/M N/M N
Lemma 3 (TCE Partition). The True Calibration Error can be decomposed over the M disjoint bins as: M X TCE = E[|cf (f (X)) − f (X)| | f (X) ∈ Bm ]P(f (X) ∈ Bm ) m=1
Proof. By definition, TCE = E[|E[Y |f (X)] − f (X)|]. Substituting the ideal calibration function cf (f (X)) = E[Y |f (X)], we can rewrite this as TCE = E[|cf (f (X))−f (X)|]. Because the disjoint bins {Bm }M m=1 form a complete partition of the probability space [0, 1], we apply the Law of Total Expectation to condition on the event f (X) ∈ Bm . Summing these conditional expectations yields the stated decomposition. 11
B.2
The 3-Term Bias Decomposition
We evaluate the macroscopic deviation of the estimator from the True Calibration Error (TCE). To do so, we first introduce the continuous ideal calibration function cf (p) = E[Y |f (X) = p] and the theoretical continuous target gap ∆∗m = E[(f (X) − Y )1(f (X) ∈ Bm )]. Expanding the expected value of our estimator by adding and subtracting the theoretical anchors |∆Tmrap | and |∆∗m |, and using Lemma 3, we relate it directly to the True Calibration Error: [ blind ] = E[ECE
M X
ˆ LC E[|∆ m |]
m=1
=
M X
E[|cf (f (X)) − f (X)| | f (X) ∈ Bm ]P(f (X) ∈ Bm )
m=1
{z
| −
}
:=TCE M X
(E[1(f (X) ∈ Bm )|cf (f (X)) − f (X)|] − |∆∗m |)
m=1
| +
M X
{z
:=Bbin
}
|∆Tmrap | − |∆∗m |
m=1
| +
{z
}
:=Btrap
M X T rap ˆ LC E[|∆ | m |] − |∆m
(7)
m=1
|
{z
}
:=Bstat
This rigorously yields the 3-term decomposition of the estimator’s expected macroscopic deviation: [ blind ) = E[ECE [ blind ] − TCE = Bstat + Btrap − Bbin Bias(ECE
(8)
Part 1: Bounding the Statistical Bias (Bstat ) With Lemma 1 we have ˆ LC |] − |∆T rap | = E[|∆ ˆ LC |] − |E[∆ ˆ LC ]| E[|∆ (9) m m m m p We also have, for any random variable Z, E[|Z|] − |E[Z]| ≤ V ar(Z). Applying this directly to Equation 9: q T rap ˆ LC ˆ LC V ar(∆ E[|∆ |] − |∆ | ≤ m m m ) Substituting the upper bound from Lemma 2 (V ar ≤ 2M/N ) and summing across all M bins yields the statistical bias limit: r M r X √ M 3/2 2M 2M Bstat ≤ =M = 2 1/2 (10) N N N m=1 Part 2: Bounding the Trapezoidal Bias (Btrap ) This bias measures the geometric error introduced by replacing the continuous prediction f (X) with the constant bin midpoint cm . Using the Reverse Triangle Inequality (|A| − |B| ≤ |A − B|) and the property that |E[Z]| ≤ E[|Z|]: |∆Tmrap | − |∆∗m | ≤ |∆Tmrap − ∆∗m | = |E[(cm − Y )1Bm ] − E[(f (X) − Y )1Bm ]| ≤ E[1(f (X) ∈ Bm )|cm − f (X)|] Because the prediction f (X) is strictly constrained to the bin Bm (total width 1/M ), and cm is its exact geometric midpoint, the absolute distance between them can never exceed half the bin’s width: 12
|cm − f (X)| ≤ 1/(2M ). Substituting this absolute geometric limit: M X
X M 1 1 1 Btrap ≤ 1(f (X) ∈ Bm ) = 1(f (X) ∈ Bm ) = E 2M 2M 2M m=1 m=1
(11)
Part 3: Bounding the Binning Bias (Bbin ) The binning bias Bbin represents the error caused by dividing the continuous probability space into M discrete bins. This error depends only on the bin width and the true calibration function, making it identical to the discretization bias in standard ECE. Assuming the true calibration function is L-Lipschitz, the variation within any bin of width 1/M is strictly limited. Following the theoretical analysis of binned estimators by [4] (Theorem 3.) , this bias is bounded by: 1+L Bbin ≤ (12) 2M B.3
Variance Bound via Efron-Stein Inequality
[ blind Theorem 2 (Variance Bound of the Estimator). The variance of the empirical estimator ECE evaluated on N samples split into M disjoint subsets is strictly bounded by 8M 2 /N . PM ˆ LC Proof. Let Φ(S) = m=1 |∆ m | be our estimator acting on the dataset S. To apply the Efron-Stein inequality [3], we evaluate the maximum absolute perturbation |Φ(S) − Φ(S (k) )| when a single sample Zk ∈ S is replaced by an independent copy Zk′ . The modified sample Zk belongs to exactly one disjoint subset, Sm′ . By definition, this subset is ˆ LC′ and ∆ ˆ LC′ . Therefore, replacing Zk with Z ′ utilized in exactly two local empirical gaps: ∆ m m −1 k leaves the other M − 2 gaps perfectly unchanged. Because the observation function W (Z, t, cm ) ∈ [−1, 1], the maximum absolute difference caused by replacing one sample is bounded by |W (Zk ) − W (Zk′ )| ≤ 2. Consequently, the perturbation on ˆ LC′ is bounded by 2/Nm′ = 2M/N . The same bound applies to ∆ ˆ LC′ . the raw gap ∆ m m −1 Using the reverse triangle inequality (||a| − |b|| ≤ |a − b|), the change in the absolute gaps cannot exceed the change in the raw gaps. Summing these perturbations, the maximum total change to the estimator is: 2M 2M 4M |Φ(S) − Φ(S (k) )| ≤ + = N N N The Efron-Stein inequality limits the variance of Φ(S) by half the expected sum of squared perturbations: N i 1X h V ar(Φ(S)) ≤ E (Φ(S) − Φ(S (k) ))2 2 k=1
Substituting our deterministic worst-case bound: N
1X V ar(Φ(S)) ≤ 2
k=1
4M N
2
1 = 2
16M 2 N· N2
=
8M 2 N
This concludes the proof. B.4
Final Proof of Theorem 1 (Estimator Consistency)
We now combine the bounds for the variance and the three bias terms to prove Theorem 1. The Mean Squared Error (MSE) of any estimator is the sum of its squared bias and its variance: [ blind ) = Bias(ECE [ blind )2 + V ar(ECE [ blind ) MSE(ECE 13
(13)
From our 3-term decomposition, the total absolute bias is bounded by the sum of the individual bounds: 3/2 M 1 |Bias| ≤ Bstat + Btrap + Bbin ≤ O + (14) M N 1/2 Squaring this total bias gives: 2
Bias ≤ O
M3 1 + 2 N M
(15) 2
From the Efron-Stein inequality, we established that the variance is strictly bounded by 8M N , which M2 M3 M2 is O( N ). Because N dominates N , the variance term is absorbed into the squared bias bound. This yields the final MSE bound: 3 1 M [ (16) MSE(ECEblind ) ≤ O + 2 N M To ensure the estimator is consistent, the MSE must vanish as the dataset size N goes to infinity. If we scale the number of bins M as M ∝ N α , we can substitute this into the MSE bound: MSE ≤ O N 3α−1 + N −2α (17) For both terms to approach zero as N → ∞, their exponents must be strictly negative. This requires: 1. 3α − 1 < 0 =⇒ α < 31 2. −2α < 0 =⇒ α > 0 Therefore, for any scaling exponent 0 < α < 13 , the MSE strictly vanishes: 2 [ blind − TCE lim E ECE =0 N →∞
(18)
This concludes the proof. Discussion on the Lipschitz Assumption. Theorem1 assumes the true calibration function cf (p) = E[Y |f (X) = p] is L-Lipschitz. While the underlying neural network f (X) mapping inputs to probabilities is typically Lipschitz continuous, this does not guarantee that cf (p) is Lipschitz. The calibration curve depends on the conditional density of the dataset, meaning sharp transitions in the data distribution could create discontinuities in cf (p). Nevertheless, assuming a Lipschitz continuous true calibration function is a standard premise in theoretical calibration literature [4, 14] to derive finite-sample bounds.
C
Derivation of the Optimal Bin Scaling
[ blind estimator is From Theorem 1, the upper bound on the Mean Squared Error (MSE) of the ECE given by: 3 M 1 MSE ≤ O + 2 (19) N M To determine the optimal scaling of the number of bins M with respect to the dataset size N that minimizes this macroscopic error, we differentiate the upper bound with respect to M and set the derivative to zero: 3 ∂ M 1 3M 2 2 + 2 = − 3 =0 (20) ∂M N M N M Solving for M , we obtain: 3M 2 2 2N = 3 =⇒ M 5 = =⇒ M ∝ N 1/5 (21) N M 3 Thus, to asymptotically minimize the estimation error while balancing statistical subset noise against discretization bias, the optimal number of bins should scale as N 1/5 . 14
Bound Tightness. Because our derivation establishes an upper bound on the estimation error, we lack a theoretical guarantee on its tightness. If the bound is loose, the derived optimal bin scaling (M ∝ N 1/5 ) may be suboptimal in practice, and a different scaling exponent might yield lower empirical errors. However, this lack of tightness does not impact the asymptotic consistency of the estimator; any scaling exponent 0 < α < 1/3 guarantees that the MSE strictly vanishes as N → ∞.
15
D
Methodology and Baseline Implementations
D.1
Standard ECE Formulation and the Optimal Number of Bins
For a binary classification task over a dataset of N samples, the standard empirical binned Expected Calibration Error (ECE) is computed by partitioning the continuous probability space [0, 1] into M disjoint bins, denoted B1 , . . . , BM . Let Nm = |Bm | be the number of samples whose predicted positive class probability f (Xi ) falls into the m-th bin. The estimator is defined as: [ bin = ECE
M X Nm m=1
N
|acc(Bm ) − conf(Bm )|
(22)
P P where acc(Bm ) = N1m i∈Bm Yi is the empirical accuracy, and conf(Bm ) = N1m i∈Bm f (Xi ) is the average model confidence within the bin. The selection of the bin count M introduces a fundamental bias-variance trade-off: a small M obscures local calibration errors (high binning bias), while a large M leaves bins sparsely populated (high statistical variance). Recent information-theoretic analyses [4] have formalized this estimation bias, deriving the optimal number of bins to minimize the error of the binned estimator. In our [ bin oracle exactly as derived by [4]. This results in experiments, we scale M ∝ N 1/3 for the ECE M = 15 for BoolQ (N = 3, 270) and M = 38 for MMLU (N = 56, 168). D.2
Empirical Implementations and the Restricted Target Space
Prompting Strategy for Binary Tasks. To evaluate the models, we format all dataset queries as strict binary classification tasks. For example, a sample from the BoolQ dataset is formulated as: “Passage: All biomass goes through at least some of these steps: it needs to be grown, collected... [. . . ] Question: does ethanol take more energy to make than it produces? Answer True or False.” Similar binary constraints are applied to other datasets like MMLU (e.g., “Reply only with Yes or No.”). The Restricted Target Space. When evaluating zero-shot LLMs via APIs, standard evaluation risks “vocabulary bleeding,” where the model distributes probability mass across synonymous tokens (e.g., “Yes”, “ Correct”) instead of the strict binary targets (e.g., “True” vs. “False”). To ensure our baselines fail strictly due to query constraints rather than prompt formatting noise, we clamp the decision space for all logit-based estimators (Oracle, Ours, Sampling, and Iterative Extraction). To account for tokenizer fragmentation (e.g., "True", " true", "TRUE"), we define V + as the set of all valid token variations for the positive answer and V − for the negative answer. We apply a large positive constant bias (C = 50) to all tokens in V + ∪ V − , effectively suppressing all irrelevant vocabulary to zero. 1. The Oracle (White-Box Ground Truth). The true continuous confidence f (Xi ) is computed directly from the model’s native hidden logits. We extract the exact pre-softmax logit zv for every target token in V + ∪ V − . The continuous probability is calculated using a restricted softmax over these sets: P v∈V + exp(zv ) P P f (Xi ) = exp(z + v) + v∈V v∈V − exp(zv ) [ bin equation. This exact probability is passed to the standard ECE [ blind (Ours, 1 Query). As defined in Section 4, the dataset is partitioned into M disjoint 2. ECE subsets. For a given sample Xi ∈ Sm evaluated against threshold tm , we shift the model’s decision boundary by computing bm = − ln(tm /(1 − tm )). As described above, we apply the large constant bias C = 50 to all tokens in V + ∪ V − to prevent vocabulary bleeding. To evaluate the threshold, we apply the additional shift bm exclusively to the 16
positive tokens in V + . The modified logits z̃v sent to the API are therefore: + zv + C + bm if v ∈ V z̃v = zv + C if v ∈ V − zv otherwise Because C is sufficiently large, the probability of generating any token outside of V + ∪ V − becomes negligible. The modified probability of outputting a positive token is: P v∈V + exp(zv + C + bm ) P P̃ (pos | Xi ) = P v∈V + exp(zv + C + bm ) + v∈V − exp(zv + C) Notice that exp(C) factors out of both the numerator and the denominator, perfectly preserving the relative probabilities between the sets V + and V − while restricting the vocabulary. Factoring out exp(C) leaves: P ebm v∈V + exp(zv ) P P̃ (pos | Xi ) = bm P e v∈V + exp(zv ) + v∈V − exp(zv ) When querying the API using greedy decoding (T = 0), the model outputs a token from V + if and only if P̃ (pos | Xi ) ≥ 0.5. Substituting our shift bm and the original continuous probability p = f (Xi ), this condition simplifies exactly to p ≥ tm . Thus, observing any positive token evaluates the indicator 1(f (Xi ) ≥ tm ) = 1 in a single query. 3. Verbalized Confidence (1 Query). The model is prompted to explicitly verbalize its certainty (e.g., “Answer True or False, and state your confidence as a percentage between 50 and 100”). We parse the output text to extract the stated probability p̂i . Because this relies on free-text generation, the C = 50 target bias cannot be applied. 4. Sampling / Self-Consistency (K Queries). We query the API K times per sample at temperature T = 1.0. To strictly evaluate the mathematical variance of sampling (rather than formatting failures), we apply the restricted space bias C = 50 to the sets V + and V − . The estimated probability p̂i is the empirical frequency of the positive tokens across the K generations. 5. Iterative Logit Extraction [1] (K Queries). Given a strict budget of K queries per sample, we perform a binary search over the threshold shift b ∈ [−B, +B] (with maximum logit gap bounds B = 15) to find the decision boundary b∗ where the model’s argmax output flips from a positive token in V + to a negative token in V − . We evaluate each step at T = 0 applying the shift b to the set V + exactly as described for our method. At the boundary, the shifted probability mass is perfectly balanced, allowing us to estimate the continuous probability as p̂i = σ(−b∗ ).
E
Extended Results and Breakdown
Table 1 details the Expected Calibration Error (ECE) estimates for every method across all evaluated models and datasets. To facilitate readability, all ECE values are rounded to three decimal places. Detailed Observations. The breakdown reveals several key insights regarding the behavior of the estimators across different architectures and tasks: • Exceptional Accuracy on Specific Pairs: Despite operating under a strict 1-query budget, [ blind is remarkably precise on specific configurations. It nearly perfectly matches ECE the Oracle for Llama-3.1-8B on both datasets (e.g., 0.085 vs 0.084 on BoolQ). It also shows outstanding accuracy on the MMLU dataset for both Mistral-7B (0.340 vs 0.339) and Gemma-2-9B (0.314 vs 0.313). • Failure of Generative Proxies: Verbalized confidence is highly erratic and often entirely disconnected from the model’s true calibration (e.g., overestimating Llama’s MMLU error by nearly 3×). • Inefficiency of Sampling: Self-consistency sampling slowly converges, but it systematically overestimates the calibration error. Even at K = 8, it remains less accurate than our 1-query estimator on almost every model-dataset pair. 17
• Convergence of Iterative Extraction: The Carlini et al. binary search approach is highly inaccurate at low budgets (K ≤ 3) because the search space is unresolved. It typically [ blind achieves in a single query, requires 4 to 5 queries to match the precision that ECE before finally converging to the Oracle at K = 8. Qwen2.5-7B
Llama-3.1-8B
Boolq
Mmlu
Boolq
Mmlu
Boolq
Mmlu
Boolq
Mmlu
Oracle ECEblind (Ours, K = 1)
0.140 0.111
0.156 0.133
0.084 0.085
0.070 0.074
0.095 0.074
0.339 0.340
0.245 0.223
0.313 0.314
Verbalized (K = 1)
0.165
0.064
0.118
0.198
0.154
0.279
0.086
0.275
Sampling (K = 1) Sampling (K = 2) Sampling (K = 3) Sampling (K = 4) Sampling (K = 5) Sampling (K = 6) Sampling (K = 7) Sampling (K = 8)
0.160 0.150 0.145 0.143 0.145 0.141 0.143 0.143
0.237 0.204 0.185 0.175 0.173 0.170 0.167 0.166
0.332 0.179 0.117 0.101 0.087 0.079 0.069 0.067
0.356 0.269 0.200 0.164 0.158 0.146 0.136 0.129
0.201 0.152 0.133 0.119 0.113 0.106 0.104 0.105
0.459 0.403 0.382 0.370 0.361 0.358 0.355 0.353
0.317 0.257 0.246 0.245 0.244 0.245 0.243 0.243
0.435 0.357 0.336 0.329 0.323 0.321 0.318 0.317
Carlini et al. (K = 1) Carlini et al. (K = 2) Carlini et al. (K = 3) Carlini et al. (K = 4) Carlini et al. (K = 5) Carlini et al. (K = 6) Carlini et al. (K = 7) Carlini et al. (K = 8)
0.156 0.150 0.140 0.138 0.139 0.140 0.140 0.140
0.220 0.205 0.164 0.147 0.156 0.155 0.156 0.156
0.194 0.172 0.062 0.063 0.073 0.085 0.085 0.084
0.272 0.249 0.141 0.051 0.080 0.072 0.071 0.071
0.170 0.148 0.101 0.088 0.094 0.094 0.094 0.094
0.449 0.427 0.338 0.335 0.336 0.338 0.338 0.339
0.325 0.303 0.222 0.237 0.242 0.243 0.244 0.245
0.416 0.393 0.332 0.303 0.308 0.311 0.312 0.313
Estimator
Mistral-7B
Gemma-2-9B
Table 1: Detailed Calibration Evaluation (ECE) across Models and Datasets. All ECE values are derived using 5 uniform bins and are rounded to three decimal places. For each column, the estimator closest to the Oracle ECE is in bold, and the second closest is underlined.
18
F
Extended Blind Calibration Curve and Interpretation
In standard white-box ECE evaluation, researchers typically use reliability diagrams. These diagrams plot the model’s predicted confidence against its empirical accuracy across discrete bins. However, standard reliability diagrams can be visually misleading: a massive gap between accuracy and confidence in a specific bin only impacts the final metric if a large percentage of the dataset actually falls into that bin. To calculate the final ECE, one must manually reweight the visible gap in each bin by its underlying (and often visually hidden) sample mass. Because we operate in a strict black-box setting (K = 1), we cannot extract the exact probability pi of each sample, making it impossible to construct traditional confidence bins or calculate their exact [ blind estimator naturally bypasses this requirement by sweeping a decision mass. Instead, the ECE ˆ LC . boundary across the dataset and recording the local empirical gap, ∆ m ˆ LC Plotting ∆ m across the confidence regions yields the Blind Calibration Curve. To demonstrate its utility, Figures 3 and 4 provide a side-by-side comparison of standard reliability diagrams (computed via the white-box oracle) and our Blind Calibration Curves for all models on the BoolQ and MMLU datasets, respectively. Comparing these side-by-side highlights two key advantages of our visualization: 1. It shows Net Error (Density-Weighted): While a standard reliability diagram might show a severe accuracy drop in an edge bin, our contribution curve correctly scales this by the bin’s mass. The height of each bar in our curve represents the actual, final penalty that the specific region contributes to the total ECE. The absolute sum of these bars perfectly equals [ blind score. the final ECE 2. It shows Error Direction (Signed): • Positive Values (Overconfidence): The model crosses the threshold and predicts the positive class more frequently than the ground-truth dataset labels justify. • Negative Values (Underconfidence): The model hesitates to cross the threshold, predicting the positive class less frequently than it actually occurs in the dataset.
19
Figure 3: Calibration Comparison on BoolQ. For each of the four models, the left plot shows the standard white-box reliability diagram (Accuracy vs. Confidence), and the right plot shows our ˆ LC 1-query Blind Calibration Curve (∆ m ). Notice how visually large gaps in sparse bins on the standard curve translate to negligible penalties in our density-weighted contribution curve.
20
Figure 4: Calibration Comparison on MMLU. For each of the four models, the left plot shows the standard white-box reliability diagram, and the right plot shows our 1-query Blind Calibration Curve. Positive values indicate overconfidence, while negative values indicate underconfidence.
21