How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method Konstantin Grotov* Tel Aviv University Tel Aviv, Israel [email protected]
arXiv:2609.05274v1 [cs.LG] 4 Sep 2026
Abstract LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access to logits, weights, activations, or repeated sampling. Inverting speculative decoding, a small openweight draft model scores the agent’s alreadygenerated trajectory in a single forward pass. From these speculative cross-likelihoods we extract phase-aware features by separating the reasoning and action spans, and calibrate them against a verifiable objective. SU produces a failure-likelihood score that any downstream policy, such as routing, human intervention, or extra test-time compute, can consume directly. To show the signal is actionable, we instantiate one such policy, a pre-execution veto gate, on software engineering agents Qwen3Coder-480B and closed-source Claude 3.5 Sonnet, cutting execution error rate by 6–8 percentage points and token cost by 14–19% in deployment, transferring to out-of-distribution benchmarks without retraining, and generalizing across agent models.
1
Introduction
LLM agents are increasingly deployed in production software engineering workflows, where they autonomously draft patches, run tests, and query databases (Jimenez et al., 2024; Badertdinov et al., 2026; Yang et al., 2026). Yet behind strong benchmark scores lies an expensive failure mode: agents are often confidently wrong (Kaddour et al., 2026). They propose an action with no signal that it could fail, and the failure appears only when an environment rejects it—after the code has run, the context has been filled up, and the retry has been paid for. At production scale, these silent failures inflate API * Corresponding author.
Valentin Malykh IITU Almaty, Kazakhstan [email protected]
cost and latency, since a rejected action is appended to the context and the agent pays for the retry — frequently looping on the same problem until the budget is exhausted (Kaddour et al., 2026). Two properties of realistic deployments constrain the space of admissible methods. First, the most capable deployed agents are closed APIserved models exposing no logits, weights, or activations, so any method that needs internal signals is something that we cannot use for models. Second, sampling-based uncertainty estimators require many generations per step, which is prohibitive at agentic horizons. Taken together, these requirements rule out most of the uncertaintyquantification toolbox: single-turn calibration ignores features specific to the trajectories (Guo et al., 2017; Kuhn et al., 2023), white-box probing needs internal access, whereas sampling is too costly. The closest prior work, Holistic Trajectory Calibration (HTC) (Zhang et al., 2026), fits a calibrator on an agent’s full log-probability sequence. However, it still assumes white-box access and treats the trajectory as a homogeneous token stream. Both assumptions break in the real deployments we care about. We address the access constraint with an idea inverted from speculative decoding (Leviathan et al., 2023; Chen et al., 2023): rather than having a small draft model propose tokens to accelerate a large one, we have it score the large agent’s already-generated tokens via teacher-forced crosslikelihood. The agent doesn’t share its internal state, the draft reads only its output text and recovers a rich uncertainty signal in a single forward pass. We empirically observed that agent generation has a phase structure: a high-entropy reasoning phase is followed by a sharply lower-entropy action phase pinned down by syntax and tool-call structure. Averaging over the whole trajectory, as HTC does, dissolves this failure-relevant signal into the noise of the reasoning phase, while reading the two
phases apart surfaces it. Combining these considerations, we shape the method of Speculative Uncertainty (SU): a small draft scores the agent’s trajectory, phase-aware features are extracted separately from reasoning and action spans, and a linear calibrator maps them to a probability that the next action will follow the verifiable objective. We study code execution in software engineering agents, but the approach applies to any agent producing reasoning-then-action trajectories under a verifiable objective. This study shows that the uncertainty of a blackbox agent can be recovered from its trajectory alone, before the environment is ever called, and that the recovered signal is reliable enough to act positively in deployment. We insert a lightweight gate between code generation and execution that predicts whether the next action will be executed without exceptions. When failure is likely, this triggers a cheap replan instead of an expensive execute-fail-retry cycle. On two agentic coding benchmarks (Jimenez et al., 2024; Huang et al., 2024) the gate cuts per-call execution error rate by 6-8 percentage points and average token cost by 14–19%, with gains transferring to an out-ofdistribution benchmarks. Contributions. 1. A black-box uncertainty estimation method. By inverting speculative decoding, a small draft scores the agent’s generated tokens and recovers a usable failure signal without whitebox access. This is exactly the regime where sampling is too costly and internals are hidden, making it usable with the closed-source APIs that dominate production. 2. Deployment guidance. As one downstream policy, a pre-execution veto gate cuts per-call execution error rate by 6–8 points and token cost by 14–19% on SWE-Bench Verified and DA-Code, with a tunable false-veto rate and zero-shot out-of-distribution transfer. We further find that a distilled 4B draft is capable enough, the training objective barely matters, and a draft trained on one agent transfers to others well above the untrained baseline.
2
Background and Related Work
Uncertainty quantification for LLMs and agents. Single-turn LLM calibration is well-studied (Guo
et al., 2017; Geng et al., 2024; Kuhn et al., 2023; Lin et al., 2023), spanning post-hoc scaling, semantic-consistency sampling, and verbalized self-reflection. The agentic setting adds a layer of complexity: uncertainty is no longer a property of a single output but something that accumulates across interdependent steps (Zhang et al., 2026; Duan et al., 2025; Zhao et al., 2024). HTC (Zhang et al., 2026) confronts this directly, extracting hand-crafted trajectory features from an agent’s full log-probability sequence and fitting a logistic-regression calibrator. We adopt HTC’s trajectory-feature view and depart from it in two respects that prove decisive for deployment: we separate reasoning from action spans, and we obtain the signal from a separate draft model rather than the agent’s model own logits. Hallucinations and uncertainty in tool-using agents. A line of recent work probes confidence in tool-using agents. Healy et al. (2026); Subramani et al. (2025) train probes on a target model’s internal states to anticipate erroneous tool calls. Agents that reason before acting hallucinate tool calls more often (Yin et al., 2025). While those methods are effective, but operated on white-box access. Xuan et al. (2026); Kaddour et al. (2026) documents a systematic gap between expressed and actual confidence in tool-use agents. Speculative decoding and draft models. Speculative decoding (Leviathan et al., 2023; Chen et al., 2023) pairs a small draft model with a large target, accepting drafted tokens with alignment-specific probability. Its purpose is mainly to accelerate inference (Xia et al., 2024; Leviathan et al., 2023). We retain the machinery and invert the objective: the small model does not generate, it evaluates, scoring the target’s committed tokens after the fact.
3
Method
3.1
Problem Setup
Let p be the large, possibly black-box agent model (further simply the agent), and q a small openweight draft model. Given a task with initial context x, the agent produces a trajectory of alternating reasoning and action spans: T = r1 , a1 , r2 , a2 , . . . , rN , aN , where ri ∈ Σ∗ is the i-th reasoning span and ai ∈ Σ∗ is the i-th action span. Let y = (y1 , . . . , yT ) be the in-order concatenation of all tokens; each
Black-box Agent Trajectory Reason
Action
Reason
Small Draft Model
Replan
…
Action
Phase-Aware Features
likely fail
Next Action
Calibrator
confident
Veto Gate
Reason
Execute
Figure 1: The Speculative Uncertainty pipeline. A black-box agent generates a trajectory of alternating reasoning and action spans, exposing only its output tokens. A small open-weight draft model scores these tokens in a single forward pass, producing per-token uncertainty signals. Features are extracted separately from the reasoning and action spans, and a lightweight calibrator maps them to a calibrated success probability. Before execution, a threshold gate forwards confident actions or triggers a cheaper replan.
carries a component label τt ∈ {reason, action}, read off the harness’s message structure: in OpenHands (Wang et al., 2025) the assistant’s free-form text forms the reasoning span and the subsequent structured tool call (the invoked function and its serialized arguments) forms the action span. Each agent is paired with a verifiable objective: a binary function Ω(a, x) ∈ {0, 1} that an external oracle evaluates on action a given context x. SU operates at step granularity: for each step i it predicts the oracle verdict Ω(ai , x) on that step’s action before ai executes. Rather than conditioning on the full trajectory, which grows long and increasingly noisy over many steps, in this study we condition on a fixed lookback window of the k = 3 most recent steps. Let Wi = ri−k+1 , ai−k+1 , . . . , ri−1 , ai−1 , ri , ai
be the window ending at the step under evaluation (truncated at the start of the trajectory when i < k). We seek a calibration function F that maps the window to a calibrated probability that the step’s action will pass its check, before Ω is evaluated: ŷ = F(Wi ) ∈ [0, 1], E Ω(ai , x) | F (Wi ) = c ≈ c. In our deployment Ω is code-execution success, but the framework is agnostic to the objective. 3.2
Speculative Cross-Likelihoods
For each token yt we run the draft model q in teacher-forcing mode: conditioned on the prefix (x, y<t ), we read off q’s next-token distribution and evaluate it at the token p actually produced. A
single forward pass over the trajectory yields three token-level signals at no cost to the agent pipeline: st = − log q(yt | x, y<t ), gt = log q(ŷt | x, y<t ) − log q(yt | x, y<t ), X Ht = − q(v | x, y<t ) log q(v | x, y<t ), v∈V
where ŷt = arg maxv q(v | x, y<t ). We call st the speculative surprisal, gt the speculative gap, and Ht the speculative entropy. All three depend only on the observed token stream; p’s logits are never needed at inference. Why a black-box proxy works. In speculative decoding, a drafted token is accepted with probability αt = min(1, p(yt |·)/q(yt |·)). When p is a black box, p(yt |·) is hidden, therefor if q tracks p reasonably well, a high speculative surprisal st implies p too would face a low acceptance rate there, making st a usable proxy for p’s own uncertainty. This is why aligning q to p matters. 3.3
Component-Aware Feature Extraction
An agentic trajectory alternates between a highentropy reasoning span and a sharply lowerentropy action span, so the speculative signals carry phase-dependent meaning (we analyze this phase structure in Appendix A). We therefore compute features separately for the reasoning and action tokens within the lookback window Wi . For each of the two phases (reasoning, action) and each of the three signals (s, g, H), we summarize the corresponding tokens with eight statistics—mean, variance, max, min, skewness, trend, and the mean over the first and last 10% of tokens, giving 48 features in total. Adding the two span lengths yields
the full feature vector ϕ(Wi ) ∈ R50 (full taxonomy in Appendix C). Computing the phases separately keeps the noisy reasoning signal from masking the small, failure-relevant deviations in the action span, and the short window keeps those deviations from being noised by earlier steps. 3.4
Calibrator and Downstream Policy
Following HTC (Zhang et al., 2026) framework, we fit an ℓ1 -regularized logistic regression on labeled windows. With given features ϕ(Wi ) ∈ R50 and binary labels y ∈ {0, 1}, the calibrator F(ϕ) = σ(w⊤ ϕ + b) is fit by minimizing X min L yj , σ(w⊤ ϕ + b) + λ∥w∥1 , w,b
j
where L is the binary cross-entropy, σ the logistic function, and λ the ℓ1 penalty that selects a sparse, auditable subset of the 50 features. We deliberately keep the calibrator simple, since a linear model is cheap to train and it is robust in the small-data regime typical of SWE benchmarks. The output of SU is a single scalar failurelikelihood score: the estimated probability ŷ = F ∗ (ϕ(Wi )) that the next action will follow the verifiable objective. What a system does with that probability is a separate, application-level decision, and the framework is deliberately agnostic to it. A calibrated failure signal can drive many policies, for instance, routing the step to a larger or more capable model when confidence is low, escalating to a human reviewer or allocating extra test-time compute (e.g., resampling) to risky steps. The veto gate. To make the deployment impact measurable, we study the policy we deploy and evaluate in this work: a veto gate inserted between code generation and execution (Algorithm 1). When the predicted success probability falls below a threshold τ , the gate blocks the action and triggers a cheaper replan instead of an execute-failretry cycle. We choose this policy because its effect is transparent. It maps the probability to a single, auditable accept-replan decision, so the error and cost reduction numbers in Section 4.3 are directly attributable to the quality of the underlying signal. The threshold τ trades true vetoes (blocking code that would fail) against false vetoes (blocking code that would succeed). τ is selected on the training split and frozen for all reported evaluations. We emphasize that the gate is one policy among many the signal could drive, not a requirement of the
Algorithm 1 SU Veto-Gate Policy Require: Window Wi , threshold τ , max retries K 1: Compute ϕ(Wi ) via draft teacher forcing ▷ single pass, O(|Wi |) 2: ŷ ← F ∗ (ϕ(Wi )) ▷ calibrated P (success) 3: if ŷ ≥ τ then 4: execute: send action ai to the environment 5: else if retries < K then 6: replan: append a failure hint and re-invoke the agent 7: else abstain: return with a low-confidence flag 8: 9: end if framework: any of the alternatives described above could consume ŷ in its place.
4
Experimental Results
We evaluate SU in agentic software engineering tasks, where Ω is whether generated code runs without exceptions: action spans are long, syntactically and contextually constrained, and the objective is binary verifiable feedback from the execution environment. Setup. We train both draft model and calibrator on the SWE-rebench OpenHands trajectories (Trofimova et al., 2025), real-world GitHub issues from SWE-rebench (Badertdinov et al., 2026) solved by Qwen3-Coder-480B under the OpenHands scaffold (Wang et al., 2025). Further, we evaluate out-of-distribution on a held-out split, on SWE-Bench Verified (Jimenez et al., 2024), and on DA-Code (Huang et al., 2024). Success is scored using the execution data derived from the agentic trajectories. We use Qwen3-Coder480B (Qwen Team, 2026) and the fully closedsource Claude 3.5 Sonnet (Anthropic, 2024) as agent models (p), and Qwen3-4B (Yang et al., 2025) as a draft model (q). In the study we validate base model, SFT-tuned model, and teacherforced (TF) distilled variants. As a baselines we use Verbalized Confidence (Tian et al., 2023) made by the agent; Last-TP and Global-TP, mean logprobability of the final block and whole trajectory respectively under p counting it as a white-box; and HTC (Zhang et al., 2026) on agent’s own logprobabilities. We ablate SU-Uniform, our signals without phase separation. We report AUROC, Precision at Recall 80 (P@R80) calibration metrics. For the gate mechanism evaluation we addition-
Table 1: Code-execution failure prediction (agent: Qwen3-Coder-480B; SU uses Qwen3-4B draft). † Indistribution held-out split; SWE-Bench Verified shares the agent and scaffold and differs only in task distribution; DA-Code additionally changes domain and tool interface. SWE-rebench† SWE-Bench Ver. P@R80
DA-Code
Method
AUC P@R80 AUC
AUC P@R80
Verbalized Conf.
.58
.50
.57
.49
.59
.51
Last-TP Global-TP HTC (Agent)
.70 .67 .85
.61 .57 .77
.68 .65 .82
.59 .55 .74
.69 .66 .83
.60 .56 .75
SU, no train SU-Uniform (SFT) SU, SFT SU, TF
.63 .74 .79 .80
.54 .65 .70 .71
.61 .71 .76 .77
.52 .62 .67 .68
.63 .72 .77 .78
.54 .63 .68 .69
ally report true-/false-veto rates and net error/cost change. Implementation details, training recipes, and compute budgets are described in Appendix B. 4.1
SWE-Bench Verified DA-Code No-veto error rate SU-gated error rate ∆ error Token savings / task
21% 15% −6 pp −14%
14% 6% −8 pp −19%
No-veto task success SU-gated task success ∆ success
49% 44% −5 %
64% 63% −1 %
True-veto recall False-veto rate
0.72 0.14
0.69 0.16
Table 3: Calibration and proper-scoring metrics on the held-out SWE-rebench split (agent: Qwen3-Coder480B, Qwen3-4B draft). The constant reference always predicts the base rate, it is trivially reliable and carries no information.
Failure-Prediction Quality
Table 1 shows failure-prediction quality across the three splits. First, importance of alignment: an untrained draft aligns poorly with the agent and is barely above verbalized confidence (.61 AUROC on SWE-Bench Verified), while distillation lifts SU to .77, and both SFT and TF land within 1. Second, phase separation is worth +5–6 AUROC: dropping it (SU-Uniform, .71) confirms the actionspan signal becomes worse once phases are pooled. Third, SU recovers most of the white-box ceiling: it lifts AUROC from .57 (verbalized confidence) to .77, approaching HTC (.82), using only output tokens and showing improvements over the whitebox log-probability baselines (Last-/Global-TP). It does not exceed HTC, and not expected to: a blackbox proxy approaching its white-box ceiling is the credible outcome. 4.2
Table 2: Deployment impact of the SU veto gate (Qwen3-Coder-480B agent, Qwen3-4B draft). SWEBench Verified and DA-Code are evaluated without additional fine-tuning.
Calibration
Discrimination and calibration are distinct properties, and SU delivers them unequally. Table 3 reports proper-scoring and calibration metrics on the held-out SWE-rebench split against a constant predictor that always emits the base rate. Alignment is what produces utility: distillation improves the Brier score from .1265 to .1012 and the Brier skill score from .019 to .215 over the constant reference, consistent with the AUROC gains of Section 4.1. We therefore describe SU’s raw output as a failure-likelihood score rather than a calibrated probability, and note that the veto gate of Sec-
Method
ECE↓ MCE↓ Brier↓ NLL↓ BSS↑
Constant (base rate) SU, no train SU, SFT
.008 .015 .024
.008 .054 .060
.129 .127 .101
.426 .422 .364
.000 .019 .215
tion 4.3 requires only a monotone score plus a threshold, so its results depend on ranking quality alone and are unaffected by the residual miscalibration. Standard post-hoc recalibration, such as temperature scaling or isotonic regression is orderpreserving and would leave AUROC and the gate unchanged while reducing ECE. We did not apply it here, and we flag closing the reliability gap as required work before SU’s output is consumed by any policy that reads the probability value itself rather than thresholding it. 4.3 Deployment Impact: Fewer Errors, Lower Cost We lead with the result that matters most for deployment: does the gate (Algorithm 1) actually change how the agent behaves on the two axes a practitioner is accountable for—error rate and compute? Table 2 shows that it does. Error rate. At the operating threshold, SU cuts the per-call execution error rate from 21% to 15% on SWE-Bench Verified and from 14% to 6% on DA-Code. Crucially, the DA-Code result was obtained without additional fine-tuning: the calibrator was trained on the SWE-rebench trajectories and applied without retraining.
Compute. Each failing trajectory, left alone, triggers a retry. By catching failures before execution and swapping the full retry loop for a cheaper replan, the gate cuts average tokens per task by 14– 19% on both benchmarks. At production volume this is a direct reduction in API spend and latency. Task success rate. A natural question is whether cutting errors before execution also lifts end-to-end task success. On SWE-Bench Verified the resolved rate moves slightly down, from 49% to 44%, and on DA-Code it is essentially unchanged (64% to 63%). The reason is structural: the veto gate suppresses confidently-wrong actions and swaps the execute-fail-retry loop for a cheaper replan, but a replan is not guaranteed to recover the task, and a small fraction of false vetoes block actions that would have succeeded. The deployment value of SU therefore lies in the cost and silent-failure axes, fewer wasted execute-fail-retry cycles and 14–19% lower token spend, rather than in raising the benchmark scores. Raising task success would require a downstream policy stronger than the simple replan we study here, which we leave as a promising direction for further research. 4.4
Cross-Agent Transfer
A deployment-critical question is whether a trained draft generalizes to agents it was not specifically aligned with, since teams frequently serve several agents behind a single interface. We examine this through several training configurations, summarized in Table 4. When a strong open-weight agent is available, the draft can be aligned to it via SFT or TF distillation. However, the two objectives yield closely matched scores (.76 vs. .77 on Qwen), showing that while the alignment is crucial, the particular distillation method is not as important. More importantly, a draft aligned to one agent still transfers to a different agent: a Qwen-distilled draft reaches .69 AUROC on Claude 3.5 Sonnet, well above the untrained baseline (.60). We attribute this to two factors. First, fine-tuning on agent trajectories improves alignment to the software engineering domain rather than to the specific agent alone. Second, the adapted draft is less surprised by in-domain trajectories than the base model, so its speculative signal remains informative even when the target agent differs. Most notably, training on the open-weight corpus alone is close to and competitive with mixed-
Table 4: Cross-agent transfer on SWE-Bench Verified (AUROC / P@R80; Qwen3-4B draft). Claude 3.5 Sonnet is fully black-box (no logits). Train corpus
Qwen3-480B
Claude 3.5 Sonnet
— Qwen (TF) Qwen (SFT) Claude (SFT) Mixed (SFT)
.61 / .52 .77 / .68 .76 / .67 .69 / .61 .76 / .67
.60 / .52 .69 / .61 .69 / .60 .76 / .67 .75 / .66
corpus training: a Qwen-distilled draft reaches .69 on Claude, within a few points of the mixed-corpus result (.75). This is promising, as it suggests that an open, logit-accessible agent can serve as an effective proxy for aligning drafts toward closed-source targets, motivating the post-training methods that exploit logit access or hidden representation on the open model while still transferring to black-box agents.
5
Discussion and Practical Takeaways
The deployment recipe. For teams running frontier API agents: pair the agent with a distilled 4B open-weight draft, distill it on a few thousands labeled trajectories from the target agent (or a mixed corpus if several agents are in play), fit a linear calibrator with phase-aware features, and insert the veto gate before execution. The whole pipeline adds one small forward pass per step and no changes to the agent, and in our study returns 6–8 pp lower error and 14–19% lower token cost. Transfer beyond code. SU needs only two things from a domain: trajectories with identifiable reasoning and action spans, and an external binary oracle over actions. Both hold beyond software engineering, such as SQL agents (Ω checks query results), web agents (Ω checks a target state), and tool-orchestration agents (Ω checks the right call with the right arguments). In each, the deliberationthen-commitment structure persists, and with it the entropy divergence SU relies on. We have not measured these settings and make no quantitative claim about them; confirming the transfer is the natural next step. From observation to intervention. The veto gate is not a passive monitor but an excplicit intervention: it reshapes the trajectory by injecting a replanning prompt exactly where failure is predicted. A natural extension is a streaming veto that watches the action span’s entropy token by token
and interrupts the instant it diverges and catching failure mid-generation, before the block is finished. We leave this to future work.
6
Conclusion
Production agents fail confidently and expensively, paying for each erroneous action before its consequences reveal the error. Speculative Uncertainty anticipates the failure instead: a small open-weight draft model scores a black-box agent’s trajectory in a single forward pass, and a linear calibrator turns the result into a failure-likelihood score that gates execution. Deployed this way it cut the percall execution error rate by 6–8 percentage points and token cost by 14–19%, with gains that transfer zero-shot and extend to a fully closed-source agent. Two design choices enable this: scoring the agent through a separate model rather than its internals, and respecting the reasoning–action phase structure of generation. We expect both to transfer to other verifiable-objective settings in which an agent reasons its way to a committed action. Where that structure recurs, an agent’s failure can be anticipated rather than merely discovered.
Limitations Quality of resulted models. SU consists of several interacting components, each of which influences final performance: the choice of draft model and its size, the alignment procedure, the feature construction, and the calibrator training. The combinatorial space these define is large, and we do not claim to identify its optimum. Our contribution is instead a lightweight and scalable framework together with empirical guidance on reasonable operating points. Tuning any individual component further is likely to yield additional gains we do not pursue here. A single objective and domain. Our evidence is drawn entirely from code-execution success as the verifiable objective. While the framework itself is agnostic to the objective—requiring only trajectories with identifiable reasoning and action spans and an external binary oracle—we provide no quantitative evidence for other settings. Extending SU to additional domains and objectives is an important direction for future work. Single-run estimates. All reported numbers are single-run point estimates on fixed evaluation sets;
we do not report seed variance or confidence intervals. Differences of a few AUROC points, and the 5 percentage points SWE-Bench Verified tasksuccess change, should be read with that in mind. The draft cost is objective-dependent. Throughout, we treat the draft’s per-step forward pass as effectively free relative to the agent and the avoided retries, which holds for the code-execution setting we study. This assumption may not transfer to objectives that demand a larger or more capable draft to produce an informative speculative signal, in which case the per-step overhead must be reevaluated against the downstream savings. The favorable cost balance we report is therefore specific to our workload and should not be assumed to hold universally.
Acknowledgments We acknowledge the use of LLMs for text polishing and language improvements throughout this manuscript. All technical content, ideas, and substantial writing remain the original work of the authors.
References Anthropic. 2024. Claude Sonnet 3.5. https://www. anthropic.com/news/claude-3-5-sonnet. Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich, Anton Shevtsov, Simon Karasik, Andrei Andriushchenko, Maria Trofimova, Daria Litvintseva, and Boris Yangel. 2026. Swe-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents. Advances in Neural Information Processing Systems, 38. Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318. Jinhao Duan, James Diffenderfer, Sandeep Madireddy, Tianlong Chen, Bhavya Kailkhura, and Kaidi Xu. 2025. Uprop: Investigating the uncertainty propagation of llms in multi-step agentic decision-making. arXiv preprint arXiv:2506.17419. Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024. A survey of confidence estimation and calibration in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6577–6595.
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR. Kait Healy, Bharathi Srinivasan, Visakh Madathil, and Jing Wu. 2026. Internal representations as indicators of hallucinations in agent tool selection. arXiv preprint arXiv:2601.05214. Yiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang, Fangyu Lei, Yifan Wei, Shizhu He, Lifu Huang, Xiao Liu, Jun Zhao, and 1 others. 2024. Da-code: Agent data science code generation benchmark for large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 13487–13521. Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pages 54107–54157. Jean Kaddour, Srijan Patel, Gbètondji Dovonon, Leo Richter, Pasquale Minervini, and Matt J Kusner. 2026. Agentic uncertainty reveals agentic overconfidence. arXiv preprint arXiv:2602.06948. Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664. Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org. Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023. Generating with confidence: Uncertainty quantification for black-box large language models. arXiv preprint arXiv:2305.19187. Qwen Team. 2026. Qwen3-coder-next technical report. arXiv preprint arXiv:2603.00729. Nishant Subramani, Jason Eisner, Justin Svegliato, Benjamin Van Durme, Yu Su, and Sam Thomson. 2025. Mice for cats: Model-internal confidence estimation for calibrating agents with tools. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 12362–12375. Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5433–5442.
Maria Trofimova, Anton Shevtsov, Badertdinov Ibragim, Konstantin Pyaev, Simon Karasik, and Alexander Golubev. 2025. Openhands trajectories with qwen3-coder-480b-a35b-instruct. Nebius blog. https://nebius.com/blog/posts/ openhands-trajectories-with-qwen3-coder-480b. Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, and 1 others. 2025. Openhands: An open platform for ai software developers as generalist agents. In International Conference on Learning Representations, volume 2025, pages 65882–65919. Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. 2024. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. Findings of the Association for Computational Linguistics: ACL 2024, pages 7655– 7671. Weihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao, Junjue Wang, and Naoto Yokoya. 2026. The confidence dichotomy: Analyzing and mitigating miscalibration in tool-use agents. arXiv preprint arXiv:2601.07264. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. John Yang, Kilian Lieret, Carlos Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang. 2026. Swe-smith: Scaling data for software engineering agents. Advances in Neural Information Processing Systems, 38. Chenlong Yin, Zeyang Sha, Shiwen Cui, Changhua Meng, and Zechao Li. 2025. The reasoning trap: How enhancing llm reasoning amplifies tool hallucination. arXiv preprint arXiv:2510.22977. Jiaxin Zhang, Caiming Xiong, and Chien-Sheng Wu. 2026. Agentic confidence calibration. arXiv preprint arXiv:2601.15778. Qiwei Zhao, Xujiang Zhao, Yanchi Liu, Wei Cheng, Yiyou Sun, Mika Oishi, Takao Osaki, Katsushi Matsuda, Huaxiu Yao, and Haifeng Chen. 2024. Saup: Situation awareness uncertainty propagation on llm agent. arXiv preprint arXiv:2412.01033.
A
The Phase Structure in Tool Augmented Agents
An agentic trajectory is the sequence of tokens an agent emits while solving a task. The exact format depends on the harness, but the dominant pattern across modern agent frameworks based on the reasoning LLMs is roughly the same: within each step the model first generates a reasoning span (planning, analysis) and then an action span (the tool call or code to be executed). Although both come from a single autoregressive pass, their token-level dynamics differ significantly. Figure 2 traces per-token entropy along agentic generations, with position normalized so that x = 0 marks the start of reasoning, x = 0.5 marks the transition to the action span, and x = 1 marks its end, separately for successful and failing actions over 500 SWE-rebench OpenHands trajectories (250 success, 250 failure). Two empirical findings stand out, and each shapes the method. First, reasoning is highly exploratory: entropy is high as the model weighs alternatives, then collapses sharply at the transition. Second, and most important, the failure signal is implicitly present in both phases. This drop of entropy has a direct, practical consequence. Since the failure signal points one way during reasoning and the opposite way inside the action span, any calibrator that combines the two phases into the single sequence of tokens does not merely represent the signal. Moreover, it actively cancels it: the elevated-reasoning and suppressedaction contributions offset each other. This motivates the central design choice of our method: extract features from the reasoning and action spans separately. The phase structure in Figure 2 is shown for a single step; in practice we aggregate these perphase statistics over a short window of recent steps (Section 3.1) to give the calibrator local trajectory context without the noise of the full history.
B
Reproducibility: Data, Training, and Compute
This appendix documents every step needed to reproduce the results in Section 4. We organize it as the pipeline runs: trajectory collection and labeling, draft-model training, calibrator fitting, and the gate simulation.
B.1
Trajectory Collection and Labeling
All trajectories are collected under a single agentic scaffold, OpenHands (Wang et al., 2025), so that train and test distributions differ only in task content, not in formatting or tool interface; this isolates genuine distribution shift from harness artifacts. For each task the agent emits a sequence of reasoning and action spans: we map the assistant message’s free-form text to the reasoning span and its structured tool call (the invoked function and its serialized arguments) to the action span, which we parse directly into the component labels τt of Section 3.1. Each action step is labeled by its execution outcome, the oracle verdict Ω(ai , x) the gate predicts (Section 3.1): 1 if the tool call executes without error (exit code 0), 0 otherwise. We took SWE-rebench OpenHands trajectories, which were generated by Qwen3-Coder-480B. We use 10,000 of these trajectories for training and a heldout 1,000 for testing, plus 500 SWE-Bench Verified and 500 DA-Code trajectories used only for outof-distribution evaluation. These trajectories are long-horizon: ∼65 agent turns and ∼54 thousands tokens of context on average. Across the training split this totals ∼0.54B context tokens for training. Table 5 summarizes the splits and their base tool-call success rates, this rate is the positive-class prevalence the calibrator must handle. Table 5: Dataset splits and base tool-call success rate (Qwen3-Coder-480B): the fraction of action steps whose tool call executes without error (exit code 0). Its complement is the base per-call execution-error rate. Dataset
#Traj.
#Calls
Succ. %
SWE-rebench (train) SWE-rebench (held-out)
10,000 1,000
327,098 32,585
84.7 84.7
B.2
Draft-Model Training
We compare three regimes for each draft size. No training uses the released Qwen3-4B checkpoint as-is. SFT fine-tunes the draft to reproduce the agent’s trajectories with a standard nexttoken cross-entropy loss on the agent’s emitted tokens. Teacher-forced (TF) distillation additionally aligns the draft’s full next-token distribution to the agent’s. We keep training-time and inferencetime access separate: at inference SU is strictly black-box, reading only the agent’s realized output tokens, whereas draft training is a one-time offline step. Because the agent Qwen3-Coder-480B is
Entropy of top-5 logprobs (nats)
1.0
mean ±1 SD
0.8
Mean entropy (nats)
Entropy profile: successful vs failed steps Qwen3-Coder-480B
Aggregated over all trajectories
0.6 0.4 0.2 0.0 1
2
3
Step index (each step unit width)
Reasoning
1.0
Tool call
successful step failed (non-zero exit / Traceback)
0.8 0.6 0.4 0.2 0.0
4
0
0.25
0.5 (transition)
0.75
1
Normalized position in step generation
Figure 2: Speculative entropy of agentic generations. Left: mean entropy aggregated over all steps—the per-step reasoning-to-action entropy drop is washed out and not visible. Right: entropy along a normalized step (reasoning x < 0.5, action x > 0.5) for successful (navy) vs. failing (red) actions, revealing the sharp collapse at the transition. The failure signal appears in both phases with opposite sign, failures are more uncertain during reasoning, less so in the action span, which is why phase-aware features help. Bands show ±1σ over 500 training trajectories.
open-weight, we recover its logits in a single forward pass over the collected trajectories and distill the smaller draft against this full distribution. This soft-distribution supervision yields only marginal gains over realized-token SFT (Table 1), consistent with our finding (Section 4.4) that the training corpus matters more than the distillation objective. The fully closed Claude 3.5 Sonnet exposes no logits and therefore cannot be TF-distilled; the draft used there is aligned by SFT on realized tokens alone. Table 6 shows hyperparameters, which were used for draft model training. Table 6: Draft-model fine-tuning hyperparameters, held fixed across sizes. Hyperparameter
Value
Base model Optimizer Learning rate Weight decay Gradient clipping Epochs Effective batch size Max sequence length Context extension Loss masking Precision LoRA / full FT
Qwen3-4B AdamW (β1 =0.9, β2 =0.95) 1e-5 (cosine, 5% warmup) 0.1 1.0 2 32 sequences 32768 tokens YaRN assistant tokens only bf16 full fine-tuning
The regularization strength λ is selected by 5-fold cross-validation on the SWE-rebench training split over the grid {10−3 , . . . , 101 }, optimizing validation AUROC; the selected value is then frozen for all test evaluations, including the zero-shot DACode transfer. B.4
Table 7 reports the cost of each pipeline stage. The figure most relevant to deployment is the per-step draft latency, since it is the only overhead SU adds on the critical path: a single teacher-forced forward pass over the trajectory-so-far, with no generation. For the 4B draft model this is ≈100 ms per step on a single A100 GPU, against much larger latencies of agent model, so the gate’s overhead is a small fraction of a step. Table 7: Compute budget by pipeline stage. Per-step latency is the only overhead on the agent’s critical path. Stage
Calibrator Fitting
The calibrator is an ℓ1 -regularized logistic regression over the ϕ(Wi ) ∈ R50 features of Appendix C. Features are standardized (zero mean, unit variance) using statistics computed on the training split only, to avoid leakage into the OOD evaluations.
Hardware Cost
Draft SFT/TF (4B) 4×A100 Calibrator fit 64×CPU Draft scoring (per step, 4B) 1×A100
C B.3
Compute and Latency
∼50 GPU-h <1 min ≈100 ms
Feature Taxonomy
Table 8 enumerates all 50 features in ϕ(T ). The construction is deliberately uniform: for each of the two component types (R = reasoning, A = action) and each of the three speculative signals (s = surprisal, g = gap, H = entropy), we compute the same eight aggregate statistics over the corre-
sponding token subset, giving 2 × 3 × 8 = 48 statistical features; the final two are the reasoningand action-span lengths. The eight statistics are chosen to capture complementary aspects of a signal’s distribution within a span: location (mean), spread (variance), extremes (max, min), asymmetry (skewness), drift (trend ∆ = µlast − µfirst ), and boundary behavior (mean over the first/last 10% of tokens). The boundary and trend statistics are what let a linear calibrator pick up the late actionspan entropy uptick of Figure 2: that uptick raises µlast and the trend ∆ specifically within the action component, while leaving the reasoning component untouched, which is only possible because the two components are featurized separately (Section 3.3). Table 8: Feature taxonomy for ϕ(T ) ∈ R50 . Group
Signal
# stats
Reasoning
Surprisal s Gap g Entropy H
8 8 8
Action
Surprisal s Gap g Entropy H
8 8 8
Structural
|I R |, |I A |
2