Conceptio › Archive › arXiv CS
arXiv CSopen access

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Chaoqian Ouyang et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

TokenCast: Forecasting Token Consumption During LLM Agent Execution

arXiv:2609.35760v1 [cs.LG] 28 Sep 2026

Chaoqian Ouyang1* Ling Yue2* Libin Zheng1† Huanghui Guo3 Shengxiang Xu3 YiShu Wang3 Ran Li4 Jian Yin1 Shaowu Pan2 Shimin Di3† 1

Sun Yat-Sen University 2 Rensselaer Polytechnic Institute 3 Southeast University 4 Hong Kong University of Science and Technology [email protected] [email protected]

Abstract When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast’s mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.

1

Introduction

The applications of large language models (LLMs) are expanding from simple question answering to complex tasks such as software engineering and deep research. Completing these tasks typically relies on LLM-based agents that repeatedly plan, modify code, invoke tools, and verify results [Yao et al., 2023, Yang et al., 2024, Wang et al., 2025]. A given task may be resolved in a single edit or may require multiple rounds of retries before converging, and the dialogue and tool outputs generated in each round accumulate into the context of subsequent requests [Yang et al., 2024, Xiao et al., 2026, Zhu et al., 2026]. The final token consumption of a task is therefore difficult to predict before execution completes, making it hard for users to forecast costs and plan usage [Bai et al., 2026]. * Equal contribution. † Corresponding authors.

This work is conducted during the internship in Prof. Di’s group.

1

The challenge of predicting tokens lies in the fact that every action an agent takes can trigger new model calls, and consumption across steps is interdependent: the longer the context left by earlier steps, the larger the input to every subsequent call, causing consumption to accumulate and amplify [Salim et al., 2026, Zhu et al., 2026, Xiao et al., 2026]. The model’s generative behavior further compounds the difficulty. Given the same input, the model may choose different courses of action, generate outputs of different lengths, and consequently undergo different numbers of verification or retry cycles. In practice, token consumption across different executions of the same task can differ by up to 30× [Bai et al., 2026], making accurate prediction challenging. Existing work can be organized by prediction scope and observation time. The scope ranges from a single response to an entire agent task; the forecast can be made before execution or updated during it. These dimensions give four settings (Figure 1): response length before generation, remaining response length during generation, total task consumption before execution, and remaining task consumption as an agent runs. Prior methods address these settings with different targets and access assumptions [Shahout et al., 2025, Xie et al., 2026, Bai et al., 2026].

1

Before generation

2

During generation

call start

input context

future output

Generation 2 now

2

1 generated prefix

3 Task

task start

3

4

remaining output

Before execution Generation 1

Generation 2

Generation K

During execution input tokens

output

output tokens

now

4

short path medium path

input

Gen 1 Gen 2 Gen 3 At the request level, prior work studlong path context accumulates ies how to predict the response length of a single model call. Because the in- Figure 1: Four token-consumption prediction settings. At put is known before generation begins, the call level, the input context supports a forecast before these methods typically estimate length generation (1), and the output prefix updates it during genfrom input features for resource alloca- eration (2). At the task level, total consumption is forecast tion [Jin et al., 2023, Qiu et al., 2024, Fu before execution (3), and remaining consumption is upet al., 2024, Zheng et al., 2026], while dated as calls complete and context accumulates (4). some also refine the estimate during generation using intermediate states or entropy statistics [Shahout et al., 2025, Xie et al., 2026, Merzouk et al., 2026].

At the agent-task level, prior work estimates total token consumption before execution, either from the task description or through the model’s own cost assessment [Bai et al., 2026]. Multi-step LLM execution frameworks organize or schedule requests according to program structures, semantic dependencies, or workflow paths available before the corresponding requests are executed [Khattab et al., 2023, Lin et al., 2024, Zheng et al., 2024, Ni et al., 2026, Yu et al., 2026]. Open-ended agents present a different situation: their behavior depends on tool feedback, environment state, and intermediate results, and no complete dependency graph exists before execution begins [Yao et al., 2023, Yang et al., 2024, Wang et al., 2025]. Multi-step forecasting has also been studied in structured workflows [Ni et al., 2026]. We focus on remaining provider-accounted input and output consumption in sequential agent execution. Each future request can bill retained context again, coupling the remaining cost to both future call count and input lengths. TokenCast updates this forecast from observed execution information without querying an LLM for the cost estimate. Figure 2 summarizes TokenCast. It represents each execution segment by call count, net input-length change, and a cost residual. An exact composition identity exposes how context growth in one segment affects the cost baseline of later calls. The task predictor combines a direct forecast with a prefix–suffix forecast conditioned on a predicted boundary state, and refreshes its features as 2

Introduction

第1页

• 应对:分层预测,并在执行过程中滚动更新 Observed execution

Unresolved future

visible prefix

LLM

Task

Agent

Tool

Call 1

Feedback

Task start Task total

Current Call

LLM

Tool

Future-call prediction

Task-level prediction

Call-level prediction Context

LLM

Task state

Response

LLM

Future-call tokens

+ Remaining token forecast

Generation Call start

Call total

Current-call remaining

Resource allocation

Budget control

Figure 2: Overview of TokenCast. Observed execution evidence supports call- and task-level forecasts, while the compositional path propagates predicted context growth into later-call costs. execution proceeds. Calibrated quantile models provide prediction intervals. The predictor makes no additional LLM calls. Our contributions are as follows: • We introduce a segment-cost factorization and an exact composition identity that accounts for repeated input consumption across calls. • We use this identity in a staged prefix–suffix predictor that combines direct and compositional forecasts and updates from execution evidence without additional LLM calls. • Across four benchmarks and six agent models, TokenCast’s MAE reduction averages 14.5% across 96 comparisons with the strongest comparator in each. Budget-control replay saves 21.3% of tokens at matched trace completion.

2

Related Work

Request-level output length prediction. Early methods extract features from the prompt to estimate response length for memory reservation, batching, or shortest-job-first scheduling [Jin et al., 2023, Qiu et al., 2024]. Prompt-only estimates remain highly uncertain, and prompt-conditioned response lengths can exhibit broad or heavy-tailed distributions [Perez-Ramirez et al., 2025, Wang et al., 2026]. Subsequent work therefore models more than a single point estimate: one line of work learns pairwise length orderings among requests to set scheduling priorities [Fu et al., 2024], and another fits a heavy-tailed log-t distribution that lets the scheduler trade off average and tail latency [Zheng et al., 2026]. Once generation begins, the decoder’s own intermediate states become available. Several recent studies show that mid-layer representations or entropy statistics can continuously refine the remaining-length estimate during decoding, narrowing the gap between the initial guess and the actual length [Shahout et al., 2025, Merzouk et al., 2026, Xie et al., 2026]. These methods share a common set of assumptions: the prediction target is a single response, the prompt is known, and the system has access to the prompt or to model internals. In agent tasks, later requests do not exist until earlier actions and tool calls complete, so none of these assumptions holds.

3

Task-level and multi-step prediction. Prior work has extended prediction to entire tasks. SelfPrediction has a coding agent inspect its environment before execution and estimate input, output, and total token consumption by stage [Bai et al., 2026]. DSPy [Khattab et al., 2023], Parrot [Lin et al., 2024], and SGLang [Zheng et al., 2024] represent multi-step LLM applications through program structures or semantic dependencies available to serving runtimes. Open-ended agents expose no such structure before execution. Chimera predicts remaining workflow output with a CPU-based quantile random forest [Ni et al., 2026]. Pythia profiles historical traces to infer likely workflow paths and role-level output lengths [Yu et al., 2026]. TokenCast’s segment identity addresses the repeated input cost induced by carried context, and its forecasts use no additional LLM calls. Trace studies of deployed agents report extensive iterative review loops in multi-agent software pipelines, and long contexts with short outputs and heavy prefix reuse in real coding-agent sessions [Salim et al., 2026, Zhu et al., 2026]. Broader analyses organize token use and efficiency across single-agent, multi-agent, and agent-ecosystem settings [Chen et al., 2026]. These studies characterize consumption after the fact and do not forecast it for a running task.

3

Method

3.1

Problem Formulation

An agent executes a task through a sequence of LLM calls interleaved with tool interactions. Let Ck denote the provider-accounted input and output token consumption of call k. A task terminates upon completion, failure, or when an execution limit is reached. For a task with K calls, its total PK Pk consumption is T = j=1 Cj . After call k completes, the confirmed consumption is Sk = j=1 Cj , and the remaining consumption is Rk = T − Sk . Each forecast uses only the task and execution information available at its prediction point. We consider four prediction settings along the execution trajectory. • Task Start predicts the total consumption T before execution begins. • Call Start predicts the consumption Ck of the current call after its request has been assembled and before any output token is generated. • In-call Update updates the prediction of Ck as the generated prefix becomes available. • Task Update predicts the remaining consumption Rk after call k completes, yielding an updated bk . forecast of the task total, Tbk = Sk + R 3.2

Call–Task Forecasting

TokenCast represents segment cost relative to its starting input length. The next segment inherits the context produced by the preceding one, yielding an exact composition rule for adjacent segments. Segment representation. A contiguous block of one or more calls forms a segment. Let segment A start with input length LA , span nA calls, and consume CA tokens in total. Its representation is ϕA = (nA , gA , bA ), where gA is the net change in input length across the segment and bA = CA − nA LA is the residual after subtracting the starting-input baseline nA LA , covering generation and within-segment context growth. The next segment starts with input length LA + gA . Reasoning tokens billed by the provider contribute to bA ; when they do not persist in the conversation context, they do not contribute to the context change gA . When segment B immediately follows A, the combined representation is  ϕA◦B = nA + nB , gA + gB , bA + bB + nB gA .

4

(1)

This identity follows from the definitions and the boundary condition LB = LA + gA . The third term nB gA arises from realigning the baseline: bB is defined relative to B’s actual starting point LA + gA , whereas the combined residual is relative to LA . The difference on each call is exactly gA , and nB calls accumulate to nB gA . Forecasting with composition. The decomposition converts aggregate remaining-cost prediction into three sub-problems: the prefix segment’s representation, the boundary state, and the suffix segment’s representation. TokenCast fits LightGBM models for these predictions and updates after completed calls. Appendix D.6 compares alternative base predictors. Input features are drawn from the task and execution information available at the prediction point. Appendix B.1 details these features and when each becomes available. Call-level forecasts predict the current call’s cost from the features visible at Call Start or In-call Update. Task-level forecasts maintain two paths. The direct path predicts remaining total cost as a single target. The compositional path separates the current or next segment from the subsequent suffix. Their cost representations are predicted separately and combined using Eq. (1). The compositional path predicts the prefix representation (b nA , gbA , bbA ) and its ending state. The predicted context change sets the suffix input baseline, while the predicted ending state and features visible at the original prediction point condition the suffix model. The suffix representation is converted to a remaining-cost forecast through Eq. (1). Updating forecasts. Each completed step during execution produces new observations, and TokenCast refreshes its forecasts accordingly. At Call Start, the current request has been assembled and the actual input length is known, so the call-level forecast can be based on it directly. At In-call Update, the committed generated prefix and streaming timing supply additional features for the current-call forecast. At Task Update, the completed call supplies confirmed cumulative cost Sk and any completed tool outcomes. The next request’s input length remains predicted until that request is bk and reports Tbk = Sk + R bk . assembled. TokenCast uses the available evidence to re-predict R 3.3

Compositional Learning

Staged fitting. TokenCast fits the direct, prefix, and suffix predictors as separate LightGBM models using labels extracted from completed traces. The direct path predicts remaining total cost. The compositional path predicts the prefix and suffix representations together with the prefix-ending boundary variables. Prefix context-change and suffix call-count errors can affect downstream cost through the composition identity, motivating the following weights on their local absolute losses: ℓg = nB |b g A − gA | ,

ℓn = LB |b n B − nB | .

(2)

Cross-fitting. The suffix model is trained on predicted prefix boundaries. Tasks are partitioned into F folds, and each fold’s prefix predictions are generated by models trained on the remaining folds. The predicted context change sets the suffix input baseline, so its residual label is recomputed as b B , where L b B = LA + gboof . The loss weights in Eq. (2) are motivated by the btrain = C B − nB L B A composition identity; Appendix B.3 distinguishes this motivation from the error decomposition under the rebased suffix training label. Correction model. After the direct, prefix, and suffix models are fixed, a correction model is trained on out-of-fold outputs of the complete forecasting pipeline: X bcomp,i − hψ (qi ) . ψ ∗ = arg min yi − C (3) ψ

i

The input qi contains the visible features, the direct and compositional forecasts, their difference, bcomp + hψ (q). and the predicted boundary variables. The final task-level forecast is C Prediction intervals. Separate LightGBM quantile models produce the 0.05 and 0.95 endpoints at each prediction point. The endpoints are widened symmetrically by a quantile of interval residuals 5

Table 1: MAE ↓ in tokens at four prediction points. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the MAE of the corresponding history-median predictor. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values. Agent LLM

Prediction point

TokenCast (Ours)

TRAIL

EGTP

TIE

Self-Pred.

SWE-bench Verified Task Start Call Start In-call Update Task Update Norm. Avg.

144.0k 64.5 38.9 80.0k 0.69

165.3k 160.9k 157.0k 71.1 70.3 69.5 77.4 78.9 74.6 115.0k 117.0k 119.0k 0.94 0.92 0.92

152.0k 70.9 80.2 126.0k 0.94

Task Start Call Start Qwen3.8-27B In-call Update Task Update Norm. Avg.

192.0k 87.0 81.9 124.0k 0.65

232.0k 209.0k 215.0k 109.0 102.9 99.0 120.0 116.0 111.4 176.0k 171.0k 164.0k 0.86 0.81 0.79

221.0k 105.5 108.9 173.0k 0.82

GPT-5.4

Search-R1 Task Start Call Start In-call Update Task Update Norm. Avg.

34.2k 34.5 31.3 17.6k 0.75

37.8k 37.3 39.4 23.1k 0.89

37.6k 36.5 38.5 23.7k 0.88

36.4k 34.0 36.8 23.3k 0.85

36.9k 39.8 39.9 24.2k 0.91

Task Start Call Start Qwen3.8-27B In-call Update Task Update Norm. Avg.

39.0k 56.7 57.0 24.0k 0.59

53.5k 65.9 84.6 39.5k 0.83

49.8k 71.4 80.7 34.0k 0.79

47.6k 67.7 75.5 35.4k 0.77

51.0k 76.1 83.2 38.4k 0.84

GPT-5.4

on held-out calibration tasks and are bounded below by consumption already confirmed within the prediction scope. Appendix B.4 gives the procedure.

4

Experiments

4.1

Experimental Setup

Tasks and execution traces. We evaluate TokenCast on SWE-bench Verified [Jimenez et al., 2024, Chowdhury et al., 2024], Search-R1 [Jin et al., 2025a], MMLU-Pro [Wang et al., 2024b], and LongBench-v2 [Bai et al., 2025], covering software engineering, retrieval-based question answering, knowledge-based reasoning, and long-context understanding. We collect 11,712 execution traces from 240 benchmark tasks with six agent LLMs: GPT-5.4 [OpenAI, 2026], Claude Opus 4.6 [Anthropic, 2026], Gemini 3.1 Pro [Google DeepMind, 2026], DeepSeek-V4-Pro [DeepSeek-AI, 2026], Qwen3.827B [Qwen Team, 2026], and Llama-3.2-3B-Instruct [Meta AI, 2024]. For reproducibility, we specify gpt-5.4-2026-03-05 for the GPT-5.4 API and qwen3.8-27b-20260815 for the self-hosted Qwen checkpoint; Table 5 lists model access and reasoning configurations. Traces are collected using DeepSeek Harness [DeepSeek, 2026] and OpenHands [Wang et al., 2025]. Repeated executions support the analysis of run-to-run variation, with additional repeats for anchor tasks. Table 2 specifies the collection design for each benchmark and harness. All runs of the same task remain in one partition. Training, validation, calibration, and test tasks are separated as described in Appendix C. Generalization experiments also use independently released trajectories (Appendix C.1). Baselines. We compare TokenCast with three output-length predictors, TRAIL [Shahout et al.,

6

2025], EGTP [Xie et al., 2026], and TIE [Zheng et al., 2026], and the agent-level consumption estimator Self-Prediction [Bai et al., 2026]. We adapt these methods to the prediction targets and observations available at the four prediction points. At Task Start, Self-Prediction inspects the task environment before estimating total consumption. At the other prediction points, it uses the observed execution prefix. Appendix C.3 details each adaptation. Appendix D.6 compares direct regression, compositional forecasting, their average, and the full correction pipeline; Appendix D.6 evaluates the segment representation and composition procedure. Metrics and implementation. We report mean absolute error (MAE) and weighted absolute percentage error (WAPE) at the four prediction points. At Call Start, the target includes the assembled request’s known input tokens. For 90% prediction intervals, we report empirical coverage, mean width, and mean interval score (MIS) [Gneiting and Raftery, 2007]. Tasks receive equal weight, with that weight distributed across their runs and evaluated checkpoints. Cross-configuration summaries divide MAE by that of a history-median predictor, which outputs the median training target for the same benchmark, agent LLM, and prediction point. The local platform provides eight NVIDIA A100 GPUs. Internal-state baselines use the agent model when accessible and a proxy for API-based models. Appendix C gives metrics, data partitions, training settings, and timing procedures. 4.2 4.2.1

Experimental Results and Analysis Main Results

Table 1 shows SWE-bench Verified and Search-R1 results for GPT-5.4 and Qwen3.8-27B; Figure 3 covers all four benchmarks and six agent LLMs. On SWE-bench Verified with GPT-5.4, TokenCast reduces MAE relative to the strongest comparator by 47.9% at In-call Update, from EGTP’s 74.6 to 38.9 tokens, and by 30.4% at Task Update, from TRAIL’s 115k to 80k tokens. At Task Start, the strongest comparator is Self-Prediction, and the reduction is 5.3%, from 152k to 144k tokens. Across the 96 benchmark–model–prediction-point combinations in Appendix D.2, its MAE reduction against the lowest comparator MAE averages 14.5%. The averages are −2.2% at Task Start, 1.9% at Call Start, 30.4% at In-call Update, and 27.8% at Task Update. TokenCast trails the strongest comparator in 24 combinations: 15 at Task Start and nine at Call Start. None of these losses occurs at In-call Update or Task Update, where execution evidence accumulates. Error relative to consumption. WAPE complements MAE by expressing absolute error relative to mean target consumption. For GPT-5.4 at Task Start, TokenCast’s MAE of 32.2k tokens on LongBench-v2 corresponds to a WAPE of 4.2%, whereas its MAE of 8.0k on MMLU-Pro corresponds to 15.8%. Their mean target consumptions are 771.2k and 50.6k tokens, respectively. Table 11 in Appendix D.3 reports the full WAPE results. 4.2.2

Prediction Reliability

Run-to-run variation. Repeated GPT-5.4 executions of the same task show substantial consumption spread across all four benchmarks. Figure 6 in Appendix D.1 shows the distributions for 48 anchor tasks and gives the repetition counts per benchmark. Interval reliability. On the anchor tasks, TokenCast’s calibrated intervals reduce MIS relative to Self-Prediction’s native intervals at all four prediction points, by 32.0% on average. At Task Start, TokenCast covers 82.0% of outcomes against a nominal 90% level, while Self-Prediction covers 52.7%. Table 3 reports interval width, coverage, and MIS at every prediction point. 4.2.3

Generalization

Unseen task types. On independently released LiveClawBench trajectories [Long et al., 2026], zero-shot leave-one-domain-out transfer yields MAE ratios of 1.31 at Call Start and 1.47 at Task Update relative to Self-Prediction. With 20 target-domain tasks for adaptation, the ratios fall to 0.82 and 0.85. Appendix D.4.1 gives the protocol and intermediate results.

7

TokenCast

TRAIL

EGTP

TIE

Self-Prediction

(a) SWE-bench Verified

(b) Search-R1

(c) MMLU-Pro

(d) LongBench-v2

Normalized average MAE (↓)

1.0 0.9 0.8 0.7 0.6

Normalized average MAE (↓)

1.0 0.9 0.8 0.7 0.6

. T-5 GP

4 au Cl

de

u Op

s4

.6 ni mi Ge

3 .1

Pr

o S ep De

4 -P k -V ee

ro 3 en Qw

.8 -

27

B Ll

. a-3 am

2 -3

B

. T-5 GP

4 au Cl

de

u Op

s4

.6 ni mi Ge

3 .1

Pr

o S ep De

4 -P k -V ee

ro 3 en Qw

.8 -

27

B Ll

. a-3 am

2 -3

B

Figure 3: Normalized average MAE across four benchmarks and six agent LLMs. Each panel corresponds to one benchmark, and each bar group corresponds to one agent LLM. For each method, MAE is normalized by the corresponding history-median MAE at each prediction point and macroaveraged over the four prediction points. Lower is better. Table 2: Execution traces per benchmark and harness, over six agent LLMs. Benchmark

Harness

SWE-bench Verified SWE-bench Verified Search-R1 Search-R1 MMLU-Pro MMLU-Pro LongBench-v2 LongBench-v2

DeepSeek Harness OpenHands DeepSeek Harness OpenHands DeepSeek Harness OpenHands DeepSeek Harness OpenHands

Tasks Runs Anchors Anchor runs 144 144 32 32 32 32 32 32

Total

2 2 4 4 4 4 4 4

24 24 8 8 8 8 8 8

8 8 16 16 8 8 8 8

Traces 2,592 2,592 1,344 1,344 960 960 960 960 11,712

Unseen agent LLMs. At Call Start, transfer to a held-out agent LLM yields 1.02 times the target history-median MAE without target-model tasks, improving to 0.95 with 20 tasks. With 3–10 target tasks, transfer outperforms target-only training. Appendix D.4.2 gives the protocol and a separate remaining-output-token Task Update analysis; other generalization results appear in Appendix D.4. 4.3

Case Study

TokenCast’s forecasts are revised as execution unfolds. In a code-repair run, a verification failure raises the task-total forecast, which decreases after a compatibility workaround passes the reported assertions. In long-context QA, file operations after answer generation consume further tokens and raise the forecast. Both traces are in Appendix E. We next examine token-budget control driven by these online forecasts on SWE-bench Verified. Budget control. On 288 GPT-5.4 runs from 144 SWE-bench Verified tasks, Figure 4 shows that TokenCast uses 21.3% fewer tokens on average across seven replay budgets while matching fixedbudget trace completion at every budget. A run is trace-complete when it reaches its recorded terminal 8

TokenCast

TRAIL

EGTP

TIE

Self-Pred. 100

Completed runs (%)

Completed runs (%)

100 80 60 40 20 0

Fixed budget

0

100

200

300

80 60 40 20 0

400

0

50 100 150 200 250 300 350

Tokens per run (k)

Wall time per run (s)

(a) Tokens

(b) Wall time

Figure 4: Trace completion under stopping limits on SWE-bench Verified. Panel (a) plots trace completion against mean tokens per run, and panel (b) plots it against replay-accounted wall time per run. Prediction overhead is included. The dashed curve is the fixed-budget baseline. Table 3: 90% prediction intervals over repeated executions of anchor tasks. TokenCast uses held-out calibration; Self-Prediction uses its reported 5th and 95th percentiles. TokenCast

Self-Prediction

Prediction point

Coverage (%)

Width

MIS ↓

Coverage (%)

Width

MIS ↓

Task Start Call Start In-call Update Task Update

82.0 89.8 92.8 91.2

752.5k 315 319 379.1k

1547.9k 549 576 574.3k

52.7 73.0 75.2 77.0

354.3k 223 237 232.0k

3334.6k 638 732 941.6k

state; the replay does not measure task resolution. Prediction costs are included. Appendix D.7 gives the stopping rules, per-budget results, and overhead (Table 21). 4.4

Ablation Study

Feature analysis. Figure 5 shows that the useful evidence changes over a run. Request length leads at Task Start and Call Start, with relative importance of 0.24 and 0.30. During generation, committed prefix length dominates at 0.58. After a call completes, last-call input tokens, remaining plan items, and calls without progress carry similar importance at 0.28, 0.25, and 0.23. Segment representation. The segment triple improves both task-level and call-level forecasts. Replacing it with total token count, input length, and output length raises their normalized MAEs from 0.71 to 0.88 and from 0.68 to 0.82. The no-composition variant has a normalized average MAE of 0.78 and call-level MAE of 0.74, compared with 0.69 and 0.68 for the full method. Within the triple, omitting context growth yields 0.76, while omitting the residual yields 0.73. Forecasting strategy. Before execution, direct regression has lower normalized MAE than composition, 0.82 versus 0.87. The ranking reverses at Task Update after calls have completed: composition reaches 0.66 and direct regression 0.72. Averaging the two forecasts reduces Task Update MAE to 0.64, and the full correction model reaches 0.62 while raising interval coverage from 88.1% under direct regression to 90.6%. Appendix D.6 gives the full comparison. Pipeline components. The largest degradation comes from dropping both cross-fitting and cost weighting. Normalized average MAE rises from 0.69 to 0.78, while 90% interval coverage falls from 90.6% to 87.4%. Dropping cross-fitting alone gives 0.73 MAE, and dropping cost weighting alone gives 0.72. The correction model and predicted boundary variables also contribute: their separate ablations give 0.74 and 0.72 MAE. Among six base predictors, LightGBM gives the lowest error and

9

Task information

Execution history

Current request

Prefix and timing

Feedback

Task Start

Task Update

Task type

0.11

Reasoning mode

0.10

Code block count

0.08

Task difficulty

0.07 0.05

0.00

500

0

-500

-1,000

0.10

0.20

0.30

101

Relative importance

1,000

No-progress calls

0

Cumulative output tokens

103

104

0.23 0.09

Completed calls

-1,000

102

103

Task description length

0.07

Latest tool-result size

0.04

Calls since plan change

0.03

Input-token trend

0.02

0.00

105

104

0.10

0.20

(f) Completed calls

(j) Feature importance

0.10

Remaining plan items

0.08

Calls since plan change

Last-call output trend 0.00

0.06

Prediction contribution

0.11

Latest tool-result size

No-progress calls

Prediction contribution

Cumulative output tokens

1,000

0

-1,000

1,000

0.10

0.20

0.30

0.40

Relative importance

102

103

0 -1,000 -2,000 -3,000

105

104

0

10

20

30

Remaining plan items

(k) Committed prefix bytes

0.17

Elapsed generation time

0.11

First-output delay

500

0.06

Request input tokens

0 -500

Latest tool-result size

0.04

Open bracket count

0.02

JSON depth 0.01

-1,000

-2,000

0.03

-2,000

(l) Elapsed generation time

1,500

1,500

1,000

1,000

0.58

Committed prefix bytes

2,000

0.14

-1,000

1,000

Last-call input tokens

Prediction contribution

0.30 0.18

0

-3,000 102

0.30

1,500

Completed calls

2,000

1,000

In-call Update

(e) Request input tokens

Request input tokens

3,000

2,000

Relative importance

Initial request length

Call Start (d) Feature importance

(i) Remaining plan items

3,000 0.28 0.25

Remaining plan items

-2,000 102

(h) Last-call input tokens

Last-call input tokens

2,000

Prediction contribution

Prediction contribution

0.15

Available tool count

Context window

1,000

0.20

(g) Feature importance

Prediction contribution

0.24

Initial request length Task description length

(c) Initial request length

Prediction contribution

(b) Task description length

Prediction contribution

(a) Feature importance

500 0 -500 -1,000

500 0 -500 -1,000

Recall similarity 0.01 103

104

Request input tokens

105

-1,500

0

10

20

30

40

50

60

0.00

Completed calls

0.20

0.40

0.60

Relative importance

0.80

-1,500 128

512

2k

8k

Committed prefix bytes

32k

-1,500 10−1

100

101

102

Elapsed generation time / s

Figure 5: Feature analysis at the four prediction points. For each point, the panels show the relative importance of selected features and contribution patterns for two representative features. Colors indicate the feature groups defined in Appendix B.1. the highest coverage with 0.8 ms model time. Appendix D.6 reports the complete component and base-predictor results in Tables 19 and 16. Update frequency. Updating after every call takes 32.8 ms of mean cumulative prediction time and makes 19.7 forecasts per run. Updating every three calls cuts the time to 11.9 ms and the forecast count to 6.9, and a single forecast at Task Start takes 2.1 ms. These full-pipeline times include feature extraction and model inference, so every-call updating adds under 0.03% to the 129 s median wall time of a GPT-5.4 run. Appendix D.5 reports the full frequency sweep with per-run standard deviations for all four update intervals.

5

Conclusion

TokenCast addresses the problem of forecasting token consumption during LLM agent execution, where context accumulation causes each call’s cost to depend on the entire preceding trajectory. The method represents each execution segment with a composable cost triple that separates the startinginput baseline from incremental consumption, and composes adjacent segments to propagate context growth into downstream cost estimates. Across four benchmarks and six agent LLMs, TokenCast’s MAE reduction against the strongest comparator averages 14.5% over 96 evaluated combinations and, in offline budget-control replay, it matches a fixed budget’s trace-completion rate while consuming 21.3% fewer tokens. Its mean cumulative task-level prediction time is 32.8 ms per run on SWE-bench Verified. Because the predictors are learned from recorded traces, transfer to a new task domain or agent LLM improves with a small set of target tasks from the new setting. This work provides a lightweight prediction layer for agent token consumption, and we leave the integration of these forecasts into a runtime decision framework that actively manages execution under budget constraints as a direction for future work.

10

Ethics Statement Our experiments use public benchmarks and recorded agent trajectories to study token consumption in LLM agents. The study involves no human subjects or collection of private user data. Models and datasets are used under their respective licenses and terms of use.

Reproducibility Statement Code associated with this study is linked at https://github.com/DEFENSE-SEU/TokenCast. Section 3 presents the forecasting formulation and training objective, with implementation details in Appendix B. Appendix C documents benchmark sampling, trajectory collection, agent and harness configurations, data splits, baseline implementations, and evaluation metrics. The budget-control replay protocol and complete Self-Prediction prompt are provided in Appendices D.7 and F, respectively.

AI Use Statement During manuscript preparation, large language models were used solely as general-purpose writing assistants for grammar checking, word refinement, and improving clarity. LLMs did not contribute to the research ideation, methodological design, or experimental execution. All suggestions produced by the LLMs were reviewed, edited, and vetted by the authors, who take full responsibility for the entire content of the paper.

References Anthropic. Introducing Claude Opus 4.6, 2026. URL https://www.anthropic.com/news/ claude-opus-4-6. Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, and Jiaxin Pei. How do AI agents spend your money? analyzing and predicting token consumption in agentic coding tasks. CoRR, abs/2604.22750, 2026. doi: 10.48550/ARXIV.2604. 22750. URL https://doi.org/10.48550/arXiv.2604.22750. Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, pages 3639–3664. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.ACL-LONG.183. URL https://doi.org/10.18653/v1/2025. acl-long.183. Yuxi Chen, Junming Chen, Chenyu He, Yiwei Li, Yicheng Ji, Yifan Wu, Dingyu Yang, Lansong Diao, Lidan Shou, Hongliang Zhang, et al. Token economics for llm agents: A dual-view study from computing and economics. arXiv preprint arXiv:2605.09104, 2026. Neil Chowdhury, James Aung, Chan Jun Shern, Oliver Jaffe, Dane Sherburn, Giulio Starace, Evan Mays, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E. Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry. Introducing SWE-bench verified, 2024. URL https://openai.com/index/introducing-swe-bench-verified/. DeepSeek. DeepSeek Harness, 2026. URL https://www.deepseek.com/harness/en/. DeepSeek-AI. Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

11

Yichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao, Ion Stoica, and Hao Zhang. Efficient llm scheduling by learning to rank. Advances in Neural Information Processing Systems, 37:59006–59029, 2024. Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007. Google DeepMind. Gemini 3.1 Pro model card, February 2026. URL https://deepmind. google/models/model-cards/gemini-3-1-pro/. Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6609–6625, 2020. Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. Swe-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id= VTF8yNQM66. Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. Searchr1: Training llms to reason and leverage search engines with reinforcement learning. CoRR, abs/2503.09516, 2025a. doi: 10.48550/ARXIV.2503.09516. URL https://doi.org/10. 48550/arXiv.2503.09516. Jiajie Jin, Yutao Zhu, Zhicheng Dou, Guanting Dong, Xinyu Yang, Chenghao Zhang, Tong Zhao, Zhao Yang, and Ji-Rong Wen. Flashrag: A modular toolkit for efficient retrieval-augmented generation research. In Guodong Long, Michale Blumestein, Yi Chang, Liane Lewin-Eytan, Zi Helen Huang, and Elad Yom-Tov, editors, Companion Proceedings of the ACM on Web Conference 2025, WWW 2025, Sydney, NSW, Australia, 28 April 2025 - 2 May 2025, pages 737–740. ACM, 2025b. doi: 10.1145/3701716.3715313. URL https://doi.org/10.1145/3701716.3715313. Yunho Jin, Chun-Feng Wu, David Brooks, and Gu-Yeon Wei. S3 : Increasing GPU utilization during generative inference for higher throughput. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.cc/paper_files/paper/2023/hash/ 3a13be0c5dae69e0f08065f113fb10b8-Abstract-Conference.html. Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601– 1611, 2017. Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. Dspy: Compiling declarative language model calls into selfimproving pipelines. CoRR, abs/2310.03714, 2023. doi: 10.48550/ARXIV.2310.03714. URL https://doi.org/10.48550/arXiv.2310.03714. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466, 2019. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Jason Flinn, Margo I. Seltzer, Peter Druschel, Antoine Kaufmann, 12

and Jonathan Mace, editors, Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, pages 611–626. ACM, 2023. doi: 10.1145/ 3600006.3613165. URL https://doi.org/10.1145/3600006.3613165. Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, and Lili Qiu. Parrot: Efficient serving of llm-based applications with semantic variable. In Ada Gavrilovska and Douglas B. Terry, editors, 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 929–945. USENIX Association, 2024. URL https://www.usenix.org/conference/osdi24/ presentation/lin-chaofan. Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, et al. Liveclawbench: Benchmarking llm agents on complex, real-world assistant tasks. arXiv preprint arXiv:2604.13072, 2026. Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers), pages 9802–9822, 2023. Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, and Adam Oberman. How much is left? llms linearly encode their remaining output length. CoRR, abs/2607.05316, 2026. doi: 10.48550/ARXIV.2607.05316. URL https://doi.org/10.48550/arXiv. 2607.05316. Meta AI. Llama-3.2-3B-Instruct. Hugging Face, 2024. URL https://huggingface.co/ meta-llama/Llama-3.2-3B-Instruct. Kangqi Ni, Wenyue Hua, Xiaoxiang Shi, Jiang Guo, Shiyu Chang, and Tianlong Chen. Chimera: Latency- and performance-aware multi-agent serving for heterogeneous llms. CoRR, abs/2603.22206, 2026. doi: 10.48550/ARXIV.2603.22206. URL https://doi.org/10. 48550/arXiv.2603.22206. OpenAI. Introducing GPT-5.4, March 2026. introducing-gpt-5-4/.

URL https://openai.com/index/

Daniel F. Perez-Ramirez, Dejan Kostic, and Magnus Boman. CASTILLO: characterizing response length distributions of large language models. CoRR, abs/2505.16881, 2025. doi: 10.48550/ ARXIV.2505.16881. URL https://doi.org/10.48550/arXiv.2505.16881. Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. Measuring and narrowing the compositionality gap in language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5687–5711, 2023. Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, and Ravishankar K. Iyer. Efficient interactive LLM serving with proxy model-based sequence length prediction. CoRR, abs/2404.08509, 2024. doi: 10.48550/ ARXIV.2404.08509. URL https://doi.org/10.48550/arXiv.2404.08509. Qwen Team. Qwen3.8-Max: A new bar for coding and cowork, August 2026. URL https: //qwen.ai/blog?id=qwen3.8. Stephen Robertson and Hugo Zaragoza. The probabilistic relevance framework: Bm25 and beyond. Foundations and trends® in information retrieval, 4(1-2):1–174, 2009. Mohamad Salim, Jasmine Latendresse, SayedHassan Khatoonabadi, and Emad Shihab. Tokenomics: Quantifying where tokens are used in agentic software engineering. arXiv preprint arXiv:2601.14470, 2026.

13

Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, and Michael Mitzenmacher. Don’t stop me now: Embedding based scheduling for LLMS. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=7JhGdZvW4T. Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. musique: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554, 2022. Jing Wang, Yu-Yang Qian, Ke Xue, Chao Qian, Peng Zhao, and Zhi-Hua Zhou. Robust length prediction: A perspective from heavy-tailed prompt-conditioned distributions. CoRR, abs/2604.07931, 2026. doi: 10.48550/ARXIV.2604.07931. URL https://doi.org/10.48550/arXiv. 2604.07931. Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. Executable code actions elicit better llm agents. arXiv preprint arXiv:2402.01030, 2024a. Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, and et al. Openhands: An open platform for AI software developers as generalist agents. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=OJd3ayDDoF. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024b. URL http://papers.nips.cc/paper_files/paper/ 2024/hash/ad236edc564f3e3156e1b2feafb99a24-Abstract-Datasets_ and_Benchmarks_Track.html. Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. C-pack: Packed resources for general chinese embeddings. In Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang, editors, Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024, pages 641–649. ACM, 2024. doi: 10.1145/3626772.3657878. URL https://doi.org/10.1145/3626772.3657878. Yuan-An Xiao, Pengfei Gao, Chao Peng, and Yingfei Xiong. Reducing cost of LLM agents with trajectory reduction. Proc. ACM Softw. Eng., 3(FSE):1241–1263, 2026. doi: 10.1145/3797084. URL https://doi.org/10.1145/3797084. Huanyi Xie, Yubin Chen, Liangyu Wang, Lijie Hu, and Di Wang. Predicting LLM output length via entropy-guided representations. CoRR, abs/2602.11812, 2026. doi: 10.48550/ARXIV.2602.11812. URL https://doi.org/10.48550/arXiv.2602.11812. John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. URL http://papers.nips.cc/paper_files/paper/2024/ hash/5a7c947568c1b1328ccc5230172e1e7c-Abstract-Conference.html.

14

Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 2369–2380, 2018. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/forum?id=WE_vluYUL-X. Shan Yu, Junyi Shu, Yuanjiang Ni, Kun Qian, Xue Li, Yang Wang, Jinyuan Zhang, Ziyi Xu, Shuo Yang, Lingjun Zhu, et al. Pythia: Exploiting workflow predictability for efficient agent-native llm serving. arXiv preprint arXiv:2604.25899, 2026. Haoyu Zheng, Yongqiang Zhang, Fangcheng Fu, Xiaokai Zhou, Hao Luo, Hongchao Zhu, Yuanyuan Zhu, Hao Wang, Xiao Yan, and Jiawei Jiang. Scheduling LLM inference with uncertainty-aware output length predictions. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id=I5IMkvVKd7. Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark W. Barrett, and Ying Sheng. Sglang: Efficient execution of structured language model programs. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 15, 2024, 2024. URL http://papers.nips.cc/paper_files/paper/2024/hash/ 724be4472168f31ba1c9ac630f15dec8-Abstract-Conference.html. Kan Zhu, Mathew Jacob, Chenxi Ma, Yi Pan, Stephanie Wang, Arvind Krishnamurthy, and Baris Kasikci. Tracelab: Characterizing coding agent workloads for llm serving. arXiv preprint arXiv:2606.30560, 2026.

15

Appendix A

Comparison with Related Methods

Table 4 compares prediction targets, evidence, and update points. TokenCast forecasts call and remaining-task consumption from execution features and a segment-cost representation. Table 4: Prediction targets and estimators in related work. Appendix C.3 describes our adaptations. Work

Prediction target

Evidence and estimator

Update point

TRAIL

Response length

Before generation

EGTP

Response and remainingresponse length Response length distribution Remaining workflow output

Last-token hidden state and length bins Hidden states and entropy pooling

TIE Chimera Pythia Self-Prediction TokenCast

Workflow path and role output length Total task tokens Provider-accounted call and remaining-task tokens

B

Method Details

B.1

Execution Evidence

Pre/mid generation

Text embedding and log-t model Prompt and workflow features; quantile forest Historical trace profiles

Before generation Workflow request

Agent inspection of task environment Execution features; direct and compositional predictors

Before execution

Incoming request

Four points

prediction

The observation record v contains only information available at the prediction point. Token usage comes from provider records, and a logical call aggregates its initial request and any retries. The input length of an assembled request is measured directly; a future request’s input length remains predicted until assembly. The visible prefix contains streamed text and tool-call arguments committed so far, with a checkpoint every 128 UTF-8 bytes. Unavailable fields are masked. Prediction-point observations remain separate from the prefix-ending predictions passed to the suffix model. Task representation. The task representation d uses TF-IDF features over word unigrams and bigrams and character 3- to 5-grams, with 12,000 terms for each representation. When a benchmark provides a task-difficulty label, the label is prepended to the task statement before encoding. Action types. Reads include file reads, glob matches, and searches. Edits write or replace file content. Tracking updates the to-do list. A shell command is labeled as a test when it invokes pytest, tox, nox, or unittest, or refers to a test directory, and other shell commands are labeled as runs. An action is marked as failed when a tool reports an error or a shell command exits with a nonzero status. The stage of a call is determined by its most recent action. Execution history. The progress record contains the numbers of planned, completed, and remaining to-do items, together with the number of calls since the list was last updated. The consumption history contains the input, output, and reasoning usage of the latest completed call, their changes from the preceding call, their mean, median, and least-squares trends over the last five and ten calls, and running totals of calls, model requests, tool actions, and confirmed tokens. Task Update additionally records the numbers of failed requests and tool errors in the latest call, together with the numbers of consecutive calls that failed, required a retry, or made no progress. Progress is defined as an edit, test, tracking update, or completed plan item. A window statistic is masked when the required observations are unavailable. Recent tool actions. The w most recent tool actions fill the slots in recency order. Each slot records 16

the action type, status, result size, calls since execution, and the reported result: an exit code, the number of lines returned by a read, the number of matches returned by a search, or the passed and failed counts of a test run. We select w per fold from {3, 5, 8} using the validation tasks. Generated prefix. The lowercased prefix is hashed by words and word pairs into 256 signed buckets with weight 1 + log n for n occurrences, and its last line is hashed into 64 buckets. A single scan records whether the prefix is a valid JSON prefix, its nesting depth, whether it ends inside a string and under which tool-argument key, and, within string contents, unmatched brackets, quote parity, open code fences, heredocs, and the termination pattern of the last line. Prefix recall. Earlier completed model requests from the same run serve as references, each represented by its hashed prefix at every checkpoint and its final generated length. The current prefix is compared with these references using cosine similarity, and the record includes the final and remaining lengths associated with the most similar requests and checkpoints. Stream timing. For the current request, let s denote its start time and let c1 < · · · < cm denote the output-checkpoint times. The timing record contains the time to the first checkpoint c1 − s, elapsed time cm − s, streaming time cm − c1 , committed bytes divided by streaming time, time since the previous checkpoint, the longest gap between checkpoints, and the initial-delay fraction (c1 − s)/(cm − s). Fields with undefined denominators or too few observations are masked. B.2

Predictors and Training

Predictors. TokenCast uses separate LightGBM models for the four prediction settings. Call Start and In-call Update predict the complete consumption of the current call. Task Start and Task Update use a direct path and a compositional path. The direct path predicts remaining task consumption as one target. The compositional path predicts the current or next segment, its ending boundary, and the subsequent suffix, then combines the two segment representations using Eq. (1). Training instances. Labels are extracted from completed traces. Call-level instances use the realized complete-call consumption. Task-level instances contain the remaining-consumption target, the prefix representation and ending boundary, and the suffix representation. Model inputs contain only the evidence available at the corresponding prediction point. Task-level fitting. The direct and prefix models are fitted first. Task-level cross-fitting then produces out-of-fold prefix boundaries for suffix training. The suffix residual label is recomputed against the predicted input baseline as described in Section 3.3. After these models are fixed, the complete pipeline produces out-of-fold forecasts used to fit the correction model in Eq. (3). Inference. Call-level forecasts use the corresponding fitted model directly. Task-level forecasts compute the direct prediction, the prefix boundary, the suffix prediction, and the compositional forecast before applying the correction model. The point models remain fixed during execution. Raw interval endpoints are produced by the quantile models in Appendix B.4. B.3

Theoretical Analysis

Composition properties. Let LA and LB be the initial input lengths of adjacent segments A and B. Since LB = LA + gA and CX = nX LX + bX for X ∈ {A, B}, CA◦B = (nA + nB )LA + bA + bB + nB gA = CA + CB . The cross-term nB gA arises when the cost of B is expressed relative to the initial input length of A. For three consecutive segments A, B, and D, either grouping yields the residual bA + bB + bD + nB gA + nD (gA + gB ). The composition is therefore associative. Single-call representation. A single call has representation ϕj = (1, Lj+1 − Lj , Cj − Lj ). For the terminal call, set LK+1 = LK to close the notation; no suffix follows it. Composing these representations in execution order recovers the representation of any contiguous segment. For a

17

complete trace of K calls, the resulting residual is

PK

j=1 Cj − KL1 . Hence

T = KL1 + bϕ1 ◦···◦ϕK , regardless of how the trace is partitioned. Error propagation and loss weights. In a compositional forecast, the suffix is predicted from the estimated boundary of the prefix. Let LB = LA + gA be the true input length at the start of the suffix, and let δgA and δnB denote prediction errors in the prefix’s context change and the suffix’s call count. With LA fixed, and writing δbB for the error in the suffix residual, the suffix cost error is ∆CB = nB δgA + LB δnB + δnB δgA + δbB . The coefficients nB and LB motivate the local weights in Eq. (2) under a true-boundary residual. b B . Writing eb = bbB − btrain and en = n Suffix training instead uses btrain = CB − n B L b B − nB B B bB − CB = L b B en + eb . Boundary prediction can affect the learned residual and the suffix gives C model’s inputs. Cross-fitting exposes suffix training to out-of-fold boundary errors. B.4

Interval Calibration

After the point models are fixed, separate LightGBM quantile models are trained for the 0.05 and 0.95 endpoints at each prediction setting. They use the evidence available at that prediction point and are fitted on the training tasks. Their outputs form the raw interval [ℓi , ui ]. Intervals are calibrated separately for Task Start, Call Start, In-call Update, and Task Update using the held-out calibration tasks of each fold. For target yi , the absolute interval residual is si = max {ℓi − yi , yi − ui , 0} .

(4)

For each prediction setting, the corresponding calibration quantile of {si } is added symmetrically to the raw endpoints. At runtime, the calibrated endpoints are bounded below by consumption already confirmed within the prediction scope. Calibration parameters remain fixed during inference.

C

Experimental Setup

C.1

Trace Collection

Benchmarks. From the 500 instances of SWE-bench Verified [Jimenez et al., 2024, Chowdhury et al., 2024], we take 144. The four smallest repositories, flask, seaborn, requests, and pylint, enter in full. pytest, xarray, astropy, scikit-learn, and matplotlib contribute 15 each, and sphinx, sympy, and django 16 each. Search-R1 [Jin et al., 2025a] contributes 32 questions from seven FlashRAG evaluation sets [Jin et al., 2025b], half single-hop from NQ [Kwiatkowski et al., 2019], TriviaQA [Joshi et al., 2017], and PopQA [Mallen et al., 2023], and half multi-hop from HotpotQA [Yang et al., 2018], 2WikiMultihopQA [Ho et al., 2020], MuSiQue [Trivedi et al., 2022], and Bamboogle [Press et al., 2023]. MMLU-Pro [Wang et al., 2024b] contributes 32 test questions with two or three per category, and LongBench-v2 [Bai et al., 2025] has 32 questions spread over its six domains and three length labels. Within each stratum, we take instances in the stable-hash order of their id, and the anchors are the first two per SWE-bench Verified repository, one for flask and three for django, and the first eight in each other benchmark. Task packs. A task is a statement and a seed workspace. For SWE-bench, the seed is the repository at the base commit. The agent may not edit tests and finishes by writing submission.patch, the diff of its working tree. For the other benchmarks, the seed holds README.md and an empty answer.md, the statement provides the question and its options, for LongBench-v2 together with the context document, and the agent writes its answer into answer.md. A Search-R1 statement also names a shell command that queries a local BM25 server [Robertson and Zaragoza, 2009, Jin et al., 2025a], the official Search-R1 retriever over the 21,015,324-passage wiki-18 corpus, and prints three 18

passages. Correctness is scored after collection, with the SWE-bench Verified harness for patches, letter match for MMLU-Pro and LongBench-v2, and alias match for Search-R1. Agent LLMs. Table 5 lists the models and their reasoning configurations. GPT-5.4 [OpenAI, 2026], Claude Opus 4.6 [Anthropic, 2026], and Gemini 3.1 Pro [Google DeepMind, 2026] use the lowest reasoning setting available through their respective APIs. DeepSeek-V4-Pro [DeepSeek-AI, 2026] and Qwen3.8-27B [Qwen Team, 2026] run with thinking enabled, while Llama-3.2-3B-Instruct [Meta AI, 2024] has no reasoning mode. Temperature is 0 where supported. The self-hosted models run on vLLM [Kwon et al., 2023] in bf16 on eight A100 GPUs with full context windows. Table 5: Agent LLMs used for trace collection, with access modes and reasoning configurations. Model

Vendor

Access

Reasoning configuration

GPT-5.4 Claude Opus 4.6 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.8-27B Llama-3.2-3B-Instruct

OpenAI Anthropic Google DeepSeek Alibaba Meta

API API API API Self-hosted Self-hosted

Low reasoning effort Adaptive thinking; low effort Low thinking level Thinking enabled Thinking enabled No reasoning mode

Harnesses. DeepSeek Harness [DeepSeek, 2026] runs at version 0.1.1-rc.2, commit b150a551, in its headless profile with 12 tools: read, write, edit, str_replace_editor, glob, grep, pwsh, job_list, job_output, job_kill, todo_write, and skill. Commands are executed without a sandbox or approval prompts. Context compaction, result pruning, generated session titles, web access, and subagents are turned off, so every provider request is an agent call. Each run starts a fresh session in a fresh copy of the seed workspace on Windows. OpenHands [Wang et al., 2025] runs at version 1.18.0 with the CodeAct [Wang et al., 2024a] agent and its execute_bash, str_replace_editor, and task-tracking tools, with browsing and the condenser disabled. It runs SWE-bench Verified in the official image of each instance and the other benchmarks in a plain Linux image and calls the same retrieval script. Both harnesses cap a run at 7,200 s, 500 calls, 4 retries per call, 65,536 output tokens per model request, and 30,000,000 tokens. A run ends when the agent submits, the model gives a final response without submitting, or a cap is reached. A logical call consists of its initial model request and any retry requests issued before tool feedback is received. The observer timestamps every model request and emits a checkpoint every 128 bytes of committed output. Token usage is taken from the provider response for each request and aggregated at the logical-call level. Repeated executions. Table 2 lists the runs per task. SWE-bench Verified tasks run twice per model, and the other tasks run four times. Anchors run eight times, or 16 for Search-R1. Trace statistics. Table 6 lists the median tokens and calls per run and the share of correct runs for each benchmark and model. For GPT-5.4 on SWE-bench Verified, the input makes up 99% of the tokens, the median run takes 129 s, and the median call emits two checkpoints. The code-repair execution in Appendix E.1 illustrates the four prediction points and the usage records along a complete run (Figure 9). Additional trajectories. The generalization experiments use independently released trajectories from LiveClawBench [Long et al., 2026], which records multiple agent LLMs across task domains under a shared agent framework. We retain runs with complete interaction records and per-call input and output token counts, and construct targets with the accounting convention of Section 3.1. For the reasoning-configuration experiment in Appendix D.4.5, we rerun 69 SWE-bench Verified tasks with GPT-5.4 reasoning effort set to off and compare them with the same tasks under the low setting. The API route reports no separate reasoning-token count. Measured against visible output, the median request bills 0.46 output tokens per byte under the low setting and 0.34 under off.

19

Table 6: Median tokens in thousands and calls among runs with complete usage, and correct runs as a percentage of all attempted runs, under each harness. DeepSeek Harness

C.2

OpenHands

Model

Tokens

Calls

Correct

Tokens

Calls

Correct

SWE-bench Verified GPT-5.4 Claude Opus 4.6 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.8-27B Llama-3.2-3B-Instruct

245.4 258.1 333.5 280.9 526.0 508.8

21 21 21 20 22 31

40.7 41.7 43.5 34.3 29.6 9.5

314.5 470.7 253.5 336.8 211.5 619.7

20 22 21 28 15 33

36.3 54.2 47.5 33.8 31.2 11.8

Search-R1 GPT-5.4 Claude Opus 4.6 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.8-27B Llama-3.2-3B-Instruct

38.1 58.3 55.8 94.8 60.3 136.4

8 9 10 12 10 16

76.3 59.8 68.8 65.2 54.0 27.2

37.5 58.5 48.3 41.2 84.0 98.7

8 10 8 9 9 15

75.0 69.6 69.6 70.5 58.9 28.6

MMLU-Pro GPT-5.4 Claude Opus 4.6 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.8-27B Llama-3.2-3B-Instruct

21.6 27.3 54.6 35.2 41.9 30.1

4 5 9 6 7 5

81.9 87.5 75.6 70.6 76.9 42.5

46.1 29.3 59.7 38.6 52.7 48.7

6 4 9 5 7 5

81.9 88.1 77.5 76.9 67.5 43.8

LongBench-v2 GPT-5.4 Claude Opus 4.6 Gemini 3.1 Pro DeepSeek-V4-Pro Qwen3.8-27B Llama-3.2-3B-Instruct

762.5 361.7 832.5 523.7 884.4 645.9

6 4 8 4 8 7

58.8 66.9 64.4 60.6 45.0 25.6

608.0 716.1 610.0 1,030.7 581.0 592.1

5 6 8 9 6 4

58.1 75.0 69.4 52.5 45.6 13.1

Data Splits

Tasks are assigned to five folds, balanced over benchmarks and agent LLMs, and all runs of a task remain in their fold. Each fold serves once as the test fold. The next fold calibrates intervals, the one after selects hyperparameters, and the remaining two train. We repeat the partitioning with three random seeds. For each seed, we pool out-of-fold test predictions and compute task-weighted metrics over all evaluated tasks. Tables 1 and 7–10 report means across the three seeds. C.3

Baselines

Output-length predictors. TRAIL, EGTP, and TIE are originally designed to predict the length of a single response. We adapt each method to the target associated with a prediction point while preserving its core representation and estimator. For Qwen3.8-27B and Llama-3.2-3B-Instruct, the internal-state methods read the last layer of the agent LLM itself. For the API models, they use Llama-3.2-3B-Instruct as a proxy encoder over the text visible at the prediction point: the task statement at Task Start, the statement and transcript of completed calls at Call Start and Task Update, and the statement and generated prefix at In-call Update. The window keeps the first 384 tokens of the statement and the last 1,536 of the remaining text. TRAIL classifies the last-token state into 8 log-spaced target-length bins and predicts the expected bin median. EGTP fits a ridge regression

20

of log target length on the entropy-weighted mean state and entropy statistics of the window. TIE encodes the window with bge-small-en-v1.5 [Xiao et al., 2024] and fits the location and scale of a log-t distribution with 3.5 degrees of freedom on the concatenated CLS, mean, and max poolings. Its point prediction is the median, and its interval spans the 0.05 and 0.95 quantiles. Agent usage estimators. Self-Prediction queries the agent LLM for the 5th, 50th, and 95th percentiles of token consumption for the same prediction target. The median serves as the point prediction, and the 5th and 95th percentiles define a nominal 90% prediction interval. We use these predictions directly, without fitting or calibrating them on held-out tasks. At Task Start, the agent first inspects the seed workspace with its full tool set and estimates the tokens of a complete run. At the other prediction points, we adapt the prompt to the observed execution record and corresponding prediction target. At Call Start, the agent estimates the complete consumption of the current call. At In-call Update, it estimates the same complete-call quantity using the observed generation prefix. At Task Update, it estimates the token consumption of the remaining execution. Appendix F provides the complete prompt template for all four points. Fitting. Each learned baseline uses TokenCast’s partition: it is fitted on each fold’s training tasks, tuned on the validation tasks, and calibrated on the calibration tasks. C.4

Metrics

For each prediction setting, let i index the evaluated prediction points, with target yi , point prediction ŷi , 90% prediction interval [ℓi , ui ], and normalized weight wi . We assign equal weight to each evaluated task, divide that weight equally among itsP evaluated runs, and divide each run’s weight equally among its evaluated prediction points. Thus, i wi = 1, and MAE =

X

wi |ŷi − yi |,

ȳ =

X

i

wi yi ,

i

WAPE =

MAE × 100%. ȳ

(5)

MAE and mean target consumption ȳ use the same evaluation samples and task–run–checkpoint weights. WAPE expresses the mean absolute error as a percentage of the mean actual target consumption. For a nominal 1 − α prediction interval, we use the interval score ISi = (ui − ℓi ) +

2 2 (ℓi − yi )1{yi < ℓi } + (yi − ui )1{yi > ui }, α α

(6)

with α = 0.1. Mean interval score is MIS =

X

wi ISi .

(7)

i

Empirical coverage and mean interval width are aggregated using the same weights. The Norm. Avg. column of Table 1 is the arithmetic mean over the four prediction points of a method’s MAE divided by the MAE of the corresponding history-median predictor. This predictor outputs the median target of the training tasks in the same fold for the same benchmark, agent LLM, and prediction point, and uses no task features or execution evidence.

D

Extended Evaluation

D.1

Run-to-Run Variation and Interval Reliability

Figure 6 summarizes repeated GPT-5.4 executions from 48 anchor tasks across four benchmarks and two harnesses. Search-R1 has 16 repetitions per task–harness pair, while the others have eight. Repeated executions of a task show wide consumption spread across all four benchmarks. Table 3 compares TokenCast’s calibrated intervals with the native intervals returned by Self-Prediction at the four prediction points. MIS combines interval width with a penalty for how far outcomes fall 21

3.0

SWE-bench

101

Search-R1

MMLU-Pro

LongBench-v2

Max / min consumption

2.5

Density

2.0 1.5 1.0 0.5

4.00× 3.00×

2.00× 1.50×

0.0 0.5

0.8

1

1.2

2

1.00×

2.5

104

105

106

107

Median token consumption

Run tokens / task median

(a) Run-to-run variation

(b) Consumption spread

Figure 6: Run-to-run variation in token consumption across repeated executions of the same task. The panels show the distribution across repeated runs and the within-task consumption spread as a function of task consumption, over 48 GPT-5.4 anchor tasks. outside the interval. Lower values indicate better interval forecasts. TokenCast’s wider intervals reduce the missed-outcome penalty enough to improve MIS. This comparison includes TokenCast’s held-out calibration and Self-Prediction’s native, uncalibrated intervals.

22

D.2

Results Across Benchmarks and Agent LLMs

Tables 7 to 10 report the MAE for each benchmark and agent LLM on the same runs, with the history median as the normalization reference. MAE is reported in tokens, with k denoting thousands. Norm. Avg. follows the normalization in Appendix C.4. Table 7: MAE on SWE-bench Verified. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.

D.3

Prediction point

TokenCast

TRAIL

EGTP

TIE

Self-Pred.

History median

GPT-5.4 Task Start Call Start In-call Update Task Update Norm. Avg.

144.0k 64.5 38.9 80.0k 0.69

165.3k 71.1 78.9 115.0k 0.94

160.9k 70.3 74.6 117.0k 0.92

157.0k 69.5 77.4 119.0k 0.92

152.0k 70.9 80.2 126.0k 0.94

179.0k 74.2 80.3 130.0k 1.00

Claude Opus 4.6 Task Start Call Start In-call Update Task Update Norm. Avg.

150.7k 64.1 55.0 88.7k 0.77

157.5k 69.4 80.2 112.7k 0.92

153.8k 69.9 71.7 122.4k 0.91

146.8k 66.8 74.4 116.4k 0.88

138.9k 68.8 78.1 122.0k 0.90

171.0k 73.5 83.4 133.0k 1.00

Gemini 3.1 Pro Task Start Call Start In-call Update Task Update Norm. Avg.

157.8k 68.3 41.0 69.6k 0.62

173.5k 74.7 84.3 126.3k 0.86

170.2k 71.6 79.4 121.4k 0.83

176.4k 68.4 77.2 118.6k 0.82

163.6k 79.8 87.2 132.2k 0.88

193.0k 88.5 96.2 151.0k 1.00

DeepSeek-V4-Pro Task Start Call Start In-call Update Task Update Norm. Avg.

212.9k 80.0 65.5 106.2k 0.72

204.9k 89.0 96.1 143.4k 0.86

211.5k 84.6 94.9 155.6k 0.87

198.4k 85.9 86.3 148.1k 0.83

194.5k 91.3 104.8 158.4k 0.89

232.0k 101.5 112.8 177.0k 1.00

Qwen3.8-27B Task Start Call Start In-call Update Task Update Norm. Avg.

192.0k 87.0 81.9 124.0k 0.65

232.0k 109.0 120.0 176.0k 0.86

209.0k 102.9 116.0 171.0k 0.81

215.0k 99.0 111.4 164.0k 0.79

221.0k 105.5 108.9 173.0k 0.82

284.0k 119.0 147.0 198.0k 1.00

Llama-3.2-3B-Instruct Task Start 272.3k Call Start 133.8 In-call Update 100.3 Task Update 152.2k Norm. Avg. 0.79

277.6k 140.7 141.8 205.3k 0.93

269.7k 134.6 137.7 213.1k 0.91

258.2k 140.0 126.4 188.2k 0.87

264.8k 141.6 150.2 213.7k 0.94

302.0k 147.0 158.0 218.0k 1.00

Error Relative to Consumption

Table 11 reports WAPE and mean target consumption for GPT-5.4 and Qwen3.8-27B across four benchmarks and four prediction points. For each benchmark, agent LLM, and prediction point, we

23

Table 8: MAE on Search-R1. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values. Prediction point

TokenCast

TRAIL

EGTP

TIE

Self-Pred.

History median

GPT-5.4 Task Start Call Start In-call Update Task Update Norm. Avg.

34.2k 34.5 31.3 17.6k 0.75

37.8k 37.3 39.4 23.1k 0.89

37.6k 36.5 38.5 23.7k 0.88

36.4k 34.0 36.8 23.3k 0.85

36.9k 39.8 39.9 24.2k 0.91

44.9k 44.0 42.0 25.0k 1.00

Claude Opus 4.6 Task Start Call Start In-call Update Task Update Norm. Avg.

31.1k 30.4 22.8 12.9k 0.67

33.9k 33.6 33.4 20.8k 0.86

32.9k 32.0 31.6 18.7k 0.81

32.2k 30.4 29.7 19.5k 0.79

29.8k 34.9 34.0 21.4k 0.85

39.8k 40.2 38.5 23.6k 1.00

Gemini 3.1 Pro Task Start Call Start In-call Update Task Update Norm. Avg.

38.0k 40.0 25.5 16.1k 0.61

42.3k 44.7 46.8 27.8k 0.84

41.0k 39.8 44.2 24.6k 0.78

40.7k 42.3 42.6 25.2k 0.79

37.1k 47.0 47.5 28.6k 0.84

49.8k 54.5 57.0 31.5k 1.00

DeepSeek-V4-Pro Task Start Call Start In-call Update Task Update Norm. Avg.

46.5k 50.7 35.1 19.9k 0.64

50.0k 57.3 60.8 33.7k 0.86

47.5k 54.0 56.2 31.1k 0.81

44.2k 50.6 50.4 29.8k 0.75

42.9k 60.6 62.3 34.5k 0.86

58.5k 68.0 71.5 37.2k 1.00

Qwen3.8-27B Task Start Call Start In-call Update Task Update Norm. Avg.

39.0k 56.7 57.0 24.0k 0.59

53.5k 65.9 84.6 39.5k 0.83

49.8k 71.4 80.7 34.0k 0.79

47.6k 67.7 75.5 35.4k 0.77

51.0k 76.1 83.2 38.4k 0.84

72.9k 91.0 93.0 41.0k 1.00

Llama-3.2-3B-Instruct Task Start 59.0k Call Start 79.0 In-call Update 64.0 Task Update 31.0k Norm. Avg. 0.73

60.0k 80.9 99.4 43.8k 0.89

57.9k 70.4 91.2 40.2k 0.81

53.6k 75.4 84.7 38.6k 0.79

52.4k 84.6 95.6 45.9k 0.87

68.0k 95.0 108.0 49.0k 1.00

24

Table 9: MAE on MMLU-Pro. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values. Prediction point

TokenCast

TRAIL

EGTP

TIE

Self-Pred.

History median

GPT-5.4 Task Start Call Start In-call Update Task Update Norm. Avg.

8.0k 14.8 9.2 4.3k 0.65

8.8k 16.7 19.6 6.2k 0.87

8.5k 15.5 18.3 6.5k 0.84

8.2k 15.8 16.7 6.0k 0.80

7.4k 17.6 17.4 5.8k 0.80

9.7k 19.6 24.3 6.9k 1.00

Claude Opus 4.6 Task Start Call Start In-call Update Task Update Norm. Avg.

7.5k 14.9 11.5 4.4k 0.72

8.0k 14.8 15.5 6.0k 0.84

7.7k 15.6 16.8 6.1k 0.86

7.3k 14.3 16.0 5.8k 0.81

6.9k 16.1 18.4 5.6k 0.85

8.9k 18.3 22.0 6.3k 1.00

Gemini 3.1 Pro Task Start Call Start In-call Update Task Update Norm. Avg.

8.6k 18.0 10.5 4.7k 0.61

10.1k 18.3 24.7 7.2k 0.85

9.6k 20.1 20.6 6.8k 0.81

9.3k 19.4 21.8 6.4k 0.79

8.9k 21.3 23.6 7.5k 0.86

11.5k 23.9 28.7 8.0k 1.00

DeepSeek-V4-Pro Task Start Call Start In-call Update Task Update Norm. Avg.

11.6k 24.4 19.5 6.9k 0.73

12.3k 26.2 31.8 8.5k 0.88

11.9k 24.8 28.7 9.2k 0.86

11.0k 23.5 25.4 8.9k 0.80

10.6k 26.9 30.4 9.4k 0.87

13.7k 29.6 35.9 9.8k 1.00

Qwen3.8-27B Task Start Call Start In-call Update Task Update Norm. Avg.

13.0k 26.5 21.1 8.4k 0.69

14.3k 30.3 37.7 11.2k 0.89

12.5k 28.6 35.5 10.3k 0.82

13.7k 27.9 34.3 10.7k 0.84

13.3k 31.1 32.8 11.5k 0.86

16.2k 35.1 42.4 12.0k 1.00

Llama-3.2-3B-Instruct Task Start 16.4k Call Start 33.5 In-call Update 30.7 Task Update 10.8k Norm. Avg. 0.75

17.5k 36.5 46.0 13.2k 0.90

16.9k 33.2 43.7 13.8k 0.87

15.4k 33.8 42.3 13.5k 0.84

15.8k 37.9 47.3 12.9k 0.89

19.6k 41.8 50.2 14.4k 1.00

25

Table 10: MAE on LongBench-v2. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values. Prediction point

TokenCast

TRAIL

EGTP

TIE

Self-Pred.

History median

GPT-5.4 Task Start Call Start In-call Update Task Update Norm. Avg.

32.2k 53.2 35.5 13.9k 0.57

37.1k 60.6 73.5 26.9k 0.82

36.0k 56.9 68.8 24.2k 0.77

35.1k 53.7 62.7 25.1k 0.74

32.7k 61.2 70.1 28.2k 0.80

44.0k 74.0 91.0 33.0k 1.00

Claude Opus 4.6 Task Start Call Start In-call Update Task Update Norm. Avg.

28.7k 48.1 41.3 14.9k 0.61

33.4k 57.2 69.5 23.8k 0.84

32.0k 50.0 57.9 22.4k 0.76

30.7k 54.5 61.3 21.5k 0.77

29.4k 56.0 65.4 25.1k 0.81

40.5k 67.5 82.0 28.5k 1.00

Gemini 3.1 Pro Task Start Call Start In-call Update Task Update Norm. Avg.

39.9k 60.6 43.7 15.8k 0.59

43.2k 72.3 88.4 31.2k 0.85

41.9k 66.8 80.6 27.2k 0.78

38.2k 66.0 75.1 28.6k 0.76

38.9k 75.0 85.8 33.0k 0.84

50.0k 85.0 103.0 38.0k 1.00

DeepSeek-V4-Pro Task Start Call Start In-call Update Task Update Norm. Avg.

52.3k 84.5 77.2 26.1k 0.69

55.0k 96.0 115.6 35.7k 0.85

54.9k 89.6 110.2 38.7k 0.84

51.2k 85.6 101.7 37.4k 0.80

49.6k 97.9 118.4 42.2k 0.87

61.5k 108.0 132.0 47.5k 1.00

Qwen3.8-27B Task Start Call Start In-call Update Task Update Norm. Avg.

50.8k 94.4 95.0 26.9k 0.65

60.7k 107.8 136.4 47.6k 0.87

55.9k 94.1 128.9 42.7k 0.79

54.6k 99.7 121.2 44.5k 0.80

52.6k 110.8 118.6 49.0k 0.83

70.0k 124.0 151.0 56.0k 1.00

Llama-3.2-3B-Instruct Task Start 70.6k Call Start 116.8 In-call Update 111.9 Task Update 38.6k Norm. Avg. 0.74

73.6k 128.0 156.2 56.8k 0.90

69.2k 116.2 141.9 55.4k 0.84

66.3k 121.8 145.2 52.1k 0.83

71.4k 133.7 162.7 60.2k 0.93

82.0k 142.0 171.0 64.0k 1.00

26

use the same evaluation weights as MAE: P wi |b y i − yi | MAE = 100 WAPE(%) = 100 iP , ȳ i w i yi

ȳ =

X

wi yi ,

i

X

wi = 1.

(8)

i

Here, yi is the target consumption: total task consumption T at Task Start, full-call consumption Ck at Call Start and In-call Update, and remaining task consumption Rk at Task Update. The weights follow Appendix C.4. All methods within a row share the same target mean and evaluation weights. A WAPE of 15% means that the weighted MAE is 15% of the weighted mean target. Table 11: WAPE (%) at four prediction points. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values. Agent LLM

Prediction point Mean target (k) TokenCast TRAIL EGTP

TIE Self-Pred.

SWE-bench Verified Task Start Call Start In-call Update Task Update

475.6 19.6 20.1 286.3

30.3 0.329 0.194 27.9

34.8 33.8 33.0 0.363 0.359 0.355 0.393 0.371 0.385 40.2 40.9 41.6

32.0 0.362 0.399 44.0

Task Start Call Start Qwen3.8-27B In-call Update Task Update

611.5 25.4 25.9 366.3

31.4 0.343 0.316 33.9

37.9 34.2 35.2 0.429 0.405 0.390 0.463 0.448 0.430 48.0 46.7 44.8

36.1 0.415 0.420 47.2

GPT-5.4

Search-R1 Task Start Call Start In-call Update Task Update

53.5 6.0 6.1 31.6

63.9 0.575 0.513 55.7

70.7 70.3 68.0 0.622 0.608 0.567 0.646 0.631 0.603 73.1 75.0 73.7

69.0 0.663 0.654 76.6

Task Start Call Start Qwen3.8-27B In-call Update Task Update

98.1 9.5 9.8 57.6

39.8 0.597 0.582 41.7

54.5 50.8 48.5 0.694 0.752 0.713 0.863 0.823 0.770 68.6 59.0 61.5

52.0 0.801 0.849 66.7

GPT-5.4

MMLU-Pro Task Start Call Start In-call Update Task Update

50.6 8.2 8.4 27.9

15.8 0.180 0.110 15.4

17.4 16.8 16.2 0.204 0.189 0.193 0.233 0.218 0.199 22.2 23.3 21.5

14.6 0.215 0.207 20.8

Task Start Call Start Qwen3.8-27B In-call Update Task Update

63.8 10.0 10.3 35.2

20.4 0.265 0.205 23.9

22.4 19.6 21.5 0.303 0.286 0.279 0.366 0.345 0.333 31.8 29.3 30.4

20.8 0.311 0.318 32.7

GPT-5.4

LongBench-v2 Task Start Call Start In-call Update Task Update

771.2 135.7 136.6 391.2

4.2 0.039 0.026 3.6

4.8 4.7 4.6 0.045 0.042 0.040 0.054 0.050 0.046 6.4 6.9 6.2

4.2 0.045 0.051 7.2

Task Start Call Start Qwen3.8-27B In-call Update Task Update

856.5 112.1 112.6 434.2

5.9 0.084 0.084 6.2

7.1 6.5 6.4 0.096 0.084 0.089 0.121 0.114 0.108 11.0 9.8 10.2

6.1 0.099 0.105 11.3

GPT-5.4

27

D.4 D.4.1

Generalization Unseen Task Types

We use the publicly released LiveClawBench trajectories [Long et al., 2026], which contain tasks from 10 application domains executed under a shared agent framework. For each agent LLM, we perform leave-one-domain-out evaluation, treating one domain as the unseen task type and training on the remaining domains. Zero-shot transfer uses no trajectories from the held-out domain. For adaptation, we add k held-out-domain tasks to the training set and evaluate on the remaining held-out tasks. The target-only baseline is trained on the same k tasks without source-domain data. All runs of a task remain in one partition. Relative to Self-Prediction, zero-shot MAE is 1.31 at Call Start and 1.47 at Task Update. With 10 target tasks, the ratios fall to 0.86 and 0.91; with 20, they reach 0.82 and 0.85. The trajectory-selection criteria are in Appendix C.1. D.4.2

Unseen Agent LLMs

We hold out each agent LLM in turn and train on trajectories from the remaining models under a shared task-collection and agent framework. Adaptation uses k tasks executed by the held-out model, with task partitions shared across models and adaptation tasks disjoint from the test tasks. At Call Start, zero-shot MAE is 1.02 times the target history median; this ratio falls to 0.96 with 10 target-model tasks and 0.95 with 20. With 3–10 tasks, the transferred predictor achieves lower MAE than training on the same target-model tasks alone. Figure 7 also reports an auxiliary Task Update evaluation of remaining output tokens, a different target from total remaining consumption. Its normalized MAE ranges from 1.01 to 1.07 with up to 20 target-model tasks and reaches 0.98 when all target-model tasks are available. The target-only predictor starts at 1.22 with three tasks and approaches the history median as additional target-model trajectories are provided. History median

Target only

Transfer + adapt 1.3

MAE / history median

MAE / history median

1.3 1.2 1.1 1.0 0.9 0.8 0.7

0

3

5

10

20

1.2 1.1 1.0 0.9 0.8 0.7

All

Tasks of the target used

0

3

5

10

20

All

Tasks of the target used

(a) Call Start

(b) Task Update, output

Figure 7: Transfer to unseen agent LLMs with increasing amounts of target-model data. The panels report Call Start and remaining output-token prediction at Task Update. MAE is normalized by the target history median; lower is better. D.4.3

Cross-Harness Transfer

Table 12 reports TokenCast’s MAE relative to the history median under each test harness. It also reports transfer from DeepSeek Harness training runs to OpenHands runs of held-out tasks. D.4.4

Length Extrapolation

We sort the SWE-bench Verified tasks by their median number of calls into five bands, train on the four shorter bands, and evaluate on the longest band. TokenCast achieves lower MAE than Self-Prediction at all four prediction points on these longer tasks, with reductions of 12.4% at Task

28

Table 12: TokenCast MAE divided by the history-median MAE on the test harness. Lower is better. The two OpenHands rows use the same test runs. Training harness

Task Call In-call Task Start Start Update Update

Test harness

DeepSeek Harness DeepSeek Harness 0.90 0.71 OpenHands OpenHands 0.83 0.77 DeepSeek Harness OpenHands 1.04 0.88

0.64 0.68 0.79

0.82 0.75 0.96

Start, 15.7% at Call Start, 44.9% at In-call Update, and 26.8% at Task Update. Table 13 additionally compares this length-based split with random splits of the same fold sizes. Relative to the random splits, MAE differs by +7.8% at Task Start, +1.9% at Call Start, −2.6% at In-call Update, and +7.1% at Task Update, so the longer test band raises MAE by at most 7.8%. Table 13: Length extrapolation to tasks longer than those observed during training. MAE changes are reported relative to Self-Prediction and matched random splits.

D.4.5

Prediction point

vs. Self-Prediction ↓

vs. random split ↓

Task Start Call Start In-call Update Task Update

−12.4% −15.7% −44.9% −26.8%

+7.8% +1.9% −2.6% +7.1%

Reasoning Configuration Shift

We evaluate a shift in GPT-5.4’s reasoning configuration by rerunning the same 69 SWE-bench Verified tasks with reasoning effort set to off and comparing them with the corresponding runs under the low setting. Median run consumption decreases from 273k to 238k tokens. When trained and evaluated within the off configuration, TokenCast reduces MAE relative to Self-Prediction by 20.7% at Call Start, 50.8% at In-call Update, and 32.2% at Task Update (Table 14). We also transfer the predictor trained under the low configuration directly to the off configuration and adapt it using 20 target-configuration tasks. After adaptation, MAE decreases to 37.9 at Call Start, 23.6 at In-call Update, and 76.0k at Task Update. Table 14: MAE under a GPT-5.4 reasoning-configuration shift on SWE-bench Verified at the three online prediction points. Results include evaluation within the off configuration, direct transfer from low to off, and adaptation with 20 target-configuration tasks. Low → Off

Off

Prediction point TokenCast Self-Pred. Direct 20 tasks Call Start In-call Update Task Update

D.5

37.1 22.5 82.0k

46.8 45.7 121.0k

50.7 33.0 82.0k

37.9 23.6 76.0k

Online Update Frequency

We measure the prediction overhead of refreshing the Task Update forecast at different intervals on SWE-bench Verified with GPT-5.4. Starting from the same initial forecast, the 1-, 3-, and 5-call settings refresh the prediction after every 1, 3, or 5 completed calls, respectively, and End makes only the initial forecast. At skipped checkpoints, the latest forecast of final task consumption is retained, the tokens confirmed so far are subtracted, and the remaining estimate is clipped at zero.

29

Table 15 reports the mean number of forecasts and cumulative prediction time per run. Counts include the initial Task Start forecast, and times are reported as mean ± standard deviation across runs. Updating every call makes 19.7 forecasts and takes 32.8 ± 30.5 ms per run. The three-call and five-call schedules reduce these to 6.9 forecasts and 11.9±10.4 ms, and 4.3 forecasts and 7.7±6.4 ms, respectively. End makes only the initial forecast and takes 2.1 ± 0.8 ms. Every-call updating therefore remains a small fraction of the 129 s median run time reported in Appendix C.1. Table 15: Prediction counts and cumulative overhead at online update intervals on SWE-bench Verified. Predictions/run includes the initial Task Start forecast and later Task Update refreshes. Update interval

Predictions/run ↓

Time/run (ms) ↓

19.7 6.9 4.3 1.0

32.8 ± 30.5 11.9 ± 10.4 7.7 ± 6.4 2.1 ± 0.8

1 call 3 calls 5 calls End

D.6

Ablation Study

Base predictor comparison. Table 16 compares base predictors in the same pipeline. Ridge regression and KNN reach normalized average MAEs of 0.93 and 0.97. Random Forest reaches 0.81, and the MLP reaches 0.78 with 87.2% interval coverage. The three gradient-boosting methods reach 0.75 or lower, and LightGBM gives the lowest error, 0.69, and the highest coverage, 90.6%. The base-predictor comparison reports 0.8 ms for LightGBM inference, compared with 3.9 ms for XGBoost and 5.2 ms for CatBoost. Under every-call updating, the complete pipeline, including feature extraction and composition, makes 19.7 task-level forecasts and takes 32.8 ms per SWE-bench Verified run on average. We use LightGBM as the base predictor. Table 16: Base predictor comparison on SWE-bench Verified (GPT-5.4). Norm. Avg. is the normalized MAE averaged over the four prediction points. 90% Cov. is the empirical coverage of the 90% prediction interval. Best and second-best results in each column are in bold and underlined, respectively. Coverage is ranked high to low and other metrics low to high. Norm. Avg. ↓

90% Cov. (%)

Model time (ms)

Ridge Regression KNN Random Forest XGBoost CatBoost MLP

0.93 0.97 0.81 0.73 0.75 0.78

82.4 79.6 86.8 89.3 88.7 87.2

0.2 14.3 2.7 3.9 5.2 6.1

LightGBM

0.69

90.6

0.8

Predictor

Prediction strategy. At Task Start, no segment has finished, so the prefix–suffix decomposition has no observed anchor. The compositional path has a normalized MAE of 0.87, compared with 0.82 for direct regression. At Task Update, observed segment boundaries provide information about the input baseline and context growth. The compositional path improves to 0.66, below direct regression’s 0.72. Averaging the two paths reduces MAE to 0.81 at Task Start and 0.64 at Task Update. The correction model uses the direct and compositional forecasts together with the predicted boundary variables, reducing these values to 0.80 and 0.62, with 90.6% coverage (Table 17). It is trained on out-of-fold outputs of the complete pipeline. Segment representation ablation. Regressing directly on raw execution features raises the normalized average MAE to 0.85, with task-level and call-level MAE of 0.88 and 0.82. Retaining the triple while removing composition gives 0.78. This intervention also changes call-level MAE from 0.68 to 30

Table 17: Effect of forecasting strategy on SWE-bench Verified (GPT-5.4). Task Start and Task Update columns report the normalized MAE at these two task-level prediction settings. Best and second-best results in each column are in bold and underlined, respectively. Coverage is ranked high to low and other metrics low to high. Strategy

Task Start ↓

Task Update ↓

90% Cov. (%)

0.82 0.87 0.81 0.80

0.72 0.66 0.64 0.62

88.1 89.3 89.8 90.6

Direct only Compositional only Direct–compositional average Full correction pipeline

0.74, so the ablation affects more than the task-level composition readout. Removing g raises the normalized average MAE to 0.76 and task-level MAE from 0.71 to 0.79. Removing b raises it to 0.73 (Table 18). The larger effect of removing g is consistent with its role in the cross-term nB gA ; the ablations also change call-level predictions and do not isolate this term. Table 18: Ablation of the segment representation on SWE-bench Verified (GPT-5.4). Task-level and Call-level columns report normalized MAE averaged over the two prediction points within each level. Best and second-best results in each column are in bold and underlined, respectively. Variant Raw features No composition Drop g Drop b Full

Norm. Avg. ↓

Task-level ↓

Call-level ↓

0.85 0.78 0.76 0.73 0.69

0.88 0.81 0.79 0.75 0.71

0.82 0.74 0.72 0.70 0.68

Component ablation. The suffix model takes predictions from the prefix model as input. A shift in this input between training and inference can therefore propagate errors through the cascade. Cross-fitting reduces the shift, while the cost-weighted loss scales errors in the coupled variables g and n by their downstream impact. Removing either component raises the normalized average MAE by 0.03–0.04. Removing both raises MAE to 0.78 and lowers coverage to 87.4% (Table 19). Without the correction model, MAE rises to 0.74 and coverage falls to 89.0%. Removing the boundary variables raises MAE to 0.72, consistent with their role in conditioning the suffix model. Table 19: Component ablation on SWE-bench Verified (GPT-5.4). Each row removes one component from the full pipeline. Best and second-best results in each column are in bold and underlined, respectively. Coverage is ranked high to low and other metrics low to high. Configuration

Norm. Avg. ↓

90% Cov. (%)

0.69 0.73 0.72 0.74 0.72 0.78

90.6 88.9 89.4 89.0 90.1 87.4

Full − Cross-fitting − Cost weighting − Correction − Boundary vars − Cross-fit. & cost-wt.

D.7

Budget Control

Replay and stopping rules. We replay 288 GPT-5.4 runs from 144 SWE-bench Verified tasks, with two recorded runs per task. Seven budgets are defined by the 0.3 to 0.9 quantiles of recorded run

31

consumption. At each Task Update checkpoint, the controller adds confirmed consumption to a selected quantile of predicted remaining consumption and stops the run when the sum exceeds the budget. A replay is trace-complete if it reaches its recorded terminal state. For each method, budget, and test fold, the controller selects the 0.05, 0.5, or 0.95 forecast quantile on runs outside the test fold, choosing the one using the fewest tokens while reaching at least the fixed-budget trace-completion rate. Results are pooled over out-of-fold test runs and averaged over three split seeds. Cost accounting. A controller-stopped run is charged the consumption recorded at its stopping checkpoint. A run that finishes before the fixed cap is charged its recorded total; otherwise it is stopped at the cap, with wall time scaled by the recorded token progress. Prediction processing and time are included in the replay accounting. For encoder-based baselines, prediction tokens count locally processed text; Self-Prediction’s count includes additional provider-recorded LLM usage. Savings are relative to the fixed-budget strategy at the same budget. Table 20 lists the complete results, Figure 8 shows the budget-wise response, and Table 21 compares prediction overhead. TRAIL

EGTP

75

TIE

Completion change (pp)

Net token saving (%)

TokenCast

50 25 0 −25 −50 −75 200

300

400

500

600

Self-Pred. 0 −5 −10 −15 −20 −25 −30 200

Budget (k tokens)

300

400

500

600

Budget (k tokens)

(a) Net token saving

(b) Completion-rate change

Figure 8: Trace completion and token consumption across seven budgets in the offline replay.

32

Table 20: Budget control on 288 GPT-5.4 runs from 144 SWE-bench Verified tasks. Tokens are in thousands per run, wall time is in seconds per run, and trace completion and savings are in percent. Only TokenCast matches the fixed-budget trace completion at every budget. Budget (k tokens)

Method

Trace-complete Execution Prediction Total Time (%) (k/run) (k/run) (k/run) (s/run)

Saving (%)

174

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

30.2 30.2 29.5 13.5 27.4 19.4

154.7 101.1 91.0 25.3 59.8 70.0

0.0 0.0 12.5 3.6 8.0 92.0

154.7 101.1 103.5 28.9 67.8 162.0

113.7 76.5 73.0 22.6 45.6 303.5

– 34.6 33.1 81.3 56.2 -4.7

203

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

39.9 39.9 39.2 17.0 34.0 24.7

173.5 121.2 112.6 31.8 73.5 84.5

0.0 0.0 14.5 4.8 9.7 109.2

173.5 121.2 127.1 36.6 83.2 193.7

125.3 90.0 87.6 27.6 54.9 355.6

– 30.1 26.7 78.9 52.0 -11.6

234

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

50.0 50.0 49.7 24.7 43.4 29.5

190.7 141.2 132.3 42.7 92.3 100.2

0.0 0.0 16.7 5.6 11.4 127.5

190.7 141.2 149.0 48.3 103.7 227.7

135.8 103.1 101.0 34.8 67.1 413.3

– 26.0 21.9 74.7 45.6 -19.4

277

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

59.7 59.7 59.7 30.9 52.8 38.2

210.4 165.0 157.3 53.4 111.5 120.4

0.0 0.0 19.9 6.9 13.8 151.7

210.4 165.0 177.2 60.3 125.3 272.1

148.5 118.1 118.7 42.1 80.1 481.7

– 21.6 15.8 71.3 40.4 -29.3

357

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

69.8 69.8 69.4 39.9 63.9 52.1

237.4 197.2 193.9 76.0 141.2 154.8

0.0 0.0 23.8 9.5 17.1 190.5

237.4 197.2 217.7 85.5 158.3 345.3

165.1 138.3 144.2 58.9 98.8 593.9

– 16.9 8.3 64.0 33.3 -45.5

440

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

79.9 79.9 79.2 51.0 73.3 64.9

258.2 226.5 222.3 100.2 171.5 184.3

0.0 0.0 26.9 12.7 20.8 222.5

258.2 226.5 249.2 112.9 192.3 406.8

177.2 156.4 163.1 77.3 119.6 688.3

– 12.3 3.5 56.3 25.5 -57.6

604

Fixed budget TokenCast TRAIL EGTP TIE Self-Pred.

89.9 89.9 89.9 66.3 86.1 79.9

280.7 258.8 254.1 140.3 218.6 224.6

0.0 0.0 30.4 16.6 25.6 268.3

280.7 258.8 284.5 156.9 244.2 492.9

191.3 176.4 184.2 104.3 150.4 820.8

– 7.8 -1.4 44.1 13.0 -75.6

33

Table 21: Budget-equal-weighted replay means. Trace completion is shown alongside overhead because lower processing totals can result from stopping more runs. Encoder-based baselines count locally processed input tokens; Self-Prediction counts additional LLM usage. These token counts therefore represent different resources. Method

Trace-complete (%)

Prediction tokens

Total tokens

Total time

59.9 59.5 34.8 54.4 44.1

0.0k 20.7k 8.5k 15.2k 166.0k

173.0k 186.9k 75.6k 139.3k 300.1k

122.7 s 124.5 s 52.5 s 88.1 s 522.4 s

TokenCast TRAIL EGTP TIE Self-Prediction

E

Case Studies

Token consumption reflects both the actions needed to finish a task and the context carried into each call. Two executions connect these quantities to observable events: verification after an edit and artifact preparation with a long context. Task Update estimates are displayed as final totals, bk , where Sk is confirmed consumption. Call-level estimates retain the full current-call Tbk = Sk + R target Ck . Dashed lines mark retrospective targets. Grouped bars show Call Start forecasts from TokenCast and Self-Prediction for full current-call consumption. E.1

Code Repair

GPT-5.4 repairs a swap_dims() mutation issue in pydata__xarray-6938 with DeepSeek Harness. The run contains 21 calls and 39 checkpoints, consumes 243,371 tokens, and lasts 130 s. The following timeline, detail plots, and table show how forecasts change during source editing, verification failure, retry, and completion. Execution timeline. Figure 9 plots confirmed consumption over wall time. Task Start is at 0 s and the first Call Start 1 ms later, when the first request is assembled. The In-call Update shown is checkpoint 5 of call 12 at 65.3 s, with 657 bytes committed and 95,591 tokens confirmed. The Task Update after call 12 is at 68.7 s, with 109,011 tokens confirmed and 134,360 still to come.

Confirmed tokens (k)

250 200 150 In-call Update

100

Task Update

50 0

Task Start Call Start

0

20

21 calls, 243,371 tokens

40

60

80

100

120

Wall time (s)

Figure 9: Confirmed tokens over wall time for one run. The shaded span is call 12. Generated prefix. Figure 10 and Table 22 report forecasts at selected prediction points. Call 12 performs the source edit and consumes 13,420 tokens. At 795 generated bytes, the prefix exposes the original source span. At 1,450 bytes, it reveals the replacement text. Across these checkpoints, the 90% interval contracts from 911 tokens at Call Start to 65 tokens at 1,450 bytes, while the forecast remains close to the recorded 13,420-token call cost. Verification feedback. At call 14, reproduction encounters the removed NumPy unicode_alias, and the predicted task total rises to 281,746 tokens. After a compatibility workaround allows the 34

Final Spent

Call cost (k tokens)

Total (k tokens)

Forecast 90% interval

450 300 150 0

Forecast 90% interval

Final

14.0

13.5

13.0 0

5

10

15

20

0

700

Completed calls

1400

Visible prefix (bytes)

Figure 10: Forecast updates during code repair. Left: task-level forecasts following verification outcomes. Right: current-call forecasts during edit call 12 as generation becomes visible. Table 22: Forecasts at selected points in the code-repair execution. Token quantities are in thousands. Recorded targets are shown retrospectively. Prediction point Available evidence

Target

Task Start Call Start, 12 In-call, 12 In-call, 12 Update, k = 14 Update, k = 15

T C12 C12 C12 T T

Issue and configuration Edit request assembled 795 bytes, old source span visible 1,450 bytes, replacement text visible NumPy alias prevents reproduction Compatibility workaround, assertions pass

Forecast [90% interval] Recorded 194.638 [104.782, 333.912] 13.368 [13.001, 13.912] 13.383 [13.201, 13.694] 13.431 [13.407, 13.472] 281.746 [191.304, 368.259] 252.381 [211.623, 298.576]

243.371 13.420 13.420 13.420 243.371 243.371

non-mutation assertions to pass at call 15, the forecast becomes 252,381 tokens, close to the recorded total of 243,371. The execution still consumes another 92,329 tokens while inspecting the diff, preparing artifacts, and submitting the response, accounting for 37.9% of the final total. E.2

Long-Context QA

Forecast 90% interval

1500

Final Spent

Call cost (k tokens)

Total (k tokens)

File operations. Qwen3.8-27B answers a LongBench question about OpenLRM and Instant3D using OpenHands. During call 1, it generates option C and its justification, then encounters a file-creation error because answer.md already exists. Inspecting and replacing the placeholder, preparing the patch, and checking the artifacts extend the execution to six calls. Successful replacement at call 3 is followed by another 395,353 tokens of consumption. Figure 11 and Table 23 show the corresponding task-level and call-level forecasts.

1000

500

0

0

1

2

3

4

5

TokenCast

120 60 0

6

Completed calls

Self-Prediction

180

1

2

3

4

5

6

Call

Figure 11: Repeated input makes additional calls expensive. Left: task updates after a failed write and artifact preparation. Right: full-call Call Start forecasts from TokenCast and Self-Prediction. Repeated input. The first call processes 129,714 input tokens. Subsequent calls process 130,181– 35

Table 23: Long-context stages. All token quantities are in thousands. The same answer is carried through the subsequent file operations. Point

Sk

Newly available evidence

Task Start Long-context question and output requirements Update, k = 1 Answer generated, file creation fails Update, k = 2 Existing answer placeholder inspected Update, k = 3 Answer file successfully replaced Update, k = 5 Answer and patch read back Termination Finish action recorded

Forecast T [90% interval]

0.000 526.384 [264.731, 919.648] 131.059 834.719 [525.193, 1204.762] 261.368 908.362 [643.881, 1268.997] 392.112 762.541 [652.967, 1041.863] 655.063 792.813 [765.208, 839.426] 787.465 —

132,209 input tokens and emit 128–350 output tokens each. Input contributes 785,146 of 787,465 tokens (99.7%). For the recorded execution, the segment decomposition gives T = nL1 + b = 6 × 129,714 + 9,181 = 787,465. The repeated-input term contributes 778,284 tokens. Each additional call therefore carries roughly 130,000 input tokens, so the remaining call count and the input boundary jointly determine the task cost. The failed write at call 1 signals extra calls at this context length, and the Task Update forecast rises from 526k at Task Start to 835k tokens.

F

Self-Prediction Prompt

The following prompt template adapts the zero-shot agent self-prediction protocol of Bai et al. [2026] to the four prediction points used in our evaluation. It retains environment inspection, workload analysis, and phase-wise input/output token estimation. At Task Start, the estimator may inspect the initial environment before predicting complete-run consumption. At the other three prediction points, it receives the observed execution record and the target-specific evidence available at that point without advancing the task. Missing fields are marked as not available. Self-Prediction prompt template Estimate the token consumption of the agent execution described below. Base your prediction on the task requirements, the specified model and agent configuration, and the evidence available at the supplied prediction point. Your estimate should describe the consumption of the agent operating under its existing workflow and execution limits. Your deliverable is a token-cost estimate. Do not implement a solution, modify task files, or submit an answer to the underlying task. Prediction target First, determine which part of the execution the estimate must cover. At Task Start, estimate all input and output tokens that a complete run would consume from the original task state until termination. The estimate includes the exploration, solution development, and verification that the run would require. Any inspection performed specifically to prepare this estimate belongs to the estimation session and is excluded from the predicted consumption. This inspection may improve your understanding of the task; however, the predicted run still begins from its original state. At Task Update, estimate the input and output tokens that will be consumed after the latest completed call until the task ends. Completed calls provide evidence of progress and consumption, but their recorded usage is outside of this remainingtask target. Their messages and tool results may still appear in later requests; reading that retained content in a later request contributes new input usage within the prediction scope. At Call Start, estimate the complete input and output consumption of the current call, including retries of its request. At In-call Update, estimate the same completecall quantity using the generation observed so far. This includes the input, supplied output prefix, continuation, and any usage attributable to retries within

36

the call. A call comprises one initial model request and any retry requests issued before tool feedback is received. A subsequent request made after receiving tool feedback belongs to a later call and is outside the call-level estimate. A task ends when the agent finishes, fails, or reaches an execution limit. Account for the termination behavior supported by the available evidence. Successful completion is one possible outcome, and repeated failures or exhausted limits may end the run earlier. Execution limits constrain the forecast, but they are not estimates of how much work the agent will perform. Understanding the task and execution evidence Read the task statement together with the agent instructions and completion conditions. Determine what the agent must establish, produce, or verify before it can complete the task. Consider how the available tools and workflows organize the work into model calls. Distinguish the number of tool actions from the number of model calls: several actions may be issued in one response, and a single unresolved issue may require multiple responses. At Task Start, use the available inspection tools to understand the initial environment while leaving task files unchanged. For a coding task, inspect the relevant source files, their dependencies, and the tests associated with the requested behavior. Assess whether the work appears localized or spans several components, whether the expected behavior is clearly specified, and how much investigation is likely to precede a change. For retrieval or document-based tasks, examine the supplied materials and resource descriptions to assess what evidence is already available and what additional information the agent would need to obtain. Focus the inspection on uncertainties that materially affect the expected workload. At the other prediction points, use the supplied execution record without advancing the task or obtaining additional tool results. Read the actions together with their observed outcomes. A proposed edit, a completed edit, and a successful test provide different evidence of progress. Identify what has been established, what remains unresolved, and which results are still pending. Treat unavailable information as unknown; the absence of a recorded failure does not establish success. For a task-level forecast, form a plausible continuation from the current state. A coding run may still need to investigate a failure, revise an implementation, run relevant checks, and prepare its submission. A retrieval task may require further searches because the current evidence covers only part of the question. A reasoning task with sufficient information may complete in one response. Use phases that fit the actual task and the observed state. At Task Update, include only phases or portions of phases that remain. Use the execution history to assess both progress and the cost of further work. Recent calls can indicate typical response lengths, request growth, and the amount of work the agent accomplishes per interaction. Compare calls with similar roles where possible. Repeated searches, recurring test failures, or several calls without resolving an outstanding issue may indicate additional investigation or revision. A sequence of completed checks may indicate that only final verification or submission remains. Relate these observations to the work still required before estimating the remaining call count. For a call-level forecast, focus on what the current request asks the model to generate . Determine whether the response is likely to contain a short tool invocation, several tool-call arguments, a substantial code fragment, an explanation, or a final answer. At the In-call Update, examine how much of that response has already been produced and what remains unfinished. Use the supplied timing information as supporting evidence when it is informative, while keeping elapsed time distinct from token counts. Constructing the token estimate For task-level predictions, the forecast execution is divided into a small number of non-overlapping phases. For each phase, the likely number of model calls, the input those calls will receive, and the output they will generate should be assessed. Let the phase estimates reflect the actual continuation you expect. Final verification or submission may require only one additional call. For the call-level prediction, a single current_call phase is sufficient. Estimate the input consumption from the content submitted for each model request. This can include system and agent instructions, task statements, tool definitions, retained conversation history, content generated from earlier calls, and tool results. Use an exact supplied input token count when one is available. When later requests have not yet been assembled, estimate their size from the current context, expected additions, and specified context retention behavior.

37

Account for retained content each time it is submitted. A tool result added early in a run may appear in several subsequent requests; therefore, its contribution depends on both its size and the number of later calls that retain it. Likewise, additional investigation can increase the call count and enlarge the context of those calls. Reflect both effects in the phase estimates. Apply context pruning, summarization, or compaction only when supported by the supplied workflow or observations. Estimate output consumption from the responses expected within each phase. Include generated text, code, tool-call arguments, and reasoning tokens according to the supplied accounting convention. The final answer may be short, even when intermediate responses consume substantial tokens. When reasoning tokens are already included in the reported output total, count them only once. Tool-generated search results, file contents, and test logs contribute model input when submitted to a later request; their production by a tool is not itself an LLM output. Use confirmed usage as an accounting anchor. Within the selected scope, include each confirmed input or output count exactly once. At In-call Update, the visible prefix is already part of the complete output being predicted; add only the expected continuation to its counted contribution. Do not add the prefix again when a supplied cumulative output count already includes it. Character and byte counts are observations about text length, not exact token counts; use the supplied token measurements or accounting information where available. Account for retries within the call according to the request history, observed failures , and the configured retry policy. The recorded usage from retries that have already occurred within the prediction scope is included. Estimate further retry consumption only to the extent supported by the current conditions. The maximum retry allowance defines a limit and does not imply that every retry will occur. Combine these estimates into the best-supported forecast. Consider whether the evidence supports a straightforward continuation, additional revision cycles, or early termination, and let this assessment inform the predicted workload. Avoid adding unexplained safety margins to each phase. Preparing the final output Return a non-negative integer for each token estimate. The phase input estimates must sum to predicted_input_tokens, and the phase output estimates must sum to predicted_output_tokens. Their combined sum must equal the predicted_total_tokens. Check that the result covers the requested scope and includes its confirmed usage without duplication. A complete call estimate must not fall below the confirmed consumption of that call. A Task Update estimate excludes completed call usage and is zero when termination is confirmed and no further model calls remain. Set predicted_total_tokens to the median (50th percentile) of token consumption for the specified prediction target. Set lower_total_tokens and upper_total_tokens to its 5th and 95th percentiles, respectively, so that the interval targets 90% coverage. These quantiles describe uncertainty in token consumption within the same prediction scope. Return non-negative integers satisfying lower_total_tokens <= predicted_total_tokens <= upper_total_tokens. When confirmed consumption is supplied for that scope, all three estimates must be at least that amount. At Task Update, the scope includes only future consumption; set all three estimates to zero when termination is confirmed and no further model calls remain. Use one breakdown_by_phase entry for each phase included in the forecast, with names that describe the expected work. For either call-level prediction point, use current_call when no further decomposition is needed. Submit a JSON object with the following fields: { "predicted_input_tokens": <integer>, "predicted_output_tokens": <integer>, "predicted_total_tokens": <integer>, "lower_total_tokens": <integer>, "upper_total_tokens": <integer>, "breakdown_by_phase": [ { "phase": "<phase name>", "input_tokens": <integer>, "output_tokens": <integer> } ] } Use the completion interface specified below. Submit the JSON estimate without a task solution or additional explanatory text. Fields marked not available provide no additional evidence and must not be treated as zero-valued observations.

38

Prediction point: {{prediction_point}} Task: {{task_description}} Difficulty, when provided: {{task_difficulty}} Model and agent configuration: {{model_and_agent_configuration}} Agent workflow and available tools: {{agent_instructions_and_tool_descriptions}} Execution limits and token accounting: {{execution_limits_and_token_accounting}} Initial environment: {{initial_environment}} Observed execution history and tool feedback: {{execution_history_and_tool_feedback}} Current request: {{current_request}} Generated prefix and stream timing: {{generated_prefix_and_stream_timing}} Confirmed usage within the prediction scope: {{confirmed_usage_in_scope}} Completion instruction: {{completion_instruction}} Estimate the consumption for the specified prediction point and submit the JSON estimate.

The prediction_point field selects one of the four scopes defined above. When the harness exposes a finish tool, completion_instruction requests submission of the JSON estimate through that tool. Otherwise, it requests the same object as the final response. Estimation calls and inspections performed solely for estimation are recorded as prediction overhead. The input and output fields refer to the selected scope, and therefore, predicted_total_tokens represents complete-task consumption at Task Start, complete-call consumption at both call-level points, and remaining-task consumption at Task Update.

39

Record · ID 1108741 · SHA-256 c52dd4fe12c71841
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.