Conceptio › Archive › arXiv CS
arXiv CSopen access

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

arXiv:2609.24985v1 [cs.LG] 21 Sep 2026

Zixiang Chen, Wenting Zhao, Zhepeng Cen, Akshara Prabhakar, Jielin Qiu, Jianguo Zhang, Zhiwei Liu, Tulika Manoj Awalgaonkar, Liangwei Yang, Shelby Heinecke, Silvio Savarese, Huan Wang Salesforce AI Research

Abstract Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect downstream randomness rather than differences between the current actions. We introduce Critical-State RL to identify trainable states in multi-turn interactions. Given task-defined candidate calls and local rewards, the method assesses whether each reward captures the action’s effect on task success and whether improvement over a reference policy is possible. It then uses nested sampling to separate action-dependent reward variation from continuation noise and optimizes the policy at the selected states using contextual-bandit training. Experiments on the Berkeley Function Calling Leaderboard (BFCL) v4 compare training at diagnostic-selected states with training at alternative states. For missing-function tasks, the diagnostic selects the response after the tool becomes available; for missing-argument tasks, it selects the response before the missing argument is supplied. Training the selected responses improves performance, including about 14 percentage points on the missing-function task, while training the alternatives leaves performance flat or worse. We further apply the recipe across models and tasks, including logged repeat-call avoidance and memory management.

1 Introduction Multi-turn tool use involves a sequence of model decisions, yet a trajectory-level reward does not identify which decisions would benefit from training. For example, when a required tool becomes available partway through an interaction, the responses before and after its availability are both candidate training targets. Which response benefits from training can vary by model and task. Prior work studied turn-level credit assignment [50] and state selection for local training [1, 49]. We focus on whether a candidate call offers an action-dependent signal for training. Observing both successful and failed rollouts does not establish this: when rewards depend on downstream interactions, the same current action can be followed by success or failure. Reward variation then mixes differences between actions with randomness in later interactions. The relevant question is whether changing the current action changes its expected reward. We introduce Critical-State RL to diagnose this signal before training. Given task-defined candidate model calls and a local reward, we assess whether the reward captures the action’s effect on task success and whether the reference policy leaves room for improvement. On the model’s own rollouts, we sample alternative actions at a fixed context, then hold each action fixed while resampling its continuation. This nested sampling separates action-dependent reward variation 1

Critical-State RL

from continuation noise. The diagnostic ranks qualifying candidates by this signal. We then optimize the policy at the selected call using contextual-bandit training. Surrounding calls provide context or reward without receiving gradient (Figure 1). We test the diagnostic’s turn selections in a four-cell Gemma study on BFCL [2], training selected and alternative turns separately in two task categories (§4). The diagnostic selects the response after tool availability for missing-function tasks (miss_func) and the response before missing arguments arrive for missing-argument tasks (miss_param). Training the selected turns improves performance, while their alternatives stay flat or decline (§4.2). Under paired deterministic evaluation, the miss_func arm records 0.14 → 0.283 ± 0.015 across four seeds; the miss_param arm records 0.435 → 0.473 ± 0.010. The same training recipe applies to logged repeat-call avoidance, Nemotron missing-function interactions, and memory management (§5). Nemotron trains the decision before tool availability rather than Gemma’s recovery response (§5.2), illustrating why the training location is chosen for each setting. Contributions. • A training-free diagnostic for selecting model calls. We separate action-dependent reward variation from continuation noise before training. The theory links this signal to unrestricted local-label improvement potential under a small KL budget and bounds finite-sample selection errors. • Local training guided by the diagnostic. We test the selected and alternative turns in a controlled four-cell study, apply the recipe across settings (§5), and establish a sufficient condition for small local updates to improve full-rollout return (Proposition 5). t0

t1

t2

t3

t4

t5

t6

Trajectory-level Fixed-role local RL

D

Critical-State RL

D

R

Missing-function example: decision (D) vs. recovery (R)

Figure 1. Gradient placement in a missing-function example t0 , . . . , t6 , with decision (D) before tool availability and recovery (R) after. Trajectory-level training updates every turn; fixed-role local RL always chooses the same role; Critical-State RL diagnoses candidate model calls and trains the selected one. Reading guide. Shaded = receives gradient. The displayed fixed-role rule chooses decision; the four-cell study also evaluates an always-recovery rule.

2 Related work State selection for local training. PivotRL reused expert trajectories collected for supervised fine-tuning (SFT), profiled their states under a frozen reference policy, and retained low-mean states with nonzero verifier variance for group relative policy optimization (GRPO) [1]. Its verifier r (s, a) rewarded locally acceptable actions rather than exact matches to demonstrations. Critical Step Optimization (CSO) identified candidate steps in failed policy trajectories with a process reward model, verified expert-proposed alternatives through policy continuations, and trained on the resulting step-level preference pairs [49]. Critical-State RL starts from task-defined candidate model calls and local labels, then separates action-dependent variation from continuation noise to select a call for training. 2

Critical-State RL

Karino et al. [47] identified critical states by action-wise expected-return variance, using learned Q-values with uniform action weighting to favor exploitation. We apply the same varianceof-conditional-means principle under the base policy, estimating occurrence-local label means by nested fixed-prefix sampling to select a call for local RL. App. E.1 analyzes the conditional recurrence regime relevant here. Turn-level rewards and credit assignment. MT-GRPO used per-turn rewards [5], while processsupervision methods graded intermediate steps [24, 25]. Group-in-Group Policy Optimization (GiGPO) grouped actions from recurring environment states to estimate step-level relative advantages alongside episode-level credit [50]. Other approaches redistributed returns [13, 14], conditioned credit on future events or counterfactuals [15, 16], or used hindsight and fixed-history comparisons, as in TCPO [6]. ArCHer, AgentPRM, and SWEET-RL learned value or reward models [10, 11, 12]. TRACE derived tool-boundary temporal-difference rewards [17], CAST obtained turn advantages from game solvers [18], and CrEST combined turn-level verified advantages with token-level modulation [20]. Executed-replay auditing estimated policy-conditional step contributions by resampling actions and continuing execution [7]; LOTAPO replaced a turn and its retrieval observation with a fixed placeholder to estimate its effect [19]. At the execution-interface level, Agent Lightning exposed agent executions as training transitions [8]; v1.0 formalized a harness-owned interaction loop in which the trainer observed large language model (LLM) request/response pairs [9]. Training objectives after state selection. RLSTA trained the final response with a reward anchored in the model’s full-information single-turn likelihood [23], whereas Critical-State RL diagnoses which call to train. For a selected call, SFT learns from target actions, direct preference optimization (DPO) from preference pairs [21], and rejection-sampling fine-tuning from the highestreward sample in a candidate batch [22]. These choices concern how to train once a location is selected. Our targeted-SFT comparison (§4.2) evaluates this post-selection training stage.

3 Critical-State RL Critical-State RL first selects which model call to train, then updates that call using a local reward. We call a concrete model call an occurrence and its local reward an occurrence-local label. The selected occurrence, not its trajectory, is the training unit: surrounding calls supply context or reward without receiving gradient. A candidate phase groups contexts with a common task structure, such as a needed tool being unavailable. A phase qualifies as a critical state for a model and task configuration when its return-aligned label meets three conditions at every retained occurrence: the label mediates the relationship between action and benchmark return (action-sufficiency); some action’s label mean exceeds the reference-policy mean (headroom); and action-conditioned label means vary under the base policy (trainability). Task structure supplies the phase, label, and sufficiency argument; reference- and base-policy rollouts measure headroom and trainability. Figure 2 separates this workflow, tested in §4, from task-specific label construction (App. D.2). A mixed-reward group does not imply trainability. GRPO [3] forms a critic-free baseline from each prompt’s completion-group rewards. DAPO-style dynamic sampling [4] filters out binarycorrectness groups whose sampled outputs are all correct or all incorrect. By the standard law of total variance, the split Var(z | x ) = Vara Q( x, a) + Ea Var(z | x, a), where Q is the label mean given context and action, separates action-dependent label variance from downstream noise [47]. 3

Critical-State RL

DAPO’s mixed-outcome condition reflects Var(z | x ) > 0 at the sample-group level. Even identical reward marginals and headroom can conceal different action signals (App. B, Figure 3). A concrete example: clean refusals with mixed rewards. In the no-think Gemma miss_func pilot, a sampled mixed-reward prompt group makes this distinction visible. In every surviving rollout, the decision response is a clean refusal with no tool call, yet reward swings 0 ↔ 1 with within-group std 0.477. A mixed-outcome filter would retain this group, although its reward differences do not distinguish the observed decision behavior. App. C.3 documents the pilot’s training setup and results. Nested within-prefix sampling. Without updating model parameters, we estimate both terms by nested sampling on the model’s own rollout distribution: sample a policy prefix, draw candidate actions at that fixed prefix, then hold each action fixed while resampling its reward-only continuation. Action means and within-action variation separate signal from continuation noise; App. B derives the finite-sampling correction and bounds ranking errors. Ground-truth teacher-forced prefixes can mask recovery-turn trainability by saturating the action distribution at the candidate occurrence. Pooling raw labels across model-generated prefixes mixes action variation with upstream randomness. After candidate occurrences are screened for action-sufficiency and headroom, the diagnostic ranks qualifying occurrences under a common label by arg maxturn t Varat E[z | xt , at ]; App. C.3 shows that different configurations can select different turns. The repeat-call action-only label and the Gemma continuation-noise case are documented in Table 4 and App. C.3. Under a small KL budget, action-dependent variance governs the best unrestricted local-label improvement to leading order at a fixed prefix and continuation kernel (Theorem 1). core workflow

task-specific label design

Candidate model calls task-defined phases and rollout contexts

Select a call

Occurrencelocal RL

reference-policy headroom; action-dependent label variance

gradient on selected call only; surrounding calls receive no gradient

(training-free)

Local reward action-sufficient label from: action, continuation, or proxy

Figure 2. Task structure supplies the candidate phase and occurrence-local label; policy rollouts then measure headroom and action-dependent variance before local training. The BFCL four-cell study in §4 tests this measured selection step in Gemma; App. D.2 presents a contrasting label construction in a memory application.

3.1 Core operation: occurrence-local RL Occurrence-local RL applies its loss to one selected model call per training unit. It supplies context either as a precomputed prompt without an environment loop or through a trajectory whose downstream, gradient-free turns provide the consequence signal. The memory applications and no-think Gemma’s missing-function recovery training use the first form (App. C.3); Nemotron’s missing-function decision training uses the second (§5.2).

4

Critical-State RL

Algorithm 1 Diagnose an occurrence, then train it with an occurrence-local label. Input: task-defined candidate phase and label; base and reference policies held fixed during diagnosis. 1. Screen candidates. Assess action-sufficiency from task structure and check reference-policy headroom for each candidate occurrence. 2. Separate signal from noise. Sample prefixes and actions from the base policy. At each fixed prefix, hold each sampled action fixed while resampling any reward-only continuation. Estimate action means and continuation variances, then apply the finite-sampling correction to estimate action-dependent variance (App. B). 3. Select the occurrence. For each candidate, average the corrected estimates across its sampled prefixes. Among candidates passing the gates under a common label, select the one with the largest average. 4. Train locally. Optimize the policy using the local label, applying the RL loss only to the selected call’s generated tokens. Its prefix supplies context; any continuation supplies reward without receiving gradient. Output: the trained policy and its selected training location, local label, and gradient boundary.

3.2 Task-specific reward construction Task structure supplies the local reward used for diagnosis and training. The following applications combine checks on the current action with scores of downstream behavior or preserved information. Missing-function interaction (Nemotron). For Nemotron’s missing-function application (§5.2), the multiplicative reward in Eq. (1) combines a behavior gate with a score from sampled recovery: ztool = no_write( adec ) × consequence(yrec ),

(1)

Here adec and yrec denote the decision action and recovery output. The gate no_write( adec ) ∈ {0, 1} is 1 iff the decision makes no state-changing write; consequence(yrec ) ∈ [0, 1] scores whether recovery contains the required tool call. The sampled recovery supplies reward without receiving gradient. Here ztool supplies the occurrence-local label z of App. B.1, while R denotes the benchmark terminal return. Memory storage. In contrast to Eq. (1), memory uses an additive reward for correct operation order and preservation of answer-critical information, with hard gates for invalid actions. A precomputed keyword proxy grades storage text against 2–5 required terms extracted by GPT-4o during preprocessing [31], including meaningful negations such as the not in “not married.” The proxy measures keyword retention in the storage response; the benchmark evaluates the full storage-to-answer chain.

4 BFCL four-cell study of occurrence selection Benchmark setting. We test whether the diagnostic identifies which candidate turn benefits from training. Evaluation uses the BFCL [2] v4 multi_turn suite: base, miss_func, long_context, and miss_param. In our configuration, the cumulative multi-turn evaluator penalizes extra writes, missing calls, and step-cap violations but not read-only queries. An earlier no-think Gemma pilot trained the miss_func decision turn and produced a mixedreward group of clean refusals with no call. That observation motivated the nested variance decomposition. Appendix C.3 gives its training setup and supplementary results; Section 5 presents applications in other settings.

5

Critical-State RL

4.1 BFCL task mechanics In BFCL multi_turn miss_func, a required tool is withheld until turn k. The substantive user request appears at turn k −1, while the tool is unavailable; turn k carries an empty user message that the harness replaces with a fixed bridge announcing tool availability. The agent must therefore defer or issue a read-only query at turn k −1, then emit the held call once the bridge re-introduces it. We call turn k −1 the decision turn and turn k the recovery turn. Appendix Figure 4 illustrates one miss_func decision-turn instance. In miss_param, by contrast, the function is present but a required argument is missing until the following turn; the decision-turn failure is a premature state-changing call. 4.2 Diagnostic assignments and four-cell intervention Task mechanics define the candidate occurrences and labels. Using rollouts from frozen no-think Gemma-4-26B-A4B [26], the pre-training diagnostic in §3 assigns recovery for miss_func and decision for miss_param. The subsequent four-cell study fixes these assignments, then trains decision and recovery separately in each category. Table 1 pairs the training outcomes with mean corrected action variance: within each category, the selected occurrence has the higher estimate. Table 1. Base-policy diagnostics and training outcomes for Critical-State RL on BFCL v4 multi_turn. Mean corrected action variance

BFCL accuracy start → trained

∆

selected occurrence

0.0267

0.14 → 0.283 ± 0.015

+14.3 pp

alternative occurrence

0.0103

0.14 → 0.095

−4.5 pp

selected occurrence

0.0361

0.435 → 0.473 ± 0.010

+3.8 pp

alternative occurrence

0.0000

0.435 → 0.445

+1 pp

Cell and diagnostic assignment miss_func / recovery miss_func / decision miss_param / decision miss_param / recovery

Note. Action variance: finite-sampling-corrected estimates from base-policy samples, averaged across clean-prefix scenarios (App. C.3). Selected cells: BFCL mean±cross-seed std, seeds {42, 123, 7, 99}; alternatives: point estimates. All use deterministic nt=1 decoding, step 30, and paired 200-item BFCL v4 multi_turn evaluation. Bases differ from the category-specific five-evaluation-seed study (App. C.3). Bold marks selected occurrences; shading marks the primary result. The miss_param result is directional.

Recovery improves miss_func by 14.3 ± 1.5 pp but leaves miss_param flat at +1 pp; decision improves miss_param by 3.8 ± 1.0 pp but moves miss_func down by 4.5 pp. Selection-rule comparison. Table 1 compares the diagnostic with always-decision and alwaysrecovery selection under the same occurrence-local RL interface. Each fixed rule improves one category but leaves the other flat or worse; the diagnostic selects the improving turn in both. We evaluate all cells at the fixed step 30 checkpoint. In the two diagnostic-selected cells, gym validation is monotone, and BFCL scores from steps 10–30 lie on a stable plateau above base. Transfer to a novel bridge. We replace the canonical bridge at evaluation with an unseen wording that names no tools or functions: “I adjusted things on my end.” Using the seed-123 step-30 checkpoint with deterministic decoding on 200 paired miss_func examples, accuracy rises from 0.095 to 0.225 after recovery-turn RL (+13.0 pp), matching the canonical bridge’s 0.140 → 0.270 gain under the same protocol. The new wording is harder for the base policy, yet preserves the recovery gain, demonstrating transfer beyond the bridge wording used in training.

6

Critical-State RL

Table 2. Monte-Carlo return-to-go (RTG) comparison on the non-saturated miss_param cell. Method Monte-Carlo RTG temporal credit; full multi-turn

Critical-State RL contextual-bandit training

Gym val.

BFCL miss_param

∆

0.997

0.45

+1.5 pp

0.77

0.473 ± 0.010

+3.8 pp

Note. The Critical-State RL BFCL result is the four-seed summary from Table 1. RTG is a single-seed comparator and uses Monte-Carlo return-to-go rather than a learned parametric critic.

Return-to-go comparison. We compare occurrence-local RL with full multi-turn training using Monte-Carlo return-to-go. RTG reaches a gym validation score of 0.997 (versus 0.77 for the bandit), but its BFCL accuracy is 0.45, below Critical-State RL’s 0.473 ± 0.010 (Table 2). Targeted-SFT comparison. We also compare occurrence-local RL with single-seed targeted SFT at the selected turn. Targeted SFT trains on fixed action targets, whereas Critical-State RL samples actions on-policy and scores them with the verifier. On miss_func / recovery, Critical-State RL gains 14.3 pp, whereas the best targeted recovery-SFT checkpoint remains 6 pp below base. The targeted-SFT runs over-bias toward firing the tool and break the required refusal; the Critical-State RL result preserves both behaviors. On miss_param / decision, targeted SFT gains 3 pp, compared with Critical-State RL’s +3.8 pp. Turn-mask control. This control varies which turns receive gradient while retaining the trajectorylevel scalar A = R − b. Full, recovery, random, and decision masks all leave the learnable miss_param cell near its starting accuracy: changing the mask alone does not recover the occurrencelocal RL gain.

5 Cross-setting applications Beyond the BFCL four-cell study (§4.2), the applications below retain the occurrence-local training unit and three diagnostic gates (§3), while adapting the training location and label. They illustrate three label sources: the current action, sampled recovery, and a precomputed proxy. 5.1 Logged interaction traces: repeat-call avoidance In logs of real interactions, Nemotron-Super-120B repeats a prior non-error (name, arguments) call after a usable result on 30% of flagged steps. On the same traces, GPT-4.1 [32] has a measured repeat rate of 0.61%, leaving substantial reference-policy headroom. Training updates the next action at each flagged trace position using the deterministic repeat_of_history label. Given the logged context, this label depends only on the current action, so its variation is action-dependent and has no downstream sampling term. Critical-State RL raises call-versus-answer agreement with GPT-4.1 from 37% to about 75% across three training seeds (72–78%). The note to Table 4 gives additional evaluation details. 5.2 Nemotron: missing-function application In this application, a needed tool is withheld and reintroduced on the next turn. The decision occurrence receives all gradient; a no-write gate scores its action, and the consequence in Eq. (1) comes from sampled reward-only recovery (§3.1; App. B.1). Critical-State RL yields a paired miss_func gain of +4.4pt over the Nemotron start under repeated deterministic evaluation. Gemma instead selects the miss_func recovery occurrence (App. C.3); App. D.1 gives the Nemotron instantiation and evaluation.

7

Critical-State RL

5.3 Memory management: xLAM and Gemma applications In the xLAM setting, a requested add fails when memory is full. Metadata selects the storage response after that failure for training; the response must remove a safe entry before adding the new one. An additive behavior–content reward scores operation order and keyword retention, using a precomputed proxy and hard gates for invalid actions. The BFCL v4 agentic/memory sub-task mean moves from 34.54% for the xLAM start to 50.54% for the selected memory-trained checkpoint (App. D.2). The no-think Gemma-4-26B-A4B application extends the memory setting to both storage and retrieval occurrences. A retuned training mixture uses separate single-occurrence examples for storage and retrieval: storage uses the additive behavior–content reward and precomputed keyword proxy, while retrieval is scored by coverage of facts needed for the expected query. Its memory sub-task mean moves from 38.1% to 52.7% (App. D.2.2).

6 Discussion Why diagnosis precedes training. The trainable call depends on the task and model configuration. The four-cell study selects different turns across categories, while the Gemma pilot shows that miss_func can shift from decision to recovery across model-and-reasoning configurations (App. C.3). Applying the method to a new setting therefore means reassessing the candidate calls, not reusing a fixed training turn. The theory explains three aspects of this workflow. Action-dependent variance measures unrestricted local-label improvement potential; finite-sample analysis bounds ranking error (App. B). Local credit removes noise from independent sibling outcomes that enters trajectory credit (App. E.1). Updating shared parameters can also change other calls; the drift analysis bounds the expectedreturn gap between using the updated policy only at the selected call and using it throughout the rollout (App. E.2).

7 Limitations The four-cell study tests categorical turn selection rather than calibrating variance magnitudes or diagnostic thresholds. Repeat-call, Nemotron, and memory are applications rather than matched tests of the diagnostic or design choices. The four-cell, Nemotron missing-function, and memory settings have K =1. Logged repeat-call training uses single-step examples from flagged trace positions. The recurrence and drift results are conditional (Apps. E.1 and E.2). App. C.3 reports evaluation repetition, serving variability, and training-unit provenance.

8 Conclusion Critical-State RL diagnoses action-dependent learning signal and uses it to select model calls for local training. The BFCL four-cell study on no-think Gemma-4-26B-A4B supports the assignments miss_func → recovery and miss_param → decision: selected turns improve, while alternatives stay flat or decline (§4.2). Across applications, the training location and reward adapt to the task, while the workflow remains the same: diagnose where to train, then optimize the selected call.

8

Critical-State RL

References [1] Junkeun Yi, Damon Mosk-Aoyama, Baihe Huang, Ritu Gala, Charles Wang, Sugam Dipak Devare, Khushi Bhardwaj, Abhibha Gupta, Oleksii Kuchaiev, Jiantao Jiao, Jian Zhang, and Venkat Srinivasan. PivotRL: High accuracy agentic post-training at low compute cost. arXiv:2603.21383, 2026. [2] Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The Berkeley Function Calling Leaderboard (BFCL): From tool use to agentic evaluation of large language models. In International Conference on Machine Learning (ICML), PMLR 267:48371–48392, 2025. [3] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300, 2024. [4] Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, et al. DAPO: An open-source LLM reinforcement learning system at scale. arXiv:2503.14476; published in Advances in Neural Information Processing Systems (NeurIPS), 2025. [5] Quan Wei, Siliang Zeng, Chenliang Li, Zhongruo Wang, William Brown, Oana Frunza, Wei Deng, Anderson Schneider, Yuriy Nevmyvaka, Yang Katie Zhao, Alfredo Garcia, and Mingyi Hong. Reinforcing Multi-Turn Reasoning in LLM Agents via Fine-Grained Reward Structure and Credit Assignment. arXiv:2505.11821, 2025. [6] Sicong Liao, Zhi Chen, and Yaohua Tang. TCPO: Turn-Level Credit Policy Optimization. arXiv:2608.01667, 2026. [7] Haiyue Zhang. Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay. arXiv:2608.19760, 2026. [8] Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K. Qiu, and Yuqing Yang. Agent Lightning: Train ANY AI Agents with Reinforcement Learning. arXiv:2508.03680, 2025. [9] Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, and Chong Luo. Agent Lightning v1.0: Towards Harnessed Agentic RL. arXiv:2608.17528, 2026. [10] Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, and Aviral Kumar. ArCHer: Training language model agents via hierarchical multi-turn RL. In International Conference on Machine Learning (ICML), PMLR 235:62178–62209, 2024; arXiv:2402.19446. [11] Zhiheng Xi, Chenyang Liao, Guanyu Li, Yajie Yang, Wenxiang Chen, et al. AgentPRM: Process reward models for LLM agents via step-wise promise and progress. arXiv:2511.08325, 2025. [12] Yifei Zhou, Song Jiang, Yuandong Tian, Jason Weston, Sergey Levine, Sainbayar Sukhbaatar, and Xian Li. SWEET-RL: Training multi-turn LLM agents on collaborative reasoning tasks. arXiv:2503.15478, 2025. [13] Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter. RUDDER: Return decomposition for delayed rewards. In Advances in Neural Information Processing Systems (NeurIPS), 2019. [14] Zhizhou Ren, Ruihan Guo, Yuan Zhou, and Jian Peng. Learning long-term reward redistribution via randomized return decomposition. In International Conference on Learning Representations (ICLR), 2022. [15] Anna Harutyunyan, Will Dabney, Thomas Mesnard, Mohammad Gheshlaghi Azar, Bilal Piot, Nicolas Heess, et al. Hindsight credit assignment. In Advances in Neural Information Processing Systems 32 (NeurIPS), 2019. [16] Thomas Mesnard, Théophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, et al. Counterfactual credit assignment in model-free reinforcement learning. In International Conference on Machine Learning (ICML), PMLR 139:7654–7664, 2021. [17] Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, Hao Cheng, Mike Hang Wang, Jianfeng Gao, and Sharon Li. TRACE: Turn-level reward assignment via credit estimation for long-horizon agents. arXiv:2607.13988, 2026. [18] Yu Wang, Yi-Kai Zhang, Wentao Shi, Ziang Ye, Yuchun Miao, Yueqing Sun, Qi Gu, Xunliang Cai, Lan-Zhe Guo, Han-Jia Ye, and Fuli Feng. CAST: Game Solvers as Turn-Level Teachers for LLM Agents. arXiv:2607.25308, 2026. [19] Qiang Zhu, Jiajun Wu, and Longyi Wang. LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning. arXiv:2607.13501, 2026. [20] Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, and Leilei Gan. Teach the magnitude, not the direction: Verifier-bounded credit assignment for multi-turn multi-step LLM agents. arXiv:2608.13179, 2026. [21] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems (NeurIPS), 2023. [22] Hugo Touvron et al. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023. [23] Xingwu Chen, Zhanqiu Zhang, Steven Y. Guo, and Difan Zou. Breaking contextual inertia: Reinforcement learning with single-turn anchors for stable multi-turn interaction. In Findings of the Association for Computational Linguistics: ACL 2026, pages 6302–6318, 2026. [24] Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process- and outcome-based feedback. arXiv:2211.14275, 2022.

9

Critical-State RL

[25] Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In International Conference on Learning Representations (ICLR), 2024. [26] Gemma Team. Gemma 4 technical report. arXiv:2607.02770, 2026. [27] An Yang et al. Qwen3 technical report. arXiv:2505.09388, 2025. [28] Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Hoang, Shirley Kokane, et al. xLAM: A family of large action models to empower AI agent systems. In Proceedings of NAACL, pages 11583–11597, 2025. [29] Zuxin Liu, Thai Hoang, Jianguo Zhang, Ming Zhu, Tian Lan, Shirley Kokane, et al. APIGen: Automated pipeline for generating verifiable and diverse function-calling datasets. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track; arXiv:2406.18518, 2024. [30] NVIDIA. Nemotron 3 Super: Open, efficient mixture-of-experts hybrid Mamba-Transformer model for agentic reasoning. arXiv:2604.12374, 2026. [31] OpenAI. GPT-4o system card. arXiv:2410.21276, 2024. [32] OpenAI. Introducing GPT-4.1 in the API. Online publication, 2025. https://openai.com/index/gpt-4-1/. [33] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), 2022. [34] John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz. Trust region policy optimization. In International Conference on Machine Learning (ICML), PMLR 37:1889–1897, 2015. [35] Ronald J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3–4):229–256, 1992. [36] Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems 12 (NIPS), pages 1057–1063, 1999. [37] Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter. Variance reduction techniques for gradient estimates in reinforcement learning. Journal of Machine Learning Research, 5:1471–1530, 2004. [38] Sham M. Kakade. A natural policy gradient. In Advances in Neural Information Processing Systems 14 (NIPS), 2001. [39] Jan Peters, Katharina Mülling, and Yasemin Altun. Relative entropy policy search. In Proceedings of the AAAI Conference on Artificial Intelligence, 24(1):1607–1612, 2010. [40] Henry Lam. Robust sensitivity analysis for stochastic systems. Mathematics of Operations Research, 41(4):1248–1275, 2016; arXiv:1303.0326. [41] Takashi Goda. Computing the variance of a conditional expectation via non-nested Monte Carlo. Operations Research Letters, 45(1):63–67, 2017; arXiv:1605.05454. [42] Thomas Spooner, Nelson Vadori, and Sumitra Ganesh. Factored policy gradients: Leveraging structure for efficient learning in MOMDPs. In Advances in Neural Information Processing Systems 34 (NeurIPS), 2021. [43] Andrew Y. Ng, Daishi Harada, and Stuart Russell. Policy invariance under reward transformations: Theory and application to reward shaping. In International Conference on Machine Learning (ICML), pages 278–287, 1999. [44] Richard S. Sutton. Learning to predict by the methods of temporal differences. Machine Learning, 3(1):9–44, 1988. [45] John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation. In International Conference on Learning Representations (ICLR), 2016. [46] George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E. Turner, Zoubin Ghahramani, and Sergey Levine. The mirage of action-dependent baselines in reinforcement learning. In International Conference on Machine Learning (ICML), PMLR 80:5015–5024, 2018. [47] Izumi Karino, Yoshiyuki Ohmura, and Yasuo Kuniyoshi. Identifying critical states by the action-based variance of expected return. In Artificial Neural Networks and Machine Learning – ICANN 2020, pages 366–378, 2020; arXiv:2008.11332. [48] Yunpeng Sun, Daniel W. Apley, and Jeremy Staum. Efficient nested simulation for estimating the variance of a conditional expectation. Operations Research, 59(4):998–1007, 2011. [49] Mukai Li, Qingcheng Zeng, Tianqing Fang, Zhenwen Liang, Linfeng Song, Qi Liu, Haitao Mi, and Dong Yu. Verified critical step optimization for LLM agents. In Findings of the Association for Computational Linguistics: ACL 2026, pages 39627–39639, 2026. [50] Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An. Group-in-group policy optimization for LLM agent training. In Advances in Neural Information Processing Systems 38 (NeurIPS), 2025.

Appendix roadmap. Appendices A and B provide notation and core analysis. Appendix C covers the recipe (C.1), candidate identification (C.2), and evaluation (C.3); D presents applications (overview: Table 4; Nemotron: D.1); E analyzes credit and drift (E.1–E.2); F provides prompts. 10

Critical-State RL

A Notation The method’s phase, occurrence, and selected occurrence are defined in §3; Def. 1 formalizes “critical state.” The table collects notation for the general formalism, recurrence analysis, and stopping analysis. “Local-label credit” denotes gradient estimation using an occurrence-local label. Term or symbol

Meaning and role

General critical-state formalism (App. B) zi , Q( x, a) occurrence-local label and its mean Q( x, a) = E[z | x, a], distinct from App. E.2’s anchor return value Qθ0 . Vact , Vacc , F at a fixed prefix under the current policy: action-dependent label variance Vara Q; its score-accessible component µ⊤ F † µ; and score second moment F = E[ gg⊤ ]. Here µ = E[ g(z − q)] and F † is the Moore–Penrose inverse. φ x ( z ), λ ( x ) φ x (z) = E[ R | x, z] (assumed affine); local slope λ( x ) = ∂z φ x relating the local-label and trajectory gradients. q ( x ), F i Label baseline. Context-conditioned q( x ) = E[z | x ] used by the local-label estimator. Conditioning field. Fi = σ ( xi , ai , zi ), where the operator denotes the generated sigma-field, not the critical-state symbol below. π0 , ∆head , h reference policy for headroom; required margin ∆head > 0; occurrence-local label window (≤ h turns, h ≪ episode), Def. 1. w return residual R − z: return variation outside the occurrence-local label in general, ∑ j̸=i z j under K-fold recurrence (Cor. 1). √ Sz fixed-prompt label share Var(z)/ Var( R); the advantage-multiplier SNR ratio is Sz under additive independence (Cor. 1). Nemotron missing-function reward (§3.2) ztool = no_write( adec ) × multiplicative decision-turn occurrence-local label z of Thm 2. consequence(yrec ) no_write( adec ) ∈ {0, 1} process gate: 1 iff the decision-turn action emits no state-changing write. consequence(yrec ) ∈ [0, 1] score for whether the recovery-turn output contains the required tool call. adec / yrec decision-turn action and recovery-turn output, respectively. Recurrence-variance symbols (App. E.1) K retained visits to the abstract critical state σ within one episode. σ abstract critical state: an equivalence class of qualifying contexts, not a fixed turn. ai , ci action taken and correct action at occurrence i, respectively; the required action may vary across occurrences. ri = 1 [ ai = ci ] local correctness at occurrence i. qi , s2i Bernoulli success probability qi = Prθ ( ai = ci ), the label baseline for local correctness; s2i = qi (1 − qi ) = Var(ri ). R = ∑i ri terminal (trajectory-level) reward. b, At trajectory-return baseline and turn t advantage. In the recurrence analysis, the baseline is the population analogue of GRPO’s group mean; the general formalism conditions that baseline on context. ĝ flat single-occurrence local-label gradient estimator (O(1) variance under the stylized recurrence assumptions); traj denotes the trajectory-return estimator. v −i , vi , f i sibling-label variance ∑ j̸=i s2j ; local gradient trace variance; score second moment

√

SNR ∝ 1/ K

E∥ gi ∥2 . Shared credit adds v−i f i to the per-occurrence trace variance. per-occurrence gradient SNR scaling under the nondegenerate, independent √ equal-variance model; the scalar advantage-multiplier SNR ratio is exactly 1/ K.

Drift-bias and stopping symbols (App. E.2) J (θ ) = Eτ [ R(τ )] true multi-turn objective. hyb

Pθ

, Jhyb

b d0 , Q θ0 , Q

hybrid trajectory law using πθ at the selected call and πθ0 elsewhere; its expected return is the exact frozen surrogate. anchor selected-context distribution, anchor continuation return value, and its frozen b proxy. b J = Ed0 ,πθ Q.

11

Critical-State RL

Notation (continued) Term or symbol

Meaning and role

Dout

KL( Pθ ∥ Pθ ): summed KL outside the selected occurrence, averaged over current-policy rollouts.

0 Dsel ε0

selected-action KL averaged over anchor contexts, equal to KL( Pθ ∥ Pθ0 ). uniform upper bound on return-value proxy error; its contribution to the gain bound vanishes √ at the anchor (Eq. (5)). δ = Dout , exact surrogate gain p = ∆Jhyb , and gain lower bound G = p − κδ, with √ κ = 2Rmax ; δ⋆ maximizes G (Cor. 3).

δ, p(δ), G (δ)

hyb

hyb

B Core analysis B.1 When a candidate phase is a critical state Fix a policy πθ and terminal return R(τ ) ∈ [− Rmax , Rmax ]. A phase σ is an equivalence class of contexts, such as missing-function defer/recover, rather than a single fixed context. After routing, an episode has K ≥ 1 retained occurrences, each with context xi , action ai , and score gi = ∇θ log πθ ( ai | xi ). Its occurrence-local segment includes a bounded reward-only continuation. Definition 1 (Critical state). Relative to a policy πθ and a reference π0 , σ is a critical state if it admits an occurrence-local label zi , supplied by the environment or an oracle/proxy and measurable within at most h turns (h ≪ the episode length). With Q( x, a) := E[z | x, a], every retained context satisfies R ⊥ ai | ( xi , zi ), | {z }

(i) action-sufficiency

max Q( x, a) − Ea∼π0 Q( x, a) ≥ ∆head > 0, |a {z } (ii) headroom

Vara∼πθ (·| x) Q( x, a) > 0 . | {z } (iii) trainability

The headroom margin is task-specific. Contexts failing either (ii) or (iii) are split or filtered before ranking. The bounded label window prevents the vacuous choice z := R for any state. Benchmark-directed training also requires a return-aligned label; Theorem 2 gives the positiveslope condition for per-prompt gradient alignment. For no-think Gemma, no_write near 0.97 leaves little headroom in this gate; the pilot’s training comparison favors recovery (App. C.3). For the missing-function decision, the full label is z = no_write( adec ) × consequence(yrec ), denoted ztool in Eq. (1); R remains the benchmark terminal return. This operational proxy checks the held call and the decision’s process gate; the benchmark also checks extraneous recovery calls. Exact sufficiency is the theoretical condition on the whole label. Sampled recovery makes z stochastic even after the decision action is fixed. B.2 From the diagnostic to local improvement Action-dependent variance measures local improvement potential. Fix a prefix x and freeze the continuation kernel that supplies z. Write p( a) = πθ ( a | x ), Q( a) = E[z | x, a], q = E p Q, and Vact = Var p Q. The score is g( a) = ∇θ log p( a), with E p g = 0; vector variance denotes trace covariance. Proposition 1 (Continuation noise and the occurrence gradient). For a fixed baseline b = b( x ) and finite second moments, E[ g(z − b)] = E p [ g( Q − q)] =: µ,     Var g(z − b) = Var p g( Q − b) + E p ∥ g∥2 Var(z | x, a) . Thus Vact = 0 implies µ = 0, even when sampled labels vary. 12

Critical-State RL

Proof. Condition on a to get mean g( Q − b) and variance ∥ g∥2 Var(z | x, a). Total expectation and variance, with E p g = 0, give the identities. Action-conditioned means Q( a) a1 a4 a3 a2 Vact 1/4 0 0 1 1 Candidate 1 1 1 1/8 0 1 Candidate 2 2 2

Same label distribution 1/2 1/2

Average over actions

z=0

z=1

Figure 3. Same marginal labels, different action signal. With uniform base and reference policies, (0, 0, 1, 1) and (0, 12 , 12 , 1) are conditional Bernoulli means. Both give Bernoulli(1/2) labels and 1/2 headroom, but Vact = 1/4 and 1/8.

No selector using only independent marginal labels and headrooms can distinguish swapped candidates, whose rankings reverse. Its worst-case error is at least 1/2, at any sample size. The small-KL sensitivity expansion [40] gives the optimization meaning below; its exponentialtilt optimizer also appears in relative-entropy policy search [39].

Theorem 1 (Local improvement at a fixed KL budget). Let p have positive mass on a finite action set, and allow any new distribution p′ on that same set. Holding Q fixed, define  G x (ϵ) := max E p′ Q − E p Q . DKL ( p′ ∥ p)≤ϵ

If Vact > 0, then, as ϵ ↓ 0,

G x (ϵ) =

√

2ϵVact + O(ϵ).

If Vact = 0, then G x (ϵ) = 0 for every ϵ. Proof. Set A = Q − q and ψ(η ) = log E p eη A . Exponential tilting gives pη ( a) = p( a)eη A(a)−ψ(η ) , with gain ψ′ (η ) and KL ηψ′ (η ) − ψ(η ). For any p′ , DKL ( p′ ∥ pη ) = DKL ( p′ ∥ p) − ηE p′ A + ψ(η ) ≥ 0. Consequently, pη maximizes gain at its own KL budget. Since ψ(η ) = Vact η 2 /2 + O(η 3 ), choosing √ that budget to equal ϵ yields η = 2ϵ/Vact + O(ϵ) and the stated expansion. When Vact = 0, Q is constant on the action set. Average improvement over model-generated prefixes. Fix a finite-support prefix distribution ρ, set V̄act = Eρ Vact ( x ), and assume Theorem 1’s conditions at each prefix. With continuation values fixed, the largest average local-label gain over unrestricted conditional policies is p Eρ DKL ( p′x ∥ p x ) ≤ ϵ. Gρ (ϵ) = 2ϵV̄act + O(ϵ), This holds for V̄act > 0 as ϵ ↓ 0; zero variance gives zero gain at every budget. Normalize the preceding tilt separately at each prefix and set Ψ(η ) = Eρ log E px eη [Q(x,a)−q(x)] . Averaging the KL identity proves optimality; Ψ′′ (0) = V̄act gives the expansion. Prefix-mean differences supply no gain because ρ is fixed. The part accessible to a parameterized policy. At a fixed prefix, let F = E p [ gg⊤ ], let F † be its Moore–Penrose inverse, and define Vacc = µ⊤ F † µ. Compatible-function projection and the quadratic KL model [38] give h i Vact = Vacc + E p ( Q − q − g⊤ F † µ)2 , √ max µ⊤ u = 2ϵVacc . u⊤ Fu/2≤ϵ

13

Critical-State RL

Indeed, µ belongs to the range of F, and g⊤ F † µ is the least-squares projection of Q − q onto the score span. Orthogonality proves the first identity; Cauchy–Schwarz in the F metric proves the second. A full categorical policy has Vacc = Vact . Under a common label and equal small KL budgets, Vact ranks leading-order unrestricted potential at a fixed prefix. Across a fixed distribution of prefixes, a shared update under an average p ⊤ quadratic KL budget uses µ̄ = Ex µ x and F̄ = Ex Fx , with potential 2ϵµ̄ F̄ † µ̄. Applying the same projection under the joint law ρ( x ) p x ( a) gives µ̄⊤ F̄ † µ̄ ≤ V̄act . Opposing prefix gradients can cancel. The diagnostic is gradient-free; the four-cell intervention tests the shared-policy response. B.3 Finite-sample estimation and occurrence selection For a balanced nested design at a fixed prefix, draw n ≥ 2 independent actions and m ≥ 2 2 continuations per action, all independent conditional on the actions. Let Sbetween be the sample 2 variance (denominator n − 1) of the n action means, and Swithin the average within-action sample variance (denominator m − 1). Writing Vcont = Ea Var(z | x, a) gives 2 ESbetween = Vact +

Vcont , m

bact = S2 V between −

2 Swithin , m

bact = Vact . EV

2 Total variance of an action mean proves the first equality; ESwithin = Vcont gives the correction. This standard nested-simulation correction [48, 41] removes continuation noise in expectation; finite estimates can be negative.

From an unbiased estimate to reliable occurrence selection. Classical nested-variance analysis separates action coverage from continuation depth [48]. Here that distinction determines how reliably the diagnostic ranks candidate occurrences. Proposition 2 (Finite-sample diagnostic precision). Under the preceding balanced design with bounded labels, let d( a) = Q( a) − q and ν( a) = Var(z | x, a). Then 2 2 2 bact ) = Vara (d ) + 4Ea [d ν] + 2Ea [ν ] En,m := Var(V n nm nm(m − 1) 2 2(Vact + Vcont /m) + . n ( n − 1)

(2)

Consider two qualifying candidate occurrences at fixed prefixes under a common label scale, estimated using independent nested batches. If V1 > V2 , write ∆ = V1 − V2 and let E1 , E2 be their estimator variances from (2). Their misranking probability satisfies b2 ≥ V b1 ) ≤ Pr(V

E1 + E2 . ∆2 + E1 + E2

In particular, E1 + E2 ≤ δ∆2 /(1 − δ) guarantees correct selection with probability at least 1 − δ, for 0 < δ < 1. Proof. Center labels at q. For action i, let Yi = m−1 ∑ j (zij − q), let Si2 be its within-action sample variance, and set ∑ j̸=k (zij − q)(zik − q) . Ti = Yi2 − Si2 /m = m ( m − 1)

14

Critical-State RL

Conditional independence gives E[ Ti | ai ] = d( ai )2 and Var( Ti | ai ) = 4d( ai )2 ν( ai )/m + 2ν( ai )2 /[m(m − 1)]. Algebra rewrites the existing estimator as YY bact = 1 ∑ Ti − ∑i̸=k i k . V n i n ( n − 1) Since the groups are independent and EYi = 0, the two terms are uncorrelated; the second has variance 2(EYi2 )2 /[n(n − 1)]. Total variance for Ti and EYi2 = Vact + Vcont /m prove (2). Finally, b2 − V b1 has mean −∆ and variance E1 + E2 ; the one-sided Chebyshev (Cantelli) inequality gives V the ranking bound. For more candidates, sum the pairwise bounds against the unique best candidate. More continuations cannot replace more actions. As m → ∞ at fixed n, continuation noise 2 / [ n ( n − 1)], which is positive when vanishes but the estimator variance tends to Vara (d2 )/n + 2Vact Vact > 0. Conversely, any fixed m ≥ 2 allows consistent estimation as n grows, and the misranking bound vanishes for a fixed positive gap. Thus independent actions can resolve the ranking without precise estimates of every action’s label mean. Averaging estimates across sampled prefixes. For P independent prefixes Xℓ ∼ ρ with indepenb̄ act = P−1 ∑ P V b dent balanced (n, m) batches, set V ℓ=1 act ( Xℓ ). Conditional unbiasedness and total variance give b̄ act = V̄act , b̄ act ) = Varρ (Vact ( X )) + Eρ En,m ( X ) . EV Var(V P The same misranking bound applies to independent candidate datasets using these averaged targets and variances. More actions or continuations reduce within-prefix error; more independent prefixes also reduce between-prefix error. B.4 From local-label credit to terminal-return credit Proposition 1 isolates noise within the local label. The next result concerns a different source: terminal-return variation outside that label. traj

Theorem 2 (Per-prompt conditioning on an occurrence-local label). Let ĝi = gi ( R − b( xi )), with b( xi ) = E[ R | xi ], and ĝiflat = gi (zi − q( xi )), with q( xi ) = E[z | xi ]. Set Fi = σ ( xi , ai , zi ). If R ⊥ ai | ( xi , zi ) and φ x (z) := E[ R | x, z] is affine with slope λ( x ), then traj

E[ ĝi

| Fi ] = λ( xi ) ĝiflat ,

Var( ĝtraj ) = Var(λ( x ) ĝflat ) + E[∥ g∥2 Var( R | x, a, z)].

Proof. Sufficiency gives E[ R | x, a, z] = φ x (z). Affinity gives φ x (z) − E[ R | x ] = λ( x )(z − q( x )). Multiply by the measurable score and apply total variance. At a fixed prompt, λ( x ) > 0 preserves gradient direction; across prompts the return gradient weights local gradients by λ( x ). Binary labels make the link affine automatically; continuous missing-function and memory scores require calibration of the full label. A small λ attenuates local credit and a negative one reverses it. If the conditional-mean sufficiency residual is bounded by ε, the conditional-gradient identity has error at most ε∥ g∥. Corollary 1 (Return residual and independent recurrence). At a fixed prompt, suppose R = z + w with w independent of ( g, z). Both estimators have mean µ = E[ g(z − Ez)], and Var( ĝtraj ) = Var( ĝflat ) + E∥ g∥2 Var(w), 15

Sz =

Var(z) . Var( R)

Critical-State RL

For nonzero µ, define the multiplier signal-to-noise ratio (SNR) as ∥µ∥ divided by the multiplier’s standard √ deviation. Then SNRtraj /SNRflat = Sz . If w sums K − 1 independent sibling labels, each with the same positive variance as z and unaffected by the selected action, then Sz = 1/K. App. E.1 gives the full per-occurrence covariance and sample-cost result. For K = 1, w can still contain terminal-return variation beyond the local label. Proof. The centered residual g(w − Ew) has zero mean and zero covariance with g(z − Ez). Its trace variance is E∥ g∥2 Var(w); scalar variance addition gives Sz . Empirical single-occurrence scope. On Gemma-4-26B-A4B home-buying rollouts (n=6560), the per-turn-share proxy gives √ a pooled variance ratio Sz = Var(z)/ Var( R) = 0.31 (95% confidence interval (CI) [0.29, 0.33]; Sz is about 0.56), with slope λ near 1.7 and correlation 0.95. The fieldsurvey analysis has a signed terminal-reward mean difference of −0.021 between correct and wrong loadouts (App. E.1.5). The ratio and its square root depend on label scale: these are descriptive √ associations, not measured SNR losses or a fitted 1/ K exponent. The prompt-centered ratio is E[Var(z | x )]/E[Var( R | x )], equal to the marginal ratio when prompt means do not vary. App. E.1 analyzes group credit; App. E.2 analyzes changes to the surrounding policy.

C Implementation and evaluation C.1 Applying Critical-State RL to a new task Algorithm 1 gives the diagnostic and local-training steps. Here we detail how to prepare its task-defined inputs and when to run additional analyses. Identify a candidate phase when it is not given. Absent identifying metadata, cross-tabulate the model’s per-sample failures against a stronger reference, then bucket residuals by a structural axis (error type, turn position); a concentration of residuals identifies a candidate phase. Within that phase, select candidate rows from scenario metadata or a deterministic rule. For example, the missing-function application uses an LLM judge to exclude rows that another tool could satisfy. Transfer monitoring. When training against a frozen single-occurrence surrogate, monitor transfer and Kullback–Leibler (KL) drift under the stated conditions (App. E.2). The selected training location varies with the model and task configuration, while the task determines whether an occurrence-local label exists (App. C.3, App. E.1). Candidate identification and validation. For the missing-function application, humans proposed the candidate phase and designed the counterfactual validation; deterministic rules selected rows. App. C.2 records this procedure and its results. For memory storage, the memory-full flag supplies candidate membership. Optional coverage audit by slice. A coverage audit tests whether a structural mismatch between training and evaluation slices explains slice-specific changes. It partitions candidate occurrences along a structural axis (a state property, not a reward or hyperparameter), measures training– evaluation mismatch on that axis, and adds rows for the under-covered slice. A strictly-additive repair leaves existing rows, the reward, and hyperparameters untouched, while a size-matched control fixes the added row count and varies only the slice. The audit diagnoses a distribution gap; trainable-occurrence selection still comes from the diagnostic.

16

Critical-State RL

C.2 Candidate identification in the evaluated tasks Candidate identification. We proposed a defer-now, act-later candidate phase in miss_func and miss_param for occurrence-local RL (§3.1). For these BFCL configurations, we compared failures from the starting checkpoint with those from a stronger reference and grouped the failure patterns by BFCL error_type (e.g. empty_turn_model_response “defer/refuse” vs. instance_state_mismatch “wrong write”). Human interpretation of these patterns identified the empty_turn miss_func phase and the refusal pattern in no-think Gemma for further investigation. In memory storage, the memory-full flag directly identifies candidate rows. Counterfactual check. We compared the unmodified control with a clean re-ask and an explicit task restatement in place of the canonical bridge. In this one-way clean-prefix counterfactual, the state checker found that a clean prefix eliminated deferral (miss_param 17 → 0, miss_func 10 → 0), supporting a “refusal + vague re-ask” explanation of continued deferral. Eight-bit floating-point (FP8) plus expert-parallel nondeterminism made a full-pass rerun too noisy, so the result remains directional. Row selection and validation. A deterministic rule identifies candidate decision rows with an empty ground-truth action (for miss_func, turn k −1 before the held call, gated on gt[k −1] = [ ] and gt[k]). For Nemotron’s turn-0 rows, an LLM judge excludes requests that the available tools can fulfill (YES), retaining only those requiring the withheld tool (App. F). C.3 Training and evaluation details Gemma evaluation protocols. The four-cell study’s selected cells report means and standard deviations across training seeds {42, 123, 7, 99}. All cells use checkpoint step 30 and paired 200-item BFCL v4 multi_turn evaluation with nt=1 deterministic decoding (§4.2). The category-specific training study reports recovery-run peaks averaged over five evaluation seeds (≤ 0.015 std); the selected no-think Gemma-4-26B-A4B checkpoint records miss_func 0.110 → 0.44 (+33pt; supplementary results below).

Four-cell diagnostic sampling. Each scenario is one BFCL test item, sampled before training with 8 b̄ act estimated from actions and 4 continuations per action. Table 1 reports mean corrected action variance V pre-training samples using the estimator in App. B. The mean includes every clean-prefix scenario, with zero and negative estimates retained. The miss_param recovery cell has 54 rather than 64 scenarios because 10 had varying model-generated prefixes; the other three cells each retain 64.

Gemma training setup and supplementary results. We trained one no-think Gemma-4-26B-A4B [26]

model with category-specific turn assignments: miss_func rows train recovery, while miss_param rows train decision. The recovery input comprises the decision context + a sampled-but-frozen refusal + the canonical bridge. Only the sampled recovery action receives gradient; its score checks the held call under the same lenient name+argument comparator. The in-domain recovery consequence score rises 0.26 → 0.49 → 0.70 at steps 0/5/10. Table 3 reports BFCL results for the category-routed checkpoint. The decision-turn pilot retained Nemotron’s task, data, agent scaffold, and multiplicative reward (Eq. (1); App. D.1), with thinking disabled rather than chain-of-thought. For miss_func, in-domain homebuying validation remains near 0.21 (start 0.215), and BFCL accuracy moves 0.110 → 0.135. Under ztool = no_write( adec ) × consequence(yrec ), the decision already passes the no-write gate at a high rate (no_write near 0.97), leaving little headroom in this gate, while consequence(yrec ) varies with recovery success. The sampled-group example in §3 shows this variation among clean refusals. Missing-argument examples instead expose premature state-changing calls at the decision turn. Across three independent recovery-turn training runs, peak BFCL missing-function accuracies are 0.359 ± 0.006, 0.351 ± 0.011, and 0.439 ± 0.012, compared with 0.110 ± 0.013 at the start (mean±std over 5 evaluation seeds). 17

Critical-State RL

Table 3. Supplementary Gemma results: BFCL v4 target-cell accuracy for the no-think Gemma-4-26B-A4B study. Training configuration

Missing function

Missing argument

0.110 0.135 0.439

0.446 0.485 0.478

Starting checkpoint Decision-turn trained Category-routed, selected step 30

Note. Starting and category-routed rows are means over 5 evaluation seeds; the decision-turn-trained row is single-eval. The category-routed checkpoint is selected by official BFCL v4 weighted full-suite Overall. Its target-cell mean±std values are 0.439 ± 0.012 and 0.478 ± 0.006.

Nemotron missing-function application. The paired gain is +4.4pt under repeated deterministic evaluation (n=3; Nemotron start 0.408 ± 0.013 → 0.452 ± 0.021) and +6.5pt under same-day serving (0.430 → 0.495). Training on an independently designed synthetic task yields a deterministic +5.0pt transfer with a 128k context window. Nemotron-Super-120B checkpoint trajectories use single evaluations; paired serving checks support the reported gains. Memory applications. The selected xLAM checkpoint lies in a cluster of about 6–7 similarly performing checkpoints; the Gemma application evaluates five data-mixture runs (App. D.2).

Nemotron serving variability. In concurrent vLLM evaluations of Nemotron, temperature 0 runs vary by about 3–5pt because of batch nondeterminism, and cross-month reruns can shift by about 2–3pt. Evaluation therefore uses the full 128k context with serial batch-size-1 decoding or an independent multiserve check; shorter windows silently fail long thinking rollouts. The independently trained Nemotron checkpoint produced the same +5.0pt difference under serial decoding and a 16-way concurrent re-serve. Training units and cost. Nemotron trains the decision occurrence with the multiplicative label ztool = no_write( adec ) × consequence(yrec ). In the Nemotron comparison, a whole-trajectory optimizer step cost 16–30× as much as an occurrence-local RL step. The xLAM application trains one metadata-selected response with the additive reward and keyword proxy in App. D.2.1; the selected checkpoint is step 240. Nemotron renamed-domain training. We adapt NVIDIA Nemotron-Super-120B [30] with LoRA [33] (rank 64, α = 256) using GRPO with DAPO dynamic sampling in NeMo-RL (Megatron) and NeMo-Gym. The learning rate is 5 × 10−6 , lr_decay_iters = 30, and the KL penalty coefficient is 0; each step samples 64 prompts with batch multiplier 4 and 32 generations per prompt. Memory model identity and training stack. The evaluated checkpoint is an internal, unreleased research checkpoint from the xLAM line [28], used only for research and not part of any product or deployment; the public xLAM series is a separate release. It is a Qwen3-235B-A22B [27] fine-tune trained in our pipeline on function-calling data produced by APIGen [29], using verl with Fully Sharded Data Parallelism (FSDP) and GRPO with DAPO dynamic sampling.

18

Critical-State RL

D Applications Table 4. Four Critical-State RL application configurations from §5. Application

Selected occurrence Label source and construction

Reported result

Logged repeat-call (Nemotron)

next response at a flagged trace position

call/answer agreement with GPT-4.1: 37% to about 75% (72–78%)

Missing function (Nemotron)

decision before tool reward-only sampled recovery; availability no-write gate; multiplicative, Eq. (1)

Memory (xLAM)

memory-full storage response

precomputed GPT-4o keyword sub-task mean proxy; additive behavior–content 34.54% → 50.54% reward with hard gates

Memory (Gemma)

storage or retrieval response

storage: keyword proxy; retrieval: sub-task mean expected-query fact coverage 38.1% → 52.7%

deterministic action-only repeat_of_history

BFCL miss_func gain: +4.4pt

Note. The shared training unit is one selected response; its occurrence-local label is task-specific (§3.2). Repeat-call. The agreement range spans three training seeds. Training uses fixed logged contexts without online exploration. On the traces summarized in §5, the trained models skip 2–3 of 20 needed calls across seeds, compared with 2 at the start.

D.1 Nemotron: missing-function application D.1.1 Scenario, selected occurrence, and label

Context (earlier turns, abbreviated) [user] Submit a high-priority inspection request titled ‘Roof leak’ . . . [call] create_inspection_request( {"title": "Roof leak", . . . }) → {"id": 1, "status": "Open"} [user] Update the inspection request to lower the priority . . . [call] edit_inspection_request( {"request_id": 1, . . . }) → updated

Decision turn

get_my_inspection_requests withheld from the tool list [user] Show me all my inspection requests

× no_write( adec )=0 (gated): any state-changing call, e.g. edit_inspection_request(. . . ) ◦ no_write( adec )=1 (not gated): read-only query to the near-name decoy get_inspection_request({"request_id": 1}); reward then rests on the recovery turn ✓ no_write( adec )=1: defer in natural language, “I currently can’t list your inspection requests . . . ”

Recovery turn

Bridge message (tool now re-added)

(reward-only rollout; output yrec ) [call] get_my_inspection_requests({})

[user] I have updated some more functions you can choose from. What about now?

Figure 4. A missing-function decision-turn instance in the Nemotron decision-turn configuration (renamed home-buying variant). Note. The gate no_write( adec ) scores the decision under the convention in §4.1; consequence(yrec ) scores the held recovery call (Eq. (1)).

19

Critical-State RL

BFCL-aligned scoring. The withheld tool is encoded as missed_function = {"k" : [tool]}. BFCL multi_turn grading (multi_turn_checker.py) is state-based and cumulative: it penalizes (i) extra state-changing calls, (ii) missing required calls, and (iii) per-turn step-cap violations, but not read-only queries. Training-only home-buying variant. Training and in-domain validation use a renamed homebuying task variant of BFCL multi_turn with identical stateful tool APIs and turn structure; it is not held out. The Nemotron benchmark results use the unmodified BFCL v4 multi_turn benchmark. Training on the independently designed hospital-ward task yields held-out BFCL transfer (Table 5). Two-axis reward. The occurrence-local label for a sampled rollout is (Figure 5) ztool =

no_write( adec ) | {z }

procedural / behavior axis

× consequence(yrec ) , | {z } content / consequence axis

where the process gate is ( no_write( adec ) =

1, 0,

if the decision-turn action emits no state-changing write, otherwise,

and consequence(yrec ) ∈ [0, 1] checks whether the recovery-turn output contains the held call, matched by function name with lenient argument comparison. The no_write gate permits deferral and read-only queries (§4.1); consequence(yrec ) supplies the reward-only recovery signal. Approximation in the name-based no-write gate. The implemented name-based gate approximates “no state-changing write” over the tool set without verifying irreversibility, so it is not an exact read/write partition; memory uses an additive alternative (App. D.2). consequence grades the action

Trainable occurrence

Downstream event

receives all gradient

reward-only rollout

behavior axis: hard-gate statechanging writes

rollout continues

content axis: did the answer or held call survive? here: Recovery turn: canonical bridge re-adds the tool; must emit the held call

here: Decision turn: needed tool withheld; no state-changing write

Figure 5. Action/consequence decoupling in the Nemotron configuration: the decision-turn occurrence receives gradient, while reward-only recovery supplies its consequence signal. Note. This figure instantiates Eq. (1); category-specific training and the precomputed keyword proxy appear in App. C.3 and App. D.2.

20

Critical-State RL

D.1.2 Results and evidence scope Table 5. Nemotron missing-function application: BFCL miss_func accuracy under two training sources. Renamed-domain results summarize repeated deterministic evaluations (n=3). Training source Renamed home-buying domain Hospital-ward

Start

RL

Change

0.408 ± 0.013 0.425

0.452 ± 0.021 0.475

+4.4pt +5.0pt

The hospital-ward result measures transfer from an independently designed training task to unmodified BFCL. Evaluation repetition and serving variability are reported in App. C.3. Here turn 0 is the initial decision occurrence in an episode; the turn-0-augmented run adds missing-function training rows at this position. A size-matched control did not support turn-0-specific coverage, so the application result does not rely on that explanation. D.2 Memory applications: design and results D.2.1 xLAM memory application Scenario and selected occurrence. When memory is full, an attempted core_memory_add fails. The agent must delete a safely removable entry before adding the new one, rather than repeat the failed add or use core_memory_clear/archival_memory_clear, a destructive clear that erases downstream information. Using the xLAM start on BFCL v4 agentic/memory, occurrence-local RL places gradient only on that memory-storage response. Candidate membership and the keys that must be preserved or may be removed are supplied by metadata; the surrounding history is fixed context rather than part of the gradient-bearing trajectory. Action/consequence example: negation loss. Figure 6 traces one observed storage-to-retrieval failure (record 69). STORAGE TIME candidate occurrence Trigger: memory-full write failure [user] “success in finance means nothing if you sacrifice the relationships that matter most.” [store] core_memory_add(. . . ) → "values relationships over financial success" Absolute negation dropped

RETRIEVAL TIME

two turns later Downstream consequence [user] “how much does success in finance mean if I sacrifice the relationships that matter the most?” [ans] correct: Nothing model: “It means very little, if anything . . . ” Wrong consequence

Failure. Storage omits “means nothing”; the later answer weakens the original absolute negation to “very little.” This trace illustrates the action/consequence gap in Fig. 5.

Figure 6. Action/consequence trace for a memory-storage occurrence. Note. The precomputed keyword proxy comes from the same record 69. This trace illustrates the action/consequence split in Fig. 5.

Occurrence-local label. The reward combines correct-sequence behavior with a precomputed GPT-4o content proxy: 2–5 required terms, including meaningful negations, matched against storage text. Training scores this content locally; BFCL evaluates the full store→retrieve→answer chain. Reward definition. The additive storage reward grants partial credit: total = keyword_match_score + 0.1 · memory_type_score + behavior_bonus + memory-full bonuses, | {z } base

21

Critical-State RL

Here base is the keyword-and-memory-type subtotal. The total is clipped to [−1, 1], with behavior_bonus awarding +0.1 per correct procedural element (e.g. replace, archival add, remove-before-add) and memoryfull bonuses adding a +0.2 result term and a ±0.2 trash-removal term. Hard gates override this sum: a destructive memory clear returns −1.0, and a memory-full naive add-without-remove returns −0.5.

D.2.2 Gemma storage-and-retrieval application The Gemma application uses a retuned storage+retrieval mixture. Storage rows reuse the memory-full candidate, additive reward, and keyword proxy above. Retrieval rows select archival_memory_search or list_keys, graded by expected-query fact coverage. Each training unit updates one selected storage or retrieval response. D.2.3 Results Table 6. BFCL v4 agentic/memory results (%): start → memory-trained checkpoint.

memory sub-tasks Application

Key–value store

Vector store

Recursive summary

Sub-task mean

xLAM Gemma-4-26B-A4B

11.48 → 41.94 36.8 → 47.1

43.35 → 45.16 28.4 → 56.8

48.77 → 64.52 49.0 → 54.2

34.54 → 50.54 38.1 → 52.7

Note. The sub-task mean is unweighted and is not BFCL weighted full-suite Overall. The xLAM starting values average 10 memory-only evaluation runs, with population standard deviation 1.93 for the sub-task mean; trained values are single evaluations of selected checkpoints. The Gemma trained result uses step 20 from the highest-memory mixture among five data-mixture runs.

The largest gain is in key–value storage for xLAM and vector storage for Gemma (Table 6).

E Extended analysis E.1 Repeated occurrences and local credit Recurrence result. Broadcasting one terminal advantage across independent occurrences preserves each occurrence’s expected gradient but adds an exact covariance penalty from the other outcomes. Under equal variance, this penalty grows linearly with K; occurrence-local credit removes it (Lemmas 1–2). The resulting per-occurrence sample-cost ratio is explicit in Corollary 2. Even the optimal occurrence-indexed constant baseline leaves the penalty unchanged (Proposition 3). Relation to the applications. The K > 1 model below studies repeated occurrences within an episode, with controlled simulations in App. E.1.4. Missing-function and memory applications use K =1; logged repeat-call training uses single-step examples from flagged trace positions. App. B examines missing-function label–return associations. The field-survey analysis studies reward aggregation across repeated evaluations of one fixed choice (App. E.1.5).

E.1.1 Model and estimators for a recurrent critical state Fix K occurrence contexts i = 1, . . . , K of an abstract critical state σ, independent of the episode’s sampled outcomes. At occurrence i, policy πθ emits ai ∈ {1, . . . , m}; the correct action ci may vary with i, so σ alone need not identify the required action. For example, σ may mean “memory is full,” while ci identifies the safe entry to drop this time. Let local correctness be ri = 1[ ai = ci ], with qi = Prθ ( ai = ci ) and s2i = qi (1 − qi ) = Var(ri ); given θ, occurrences are conditionally independent. Only a terminal scalar is observed: R = ∑iK=1 ri , so the return-to-go at every occurrence is Gi = R.

22

Critical-State RL

Let gi = ∇θ log πθ ( ai | σ, i ) be the occurrence-i score; the phase-plus-index argument represents the local context, so the policy can distinguish occurrences. We have E[ gi ] = 0. Compare the two per-occurrence gradient contributions ĝ GRPO = A gi , {z |i

A = R−b }

vs.

one scalar advantage, broadcast to all K turns

ĝ flat = (ri − qi ) gi , |i {z }

per-occurrence local label

where b = E[ R] = ∑ j q j is the idealized population analogue of GRPO’s group mean; all baselines are stop-gradient. Write Aflat := ri − qi and AGRPO := A = R − b. This is the independent-reward i i specialization of factored policy-gradient variance reduction [42], using the standard likelihood-ratio and baseline identities [35, 36, 37]. The equations use unnormalized advantages and omit finite-group self√ inclusion. Population-standard-deviation normalization, proportional to K for shared returns under equal variance, scales signal and noise equally and preserves the gradient-SNR ratios. The results concern one occurrence’s contribution; covariance between contributions also enters the gradient summed over shared parameters.

E.1.2 Exact gradient covariance cost of shared credit Decompose the shared advantage into occurrence i’s own signal plus a zero-mean contribution from the other occurrences, A = R − E[ R ] = (r i − q i ) + ξ i ,

ξ i : = ∑ (r j − q j ),

E[ξ i ] = 0, ξ i ⊥ ( gi , ri ).

j ̸ =i

With qi = 12 at every occurrence, local-label credit for i uses Aflat = ri − 12 ∈ {± 12 }, whereas shared GRPO i uses A = ∑ j (r j − 12 ). At K =2, (ri =1, r j =0) gives A = 12 − 21 = 0; at K =3, missing both others gives A = 12 − 12 − 21 < 0. The extra contribution has zero mean (Lemma 1), but the variance of ξ i grows with K (Lemma 2). Lemma 1 (Same expected gradient under the recurrence model). E[ ĝiGRPO ] = E[ ĝiflat ] = Cov( gi , ri ) =: µi , independent of K when the occurrence policy and reward distribution are held fixed. Proof. E[ gi A] = E[ gi (ri − qi )] + E[ gi ξ i ]. By conditional independence ξ i ⊥ gi and E[ gi ] = 0, so E[ gi ξ i ] = E[ gi ] E[ξ i ] = 0; and E[ gi (ri − qi )] = E[ gi ri ] − qi E[ gi ] = Cov( gi , ri ), which is exactly E[ ĝiflat ]. Lemma 2 (Exact covariance penalty from other occurrences). Let Fi = E[ gi gi⊤ ] and v−i = ∑ j̸=i s2j . Then   Cov ĝiGRPO = Cov ĝiflat + v−i Fi .

(3)

The added covariance is positive semidefinite. For the scalar multipliers,    Var Aflat = s2i , Var AGRPO = ∑Kj=1 s2j = Ks2 under s j ≡ s . i i Proof. Write ĝiGRPO = ĝiflat + ξ i gi . Independence and E[ξ i ] = 0 make the cross covariance zero, while Cov(ξ i gi ) = E[ξ i2 ]E[ gi gi⊤ ] = v−i Fi . The scalar identity follows by adding independent reward variances. Corollary 2 (Per-occurrence SNR and fixed-precision sample cost). Use trace covariance for vector variance. Set vi = Var( ĝiflat ) and f i = E∥ gi ∥2 . For vi > 0, µi ̸= 0, and independent sampled episodes, the gradient SNR p ∥µi ∥/ Var( ĝi ) and sample counts for equal mean-squared error of the gradient sample mean satisfy SNRGRPO i SNRflat i

r

=

vi , vi + v −i f i

nGRPO v f = 1 + −i i . nflat vi

Under equal positive reward variance, v−i = (K − 1)s2 . Holding the occurrence policy fixed gives gradient SNR Θ(K −1/2 ) and sample-cost ratio Θ √(K ). The scalar advantage variances have the exact ratio K, hence the advantagemultiplier SNR ratio is exactly 1/ K. Equation (3) also covers vi = 0, where a relative sample-cost ratio is undefined. 23

Critical-State RL

For a trainable occurrence (App. B), the extra estimation cost depends on sibling variance v−i and score magnitude f i , not occurrence count alone (Lemmas 1 and 2). Deterministic sibling outcomes add no penalty. Local credit removes this return noise ξ i ; a phase-only baseline retains it (App. E.1.3). App. E.2 separately analyzes changes to the surrounding policy after the update.

E.1.3 The residual penalty after optimizing the baseline Proposition 3 (Optimal constant baselines leave the covariance penalty). At a fixed model context (σ, i ), let bi be any constant baseline, including a phase-only b(σ ), and set β i = bi − ∑ j̸=i q j . Then   Var ( R − bi ) gi = Var (ri − β i ) gi + v−i f i , so even optimizing the baseline leaves the exact penalty:   min Var ( R − bi ) gi = min Var (ri − β i ) gi + v−i f i . bi

βi

For f i > 0, the local minimizer is β∗i = E[ri ∥ gi ∥2 ]/ f i . Proof. The decomposition ( R − bi ) gi = (ri − β i ) gi + ξ i gi gives the same zero cross covariance as Lemma 2. The penalty does not depend on bi . Differentiating E[(ri − β i )2 ∥ gi ∥2 ] yields the minimizer, since the estimator mean is baseline-invariant.

Phase-only baselines. For several phase types, a phase-specific baseline b(σ) centers each phase separately. Within a recurring phase it is still constant across occurrences, so Proposition 3 leaves the sibling-noise penalty intact. The constant-baseline result complements methods that use additional information or control variates: future-conditional baselines [16], hindsight credit [15], and action-dependent baselines [46]. Potential shaping [43], TD(λ)/GAE [44, 45], and learned return decomposition [13, 14] provide other credit-assignment mechanisms.

E.1.4 Controlled simulation: local versus shared credit Figure 7a–b and Table 7 hold representation fixed and vary credit; panel (c) compares value representations. Per-turn advantage variance is K s2 (ratio = K, Lemma 2). Shared-return GRPO reaches high accuracy with a 10× budget. Measured episodes-to-threshold grow 240 → 1312 over K =1 → 32. The σ-keyed value is flat (Var[V ] = 0, R2 = 0.00), versus the acyclic reference (Var[V ] = 1.31, R2 = 0.54). Even with homogeneous correct actions, local-label convergence is K-constant while shared-return GRPO slows (14 → 56 update-steps over K =1 → 16); Lemma 2 gives the variance component of this credit-source difference.

State representation. Figure 7c plots the expected remaining local-label sum, V (si ) = E[∑Kj=i r j ], rather than the value of the terminal-only payment R. Distinct occurrence states retain this progress information (R2 = 0.54); pooling them into a phase-only value gives R2 = 0. This representation comparison is separate from the shared-credit covariance penalty, which holds with occurrence-distinguishable policies.

24

8 6 4 2 0

1

2

4

8

16 32

occurrences K

(b) fixed-budget per-occurrence accuracy 1 0.9 0.8 0.7 0.6 1

0.25K shared-return GRPO local-label credit

2

4

8

16 32

occurrences K

remaining-label value V

(a) shared-return variance: ×K

per-occurrence accuracy

per-turn advantage variance

Critical-State RL

(c) remaining-label value V: distinct vs. pooled states 4 3 2 1 0

1 2 3 4 5 6 7 8

occurrence index i distinct states shared state (σ-keyed)

shared-return GRPO local-label credit

Figure 7. Recurrence simulations: shared-return versus local-label credit in variance and fixed-budget accuracy, alongside a remaining-label value comparison. Note. Toy model with a recurrent critical state (heterogeneous occurrences, terminal reward R = ∑i ri ; m=5, group size 16, 30 seeds). (a) The theoretical shared-return variance follows the dashed 0.25K line; sampled estimates track it while local-label variance stays constant (Lemma 2). (b) With a 640-episode budget, shared-return accuracy falls with K while local-label accuracy stays flat across K. (c) Expected remaining local-label values decrease across distinct occurrence states; their phase-pooled value is constant.

Table 7. Controlled recurrence simulation by credit source and sample budget. Metric

Per-turn advantage variance

Per-occurrence accuracy 640-episode reference budget

Episodes to 0.9 accuracy

critical-state occurrences K

Credit source and budget

Shared-return GRPO Local-label credit

1

2

4

8

16

32

0.25

0.50

1.00

2.00

4.00

7.97

0.25

0.25

0.25

0.25

0.25

0.25

Shared-return GRPO Local-label credit Shared-return GRPO (10×)

0.982 0.975 0.956 0.901 0.787 0.627 0.982 0.983 0.984 0.984 0.984 0.984 0.999 0.999 0.999 0.999 0.998 0.997

Shared-return GRPO Local-label credit

240 240

336 232

448 240

648 240

920 240

1312 240

Note. Heterogeneous occurrences with terminal reward R = ∑i ri , m=5 actions, group size 16, and 30 seeds. Shared-return GRPO broadcasts one scalar advantage to all K turns; local-label credit uses each occurrence’s own label. Advantage-variance rows use a fixed Bernoulli policy (q= 12 , so s2 =0.25); accuracy and episode rows are trained outcomes. The reference budget is “640-ep”; “10×” uses a 6400-ep budget. The factor K is exact analytically (Lemma 2); the K =32 entry (7.97 vs. 8.0) is finite-sample.

25

Critical-State RL

Conditions for the variance penalty. The homogeneous control confirms that different correct actions are unnecessary for the penalty (App. E.1.4). Occurrence-local credit requires an observable label ri within the bounded segment, possibly through reward-only continuation (occurrence-local RL + task-specific reward). This label supplies the information that the shared scalar omits.

E.1.5 Field-survey diagnostic: terminal versus episode return The base no-think Gemma-4-26B-A4B policy generates complete episodes in a synthetic field-survey task. The agent selects a loadout before the terrain majority is sampled, then keeps it through K field turns that evaluate its consequences. Correctness denotes whether the loadout matches that terrain majority. Terminal R omits a direct penalty for a mismatched loadout; full-episode return-to-go G0 sums per-turn rewards, including loadout-dependent scan outcomes. For either reward quantity, SNR is the absolute correct–wrong mean difference divided by that quantity’s standard deviation. Terminal saturation. On 29,982 complete field-survey trajectories, terminal R conditioned on loadout correctness is E[ R | correct] = 0.930 vs. E[ R | wrong] = 0.951 (|∆R| = 0.021, SNR = 0.16). Across field-turncount bins K, |∆R| shrinks from 0.016 at K =1 to about 10−4 for K =7–10 and K =11–20, with both classes having R near 0.999. Thus terminal discriminability erodes as K grows. Full-episode discrimination. The full-episode return-to-go remains discriminative: E[ G0 | correct] = 3.26 vs. 1.90 (SNR = 0.60, about 3.7 times that of terminal R). The episode return includes intermediate scan rewards, whereas the additive-independent credit comparison in Lemma 1 and Proposition 3 assigns Gi = R. Mixed subtask outcomes. 99.9% of trajectories with ≥ 2 field turns contain both successes and failures in at least one subtask: chore completion or scan matching.

E.2 Drift outside the selected occurrence The exact frozen surrogate evaluates an updated action distribution at the selected call while holding the surrounding policy fixed. A full rollout with the updated model can also change surrounding calls. This outside-occurrence drift controls the mismatch between the exact surrogate and the full-rollout objective. The argument uses the standard KL chain rule and Pinsker inequality, with the surrogate comparison familiar from trust-region analysis [34].

E.2.1 The exact frozen surrogate is a hybrid-policy objective Let Pθ be the distribution of finite-horizon trajectories under πθ , with the same initial distribution and environment for every policy, and let | R(τ )| ≤ Rmax . A fixed rule selects one occurrence t⋆ from the history available before its action. A trajectory without that occurrence is padded with a policy-independent terminal no-op. The updated action distributions are absolutely continuous with respect to the anchor πθ0 at relevant histories. hyb Define a hybrid policy that uses πθ at t⋆ and πθ0 everywhere else; write its trajectory law as Pθ . Let x record the selected pre-action history, including its time and any terminal-padding marker. It has the anchor distribution d0 , and its continuation has the anchor return value   Qθ0 ( x, a) = E R | x, a, continue under πθ0 . The exact frozen surrogate is therefore the hybrid policy’s expected return: Jhyb (θ ) = E hyb [ R] = Ex∼d0 , a∼πθ (·| x) [ Qθ0 ( x, a)].

(4)

Pθ

b write The full-rollout objective is J (θ ) = EPθ [ R]. Both equal J (θ0 ) at the anchor. For a frozen proxy Q, b ]. For each objective, ∆ denotes its change from θ0 . These comparisons hold within the same b J (θ ) = Ed0 ,πθ [ Q trajectory task and reward definition.

E.2.2 Only outside-occurrence drift enters the mismatch bound With ht denoting the pre-action history, define " Dout (θ ) := Eτ ∼ Pθ

∑ KL πθ (· | ht ) ∥ πθ0 (· | ht )

t̸=t⋆

26

# 

.

Critical-State RL

Proposition 4 (Outside-occurrence drift bound). Under the preceding construction, hyb

Dout (θ ) = KL( Pθ ∥ Pθ

) ≤ KL( Pθ ∥ Pθ0 ),

and ∆J (θ ) − ∆Jhyb (θ ) ≤ Rmax

q

2Dout (θ ).

hyb

Proof. In the trajectory-density ratio Pθ /Pθ , the environment factors and selected-action factor cancel. The KL chain rule therefore leaves exactly the conditional KL terms outside t⋆ . The full KL to the anchor additionally includes the nonnegative selected-action term. Boundedness and Pinsker give q hyb | J (θ ) − Jhyb (θ )| ≤ 2Rmax TV( Pθ , Pθ ) ≤ Rmax 2Dout (θ ). The anchor values agree, so the same bound holds for their gains. The square-root order is tight: changing an unselected Bernoulli action from probability 1/2 to 1/2 + u, with return ± Rmax , gives mismatch 2Rmax |u| and Dout = 2u2 + O(u4 ).

Behavioral isolation. If behavior changes only at the selected occurrence, Dout = 0 and the exact frozen objective equals full-rollout return, regardless of the selected-action change. Gradient masking alone does not enforce this: shared parameters can change other calls.

When a local update improves the full rollout. The score projection in App. B sharpens this comparison near the anchor. Consider a smooth policy on a finite trajectory tree with positive action probabilities. All expectations below use Pθ0 . Write gt = ∇θ log πθ ( at | ht )|θ0 , s = gt⋆ , and o = ∑t̸=t⋆ gt , with zero score for the terminal no-op. Define µsel = E[s( R − ER)],

Fsel = E[ss⊤ ],

Fout = E[oo ⊤ ],

R ⊤ † Vsel = µsel Fsel µsel .

R is the return variance explained by the selected score, using the same Here µsel = ∇ Jhyb (θ0 ) and Vsel projection as App. B with terminal return in place of the local label.

Proposition 5 (Selected-return signal and outside-call drift). For every parameter direction u, q  ⊤ R u⊤ F u. Var( R) − Vsel ∇ J (θ0 )⊤ u ≥ µsel u− out A strictly positive right-hand side guarantees improved return for all sufficiently small positive steps along u. Proof. Predictable selection and zero-mean conditional scores give E[so ⊤ ] = 0 and Fout = E ∑t̸=t⋆ gt gt⊤ . † µ 2 R Project centered return onto s: the residual e = R − ER − s⊤ Fsel sel has Ee = Var( R ) − Vsel and E[ oe ] = E[oR]. The policy-gradient identity [35] gives ∇ J (θ0 ) = µsel + E[oe]; Cauchy–Schwarz proves the bound. The same matrix controls infinitesimal drift: Dout (θ0 + tu) = t2 u⊤ Fout u/2 + o (t2 ). Only residual return R = Var( R ), that effect variation contributes to the first-order return effect of outside-call changes. If Vsel vanishes even when Fout is large. Label-based credit connects to µsel through the calibrated link in Theorem 2. b − Qθ ∥∞ ≤ ε 0 and define the selected-action KL Proxy error and selected-action drift. Suppose ∥ Q 0 under anchor contexts

hyb

0 Dsel (θ ) := Ex∼d0 KL(πθ (· | x ) ∥ πθ0 (· | x )) = KL( Pθ

Then ∆J (θ ) ≥ ∆ b J (θ ) − Rmax

q

∥ Pθ0 ).

 q  0 2Dout (θ ) − 2ε 0 min 1, Dsel (θ )/2 . 27

(5)

Critical-State RL

R b − Qθ , the difference of proxy and exact surrogate gains is Ed (πθ − πθ )e da. Its Indeed, for e = Q 0 0 0 magnitude is at most 2ε 0 Ed0 TV(πθ , πθ0 ); Pinsker and Jensen give (5). Both error terms vanish at the anchor. 0 follows The two KL terms use different context distributions: Dout follows current rollouts, whereas Dsel frozen contexts. Here ε 0 concerns a return-value proxy; the label-to-return scale λ( x ) in Theorem 2 is a separate quantity.

E.2.3 Optimizing the drift-adjusted gain bound

√ For the exact surrogate, parameterize an interval of increasing outside-occurrence drift by δ = Dout , and √ let p(δ) = ∆Jhyb . Proposition 4 gives the gain lower bound G (δ) = p(δ) − κδ, where κ = 2Rmax . Corollary 3 (Optimum of the drift-adjusted lower bound). If p is continuously differentiable, increasing, and strictly concave on [0, δmax ], with p′ (0) > κ > p′ (δmax ), then G has a unique interior maximizer satisfying p′ (δ⋆ ) = κ. The surrogate continues increasing beyond δ⋆ , while its drift-adjusted lower bound decreases. This follows from the sign change of the strictly decreasing derivative G ′ = p′ − κ.

Measurement and checkpoint selection. Estimating Dout requires current-policy full rollouts and summed policy KL outside the selected occurrence. A token-averaged KL on training responses is not that quantity. Evaluate the target multi-turn metric on held-out data for checkpoint selection; where the required KL and proxy-error quantities are available, (5) additionally gives a drift-adjusted gain criterion. Refreshing contexts and continuation values defines a new anchor for the same comparison.

E.2.4 A stylized stopping example J (θ ) = 1 − e−θ and sets J (θ ) = b The closed-form toy trains scalar θ on surrogate b J (θ ) − B(θ ), with B(θ ) = c θ √ and KL ∝ (θ − θ0 ). This stipulated linear penalty illustrates marginal crossing; it is not a fitted outsideoccurrence KL envelope. Training sees only b J, although the true peak is θ ⋆ = − ln c (p′ = B′ ). At c = 0.15, b J rises 0 → 0.847 at step 10→ 0.969 at step 60, whereas true J peaks at 0.565 at step 10 (θ = 1.874, near θ ⋆ = − ln 0.15 = 1.897) and falls to 0.449 by step 60 (KL 0.040 → 0.135). With patience-4, evaluation noise ϵeval = 0.02, and 400 seeds, true-metric stopping averages step 9.2 (KL 0.036, reward 0.560); surrogate stopping averages step 19.0 (KL 0.064, reward 0.547). The true peak, 0.565, is about 26% higher than the step 60 reward, 0.449. Raising c from 0.08 → 0.45 moves θ ⋆ : 2.53 → 0.80 (peak step 22 → 2), following the same marginal-crossing calculation as Cor. 3. Figure 8a shows the turnover in the fixed-parameter toy.

E.2.5 Empirical trajectories: transfer turnover and aggregate saturation Transfer turnover (Nemotron-Super-120B, turn-0-augmented run). The highest observed single-

eval BFCL miss_func transfer point is step 15 (0.545, generation-KL about 0.0058). On the 2450-sample home-buying set, decision-turn validation instead rises 0.602 → 0.718 to its sampled maximum at step 35. Later transfer checkpoints remain below the step-15 observation as KL generally rises (Figure 8b). The observed transfer peak therefore precedes the in-domain maximum, illustrating why the two metrics lead to different checkpoint choices.

Aggregate saturation (no-think Gemma-4-26B-A4B category-routed run). This run is separate from

the four-cell study. In-domain validation continues to improve during miss_func training (App. C.3), but official BFCL weighted full-suite Overall saturates at 59.2/60.2/59.6/59.4/60.1% over steps 25/30/35/40/45. Training beyond roughly step 30 brings no aggregate gain, while the untargeted base and long_context categories fall 0.72 → 0.68 and 0.61 → 0.53, respectively, over steps 30 → 45. The narrower three-category multi_turn mean therefore turns over (about 0.593 at step 25, falling to 0.556 by step 45). Because per-step generation-KL is tiny/noisy at learning rate 10−6 , the in-domain accuracy climb itself (0.40 → 0.80) is the clearer measure of training progress.

28

Critical-State RL

surrogate objective b J true objective J

in-domain validation BFCL missing-function transfer (miss_func)

(a) toy: surrogate vs. true objective

(b) Nemotron: validation vs. transfer 0.7

accuracy

objective

1 true peak

0.5

highest observed single-eval point

0.6 0.5

0

0

10

20

30

40

50

60

10

20

30

40

training step

training step

Figure 8. Checkpoint-selection trajectories in the stylized toy and Nemotron training run. Note. (a) In the stylized toy, b J rises monotonically while J turns over. (b) The highest observed transfer point in the Nemotron turn-0-augmented run precedes the in-domain validation maximum. These curves illustrate checkpoint-selection behavior; they do not estimate Dout or calibrate the bound.

F Prompts and canonical strings Prompt inventory. The Nemotron-Super-120B and xLAM blocks below are verbatim, including {placeholder} tokens.

F.1 Nemotron tool-use prompts Recovery bridge prompt (tool re-availability signal). At the decision turn k−1 the required tool is withheld; at the recovery turn k, training and evaluation use BFCL’s canonical tool-reavailability bridge, DEFAULT_USER_PROMPT_FOR_ADDITIONAL_FUNCTION_FC.

Tool-availability judge (turn-0 rows). The judge excludes turn-0 requests labeled YES, which indicates that the available tools can fulfill the request, and retains requests judged to require the withheld tool. _JUDGE_SYSTEM = """You are a careful tool-availability judge for a multi-turn function-calling assistant.

Decide whether the user's request can be reasonably fulfilled by ANY combination of the listed tools. Be strict ,→ answer YES only if a clearly applicable tool exists. Reply with exactly one word: YES or NO.""" prompt = [ {"role": "system", "content": _JUDGE_SYSTEM}, {"role": "user", "content": ( f"User request:\n {user_msg}\n\n" f"Available tools:\n{_tools_summary(tools)}\n\n" f"Can the user's request be fulfilled? Answer YES or NO." )}, ]

29

Critical-State RL

Pre-bridge assistant messages. The refusal_pre_inject variants defer or ask about tools at a decision turn with no ground-truth call. The clarification variants ask for missing parameters.

PRIOR_ASSISTANT_VARIANTS = { "refusal_pre_inject": [ "I don't have a tool that can do that right now - could you enable a function that handles it, or should I keep ,→ looking for another way?", "None of my currently available tools can do that. Would you like to make a suitable tool available, or shall I ,→ try to work around it differently?", "I'm not able to do that with the tools I have available. Could you add a tool for it, or should I explore an ,→ alternative approach?", ], "clarification": [ "I want to make sure I get this right - could you tell me the specific value you'd like me to use?", "Could you give me a bit more detail on that so I can fill in the right information?", "To do this accurately, which exact value should I use for that?", ], }

F.2 xLAM memory-storage prompts Shared memory-system instruction (training↔evaluation).

BACKEND_INSTRUCTION = ( "You have access to an advanced memory system, consisting of two memory " "types 'Core Memory' and 'Archival Memory'. Both type of memory is " "persistent across multiple conversations with the user, and can be " "accessed in a later interactions. You should actively manage your memory " "data to keep track of important information, ensure that it is up-to-date " "and easy to retrieve to provide personalized responses to the user later." "\n\n" "The Core memory is limited in size, but always visible to you in context. " "The Archival Memory has a much larger capacity, but will be held outside " "of your immediate context due to its size." ) def build_unified_system_message(memory_state: str) -> str: return f"{BACKEND_INSTRUCTION}\n\n{memory_state}"

Shared memory-full error messages (training=evaluation). The training data and BFCL evaluator use the same error JSON for failed memory-add calls (training↔evaluation). # Training data (memory-full trigger): messages.append({ "role": "tool", "content": '{"error": "Core memory is full. ,→ Please clear some entries."}', "name": "core_memory_add", "tool_call_id": f"call_{idx}_1", "tool_calls": None }) ... "content": '{"error": "Long term memory is full. ,→ Please clear some entries."}', "name": "archival_memory_add",

# BFCL eval checker: if len(self.core_memory) >= MAX_CORE_MEMORY_SIZE: return {"error": "Core memory is full. Please ,→ clear some entries."} ... return {"error": "Long term memory is full. ,→ Please clear some entries."}

30

Critical-State RL

Keyword-extraction prompt. GPT-4o extracts 2–5 required terms offline into extra_info.keywords. The prompt asks it to retain meaningful negations; RL scores storage text by deterministic exact matching. The system prompt and five examples follow. SYSTEM_PROMPT = """You are a keyword extraction assistant. Given a ground truth value for a memory entry, extract 2-5 ,→ key factual terms or phrases that MUST be present in any correct version of this entry. Focus on: - Proper nouns (names, places) - Numbers, dates, times - Key actions or states (e.g., "high cholesterol", "not married") - Essential descriptors that distinguish this entry Rules: - Return ONLY a JSON array of lowercase strings - Each keyword should be 1-3 words - For very short values (1-2 words), return the value itself as the only keyword - Avoid generic words like "user", "the", "is", "a" - Include negations when they are meaningful (e.g., "not" in "not married")""" FEW_SHOT_EXAMPLES = [ { "value": "blue", "keywords": '["blue"]', }, { "value": "Blood work 2025-11: cholesterol high (user reported). No numeric values provided. Recommend ,→ lifestyle changes and follow-up with PCP.", "keywords": '["blood work", "2025-11", "cholesterol high", "lifestyle changes", "pcp"]', }, { "value": "User likes art, particularly the impressionists and pointillism. However, user says they are not a ,→ good artist themselves.", "keywords": '["art", "impressionists", "pointillism", "not a good artist"]', }, { "value": "Router disconnects every 20 minutes; first reported by user.", "keywords": '["router", "disconnects", "20 minutes"]', }, { "value": "User describes themselves as imaginative.", "keywords": '["imaginative"]', }, ]

31

Record · ID 1028677 · SHA-256 db85907efa376316
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.