Conceptio › Archive › arXiv CS
arXiv CSopen access

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Preprint | Salesforce AI Research

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Yang Li, Semih Yavuz, Shafiq Joty

arXiv:2609.05295v1 [cs.AI] 4 Sep 2026

Salesforce AI Research {yli2,syavuz,sjoty}@salesforce.com On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while self-distillation with privileged conditioning is limited by in-context learning capacity. We propose RISE (Recursive Improvement via SelfExtrapolating Policy Distillation), which constructs a synthetic teacher directly from the model’s own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor—in parameter space or output logit space—RISE converts a sparse outcome-induced parameter update into a dense token-level target, without any external model or privileged conditioning. RISE combines RLVR and OPD in a complementary loop: outcome rewards ground the extrapolation toward correct reasoning, while the extrapolated teacher refines token-level decisions. Moreover, since the teacher is refreshed every iteration as the student improves, distillation becomes a recursive improvement mechanism rather than a one-shot compression step. Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.

1

Introduction

A central aspiration of artificial intelligence is to build systems that can recursively improve themselves—becoming increasingly capable from their own experience with minimal human supervision (Yang et al., 2026c; Tao et al., 2024). For large language models (LLMs), this means updating parameters from self-generated data via reinforcement learning from human feedback or verifiable rewards (RLHF/RLVR) (Ouyang et al., 2022; Shao et al., 2024; Guo et al., 2025; Yu et al., 2026), self-play (Chen et al., 2024; Liu et al., 2025; Huang et al., 2025; Fang et al., 2026), or rejection sampling fine-tuning (Zelikman et al., 2022; Gulcehre et al., 2023; Xiong et al., 2025). All these approaches share a common limitation: the learning signal is sequence-level—an outcome reward, a binary accept/reject label, or a scalar preference—providing no guidance on which tokens were responsible for success or failure. On-policy distillation (OPD) offers a richer signal for weight-based self-improvement. Rather than a single scalar per response, OPD provides dense per-token supervision: a teacher specifies a full next-token distribution at every position, telling the model not just whether its response was correct but how to improve each token-level decision (Song & Zheng, 2026). However, OPD’s promise hinges on a critical question: where does a reliable teacher come from? External teachers suffer from distribution mismatch (Li et al., 2026c; Zhu et al., 2026); on-policy self-distillation (OPSD) with privileged conditioning often fails to reflect token-level correctness, due to limited in-context learning (ICL) capability and uninformative privileged information (Zhao et al., 2026b; Hübotter et al., 2026; Li et al., 2026b); and subsequent heuristic remedies—token-level gating (Xu et al., 2026b; Lu et al., 2026), divergence mixing (Jung et al., 2025; Jin et al., 2026), DAgger-style sampling (Li et al., 2026a; Zhao et al., 2026a), trajectory refinement (Yang et al., 2026b)—treat symptoms rather than the root cause. All these approaches accept a flawed teacher as given; none addresses the fundamental question: what should the teacher be? We propose a shift in perspective. A natural candidate teacher for a policy πθ is its own converged future self —the policy π ∗ that training is converging toward. While π ∗ is unknown, the training trajectory reveals the direction of convergence. Let φ map a policy to a vector space where linear operations are meaningful (e.g., logits or parameters). The displacement φ(πθ′ ) − φ(πθ ) between the current checkpoint θ′ and an earlier anchor θ captures the direction

1

Preprint | Salesforce AI Research

(a) Extrapolation geometry.

(b) The RISE training loop.

′ Figure 1: RISE overview. (a) RLVR updates θn → θn+1 (blue); extrapolation amplifies this displacement to construct θfuture (red); OPD distills θfuture into the student, yielding θn+1 (green). (b) The training loop: RLVR grounds the direction, while OPD refines token-level decisions.

of recent improvement. We can extrapolate this displacement to synthesize a future teacher:  φ(πfuture ) = φ(πθ ) + β · φ(πθ′ ) − φ(πθ ) , β > 1.

(1)

When β = 1 the teacher equals the current checkpoint (no distillation signal); when β > 1 it amplifies the model’s most recent update, projecting beyond its current state along the training trajectory. Recent analyses reveal that post-training updates are dominated by a low-rank subspace and evolve near-linearly (Cai et al., 2025; Wang et al., 2026a), making extrapolation along the training direction a principled approximation (§2). This construction eliminates the core pathologies of prior OPD: reduced distribution mismatch (the teacher extrapolates the student’s own optimization path, so it remains nearby in weight and distribution space), no privileged conditioning, and no external model—only previous checkpoints, naturally available during training. The construction above is iterative by design: each distillation step refines the current policy, which in turn provides a fresh displacement for the next extrapolation. For this recursion to converge, the extrapolation direction must point toward genuine improvement—yet extrapolation itself is direction-agnostic, amplifying whatever update the model made regardless of quality. We resolve this by using RLVR to produce the update θ → θ′ : outcome rewards ground the displacement in verified improvement. On-policy distillation then applies the extrapolated teacher to refine θ′ , providing the per-token supervision that outcome rewards alone cannot. The two signals are complementary: RLVR ensures the direction is meaningful; OPD ensures the refinement is fine-grained (Fig. 1). Our contributions are threefold. (1) We identify teacher quality as the central bottleneck of OPD and propose RISE, which constructs the teacher by extrapolating the model’s own RLVR trajectory, requiring no external model and no privileged context. Because the teacher is refreshed every iteration from the student’s latest update, distillation becomes a recursive improvement loop rather than a one-shot compression step (§3). (2) We unify this construction under a representation map φ and instantiate it in two spaces—logit-space (geometric mixture of output distributions) and weight-space (task arithmetic)—characterizing their computational and statistical trade-offs and ablating both (§3, §4). (3) We show that RISE outperforms RLVR and OPSD baselines across mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks, with improved sample efficiency, at no additional sampling cost and a modest 1.3–1.6× wall-time overhead (§4).

2

Background and Related Work

2.1

Reinforcement Learning with Verifiable Rewards

RLVR optimizes a language model policy πθ to maximize expected reward on tasks with objectively checkable outputs. Given a prompt x, the model generates a response y = (y1 , . . . , yT ) ∼ πθ (· | x) and receives an 2

Preprint | Salesforce AI Research

outcome reward R(x, y) ∈ [0, 1]. Policy gradient algorithms such as GRPO (Guo et al., 2025) and DAPO (Yu et al., 2026) estimate per-token advantages from group-normalized outcome rewards and update the policy via clipped surrogate objectives: hP  i T t |x,y<t ) (2) LPG = −E , ρt = ππoldθ (y t=1 Ât · min ρt , clip(ρt , 1−ϵ, 1+ϵ) (yt |x,y<t ) , where Ât is the advantage estimate (constant across tokens within a response in GRPO/DAPO). While effective, the outcome-level reward assigns identical advantage to every token in a response, creating a credit assignment bottleneck: the gradient signal cannot distinguish helpful reasoning steps from irrelevant or harmful ones. 2.2

On-policy Distillation

On-policy distillation (OPD) augments policy gradient training with per-token distributional supervision from a teacher model πT . At each position t, the student minimizes a divergence between its distribution and the teacher’s: hP i T LOPD = Ey∼πθ (3) t=1 DKL (πθ (· | x, y<t ) ∥ πT (· | x, y<t )) . Unlike the outcome reward, this loss provides a rich gradient at every token: the teacher’s full distribution over the vocabulary specifies not only which token is best, but the relative desirability of all alternatives. In practice, computing the divergence over the full vocabulary is expensive, so most implementations resort to either a top-K truncation or a REINFORCE-style sample-based estimator (Lu & Lab, 2025; Agarwal et al., 2024). The critical challenge is the choice of teacher πT . Existing approaches fall into two categories, each with fundamental limitations. External teacher OPD. The most direct approach uses a separate, stronger model as πT —for example, distilling from a larger model in the same family (Xiaomi, 2025; Zeng et al., 2026). The teacher evaluates πT (· | x, y<t ) conditioned on the student’s generated prefix y<t . As training progresses, the student explores increasingly diverse reasoning paths, many unseen by the teacher. The teacher’s distribution conditioned on such out-of-distribution prefixes becomes unreliable: it reflects how the teacher would continue from an alien context, not whether the student’s reasoning is sound. This prefix distribution mismatch degrades teacher quality precisely when the student needs guidance most—at novel or challenging reasoning steps (Li et al., 2026c; Zhu et al., 2026). On-policy self-distillation (OPSD). An alternative formulation, on-policy self-distillation (OPSD), uses the same base model as both teacher and student, conditioning the teacher on privileged information such as correct solutions or environmental feedback (Zhao et al., 2026b; Hübotter et al., 2026). The teacher distribution becomes πT (· | x, C, y<t ), where C is the privileged context. The hope is that ICL enables the model to produce a more informed distribution when conditioned on C. However, this approach has its own limitations. First, the model’s ICL capability may be insufficient to meaningfully incorporate the privileged context (Li et al., 2026b). Second, the teacher and student see different inputs, creating a different form of distribution mismatch: the teacher’s distribution may reflect the privileged context in ways that are not transferable or even harmful to the student’s unconditioned setting (Kim et al., 2026; Harne et al., 2026). Heuristic Remedies. Recognizing these limitations, recent work has proposed several heuristic strategies. Tokenlevel gating selectively weights the per-token loss using the teacher–student probability gap (Xu et al., 2026b; Lu et al., 2026; Xing et al., 2026), token entropy (Xu et al., 2026b; Ko et al., 2026), or teacher confidence (Liu et al., 2026). Divergence mixing interpolates between forward and reverse KL or selects adaptively per token, trading mode coverage against mode seeking (Hübotter et al., 2026; Jung et al., 2025; Jin et al., 2026; Jia et al., 2026; Jang et al., 2026). DAgger-style rollout mixing (Ross et al., 2011) interleaves teacher and student tokens during generation to keep prefixes on-distribution (Li et al., 2026a; Agrawal et al., 2026; Zhao et al., 2026a). Trajectory refinement refines rollouts before distillation so the teacher conditions on higher-quality prefixes (Yang et al., 2026b; Jiang et al., 2026). All these remedies accept a flawed teacher and engineer ways to tolerate its noise, rather than questioning whether a better teacher exists. In §3, we show that the model’s own training trajectory can be used to construct such a teacher.

3

Preprint | Salesforce AI Research

2.3

Task Arithmetic and Linear Trajectories

RISE builds on convergent empirical evidence that post-training trajectories are low-dimensional and approximately linear. In the model merging literature, task vectors τ = θft − θpre encode task-specific knowledge in a composable, approximately linear fashion (Ilharco et al., 2022), and fine-tuning trajectories exhibit approximate linear mode connectivity—checkpoints along the training path can be linearly interpolated without loss barriers (Frankle et al., 2020). The related model soups and weight averaging literature (Wortsman et al., 2022; Izmailov et al., 2018) demonstrates that uniform or stochastic averaging of checkpoints along (or near) the training trajectory consistently improves generalization, reinforcing that these trajectories lie in well-behaved, low-curvature regions of parameter space amenable to linear operations. More recently, analyses of RLVR training reveal that parameter updates are dominated by a rank-1 subspace whose projection coefficient evolves near-linearly throughout training (Cai et al., 2025; Wei et al., 2026), and that these low-rank weight dynamics propagate to linear evolution in output log-probabilities (Wang et al., 2026a). A parallel finding holds for OPD, whose cumulative updates rapidly lock into a narrow low-dimensional channel (Shen et al., 2026). Together, these properties justify extrapolation along the training direction: if θ0 → θn is approximately linear and confined to a low-dimensional subspace, then θ0 + β · (θn − θ0 ) with β > 1 is a principled estimate of a more capable future checkpoint. Whereas model soups exploit linearity through interpolation (β ∈ [0, 1]) for robustness, RISE exploits the same structure through extrapolation (β > 1) for capability improvement. RISE does not, however, require strict linearity: confinement to a low-dimensional subspace limits how far extrapolation can deviate from the true trajectory, and our own runs confirm this—three directions capture ∼87% of the variance (Appendix A.3). 2.4

Joint RLVR and OPD Training

RLVR and OPD provide complementary signals, and recent work seeks to combine them. The OPD gradient is dense but bounded by teacher quality, while the RLVR gradient is sparse but can drive the student beyond the teacher’s capability ceiling (Song & Zheng, 2026). The most direct approach augments the policy gradient objective with an auxiliary KL divergence term toward a teacher distribution (Xu et al., 2025; Ramos et al., 2026; Zhang et al., 2026). ExOPD (Yang et al., 2026d) and others (Wang et al., 2026b; Oh et al., 2026) formalize an equivalence between standard OPD and dense KL-constrained RL, where the log-probability ratio between teacher and student serves as a per-token reward. RLSD (Yang et al., 2026a) uses OPD to reweight the advantage from GRPO for finer-grained credit assignment. An alternative sequential strategy allocates RL to the teacher for capability discovery and OPD to the student for compression (Xu et al., 2026a). Whether combining with an external teacher or a privileged self-teacher, all these approaches take the teacher as given and focus on how to best integrate its signal with RL. RISE departs from this paradigm by eliminating the need for an external or privileged-conditioned teacher altogether: the teacher is constructed by extrapolating the model’s own training trajectory. Among these, ExOPD’s reward extrapolation resembles RISE structurally: its optimal policy is the same geometric mixture as RISE’s logit-space formulation (Eq. 6). The key distinction is not merely teacher source but teacher dynamics. ExOPD fixes a static external policy as teacher before distillation begins; no matter how well the student trains, the teacher gap J(π ∗ ) − J(πT ) in Theorem 2 remains constant—a hard ceiling. RISE’s teacher is non-stationary by construction: as the student improves via RLVR, the displacement vector updates, and the extrapolated teacher advances in lockstep. This eliminates the static teacher ceiling and converts OPD from a one-shot compression step into a recursive improvement loop. A direct empirical comparison is inapplicable because ExOPD requires an external stronger model—precisely what RISE eliminates.

3

Method

Theorem 2 (Appendix A) formalizes §1’s intuition: the optimal teacher for πθn is π ∗ , but since π ∗ is unknown, RISE approximates it by extrapolating the model’s own RLVR-grounded training trajectory, then distills that teacher’s per-token distribution into the student.

4

Preprint | Salesforce AI Research

3.1

Self-Extrapolated Policy Distillation

′ Consider a policy πθn updated by RLVR to πθn+1 . Let φ be a representation map into a vector space where linear operations are meaningful. The self-extrapolated teacher is defined as:  ′ φ(πfuture ) = φ(πθn ) + β · φ(πθn+1 ) − φ(πθn ) , β > 1. (4) ′ The extrapolation scale β controls how far beyond the current policy we project: β = 1 recovers πθn+1 , and β > 1 continues along the improvement direction. The choice of φ yields a family of instantiations with different computational and statistical properties.

Weight-space extrapolation (φ = θ). Setting φ to the identity on parameters gives ′ θfuture = θn + β · (θn+1 − θn ),

(5)

which is precisely task arithmetic (Ilharco et al., 2022) with an extrapolation coefficient. The teacher is then the model defined by θfuture . Logit-space extrapolation (φ = log π). Let st ≜ (x, y<t ) denote the context at position t. Setting φ to map each policy to its per-token output log-probabilities gives:  ′ log πfuture (· | st ) = log πθn (· | st ) + β · log πθn+1 (· | st ) − log πθn (· | st ) + const, (6) where the constant ensures normalization. Equivalently, πfuture ∝ πθ1−β · πθβ′ —a geometric mixture that amplifies n n+1 ′ the probability ratio πθn+1 /πθn . Remark 1 (Connection to KL regularization). Given the geometric mixture, the distillation loss decomposes as ′ DKL (πθ ∥πfuture ) = −(β − 1) DKL (πθ ∥πθn ) + β DKL (πθ ∥πθn+1 ) + log Z,

(7)

′ where Z is the normalization constant of πfuture , and both πθn and πθn+1 are treated as fixed (stop-gradient) when constructing πfuture . For β > 1, the first term has a negative coefficient and is repulsive—it pushes the current policy away from the anchor, continuing the RLVR improvement direction; the second term regularizes it toward the ′ post-RLVR checkpoint πθn+1 , preventing overshoot during OPD. Both terms operate at the token level, providing fine-grained credit assignment that a scalar outcome reward cannot.

Relationships between instantiations. Let f (θ) denote the mapping from parameters to output logits. Weightspace extrapolation computes f (θn + β · ∆θ), while logit-space computes f (θn ) + β · ∆f —the first-order Taylor approximation of weight-space extrapolation around θn . The two coincide exactly when f is linear; for neural networks they diverge. Weight-space produces a coherent model whose cross-position predictions are jointly consistent, at the cost of materializing parameters and one teacher forward pass per OPD step; logit-space needs no parameter manipulation and fixes the teacher before OPD begins. 3.2

Distillation Loss Design

Computing a divergence over the full vocabulary is prohibitively expensive. Following prior works (Hübotter et al., 2026; Zhao et al., 2026b), we use a top-K approximation. For logit-space extrapolation, the unbiased procedure would extrapolate over the full vocabulary and then select the top-K of πfuture ; however, πfuture does not exist as a model—it is defined only through the logit-space formula—so obtaining its full-vocabulary logits would require ′ keeping two extra model copies in memory or caching O(V ) logits per position. Instead, letting S = TopK (πθn+1 ), ′ we project both πθn+1 and πθn onto the same (K +1)-simplex (retaining log-probabilities at S and appending a tail bucket for the remaining mass, each renormalized), then apply Eq. 6 to all K +1 log-probabilities:  ′ log πfuture (v | ·) = log πθn (v | ·) + β · log πθn+1 (v | ·) − log πθn (v | ·) , v ∈ S ∪ {tail}. (8) With K = 100, the top-K tokens account for essentially all probability mass under typical LLM distributions, so any bias from using S instead of TopK (πfuture ) is negligible (Appendix A.4). For weight-space extrapolation, a 5

Preprint | Salesforce AI Research

forward pass through θfuture (Eq. 5) yields πfuture exactly, and we evaluate it at the student’s top-K indices with a tail bucket appended and renormalized. In both cases, πfuture is held fixed (stop-gradient) throughout all OPD gradient steps. The RISE distillation loss is: hP i T LRISE = Ey∼πθ D (π (· | s ) ∥ sg[π (· | s )]) , (9) t future t t=1 KL θ where the KL divergence is computed over the (K + 1)-dimensional simplex. Choice of divergence. Our analysis (Remark 1, Theorem 2) uses reverse KL for an exact trust-region decomposition. In practice we replace KL with the Jensen–Shannon divergence (JSD) in Eq. 9, which is bounded by log 2 thus avoiding the numerical instability of reverse KL. Since JSD(P ∥Q) ≤ 12 DKL (P ∥Q), the trust-region guarantees from the KL analysis carry over as conservative bounds. The two losses also share the same unique minimizer (πθ = πfuture ), so the teacher targeted is identical. 3.3

Training Procedure

Each RISE iteration alternates between two phases: 1. RLVR phase. Sample rollouts from the current policy πθn , compute outcome rewards, and update the policy via ′ a policy gradient method (e.g., GRPO) to obtain θn+1 . ′ 2. OPD phase. Construct πfuture from πθn+1 and anchor πθn via logit-space (Eq. 6) or weight-space (Eq. 5) ′ extrapolation. Distill πfuture into πθn+1 by minimizing LRISE (Eq. 9), yielding θn+1 .

The two phases are complementary: RLVR discovers capability improvements via sparse outcome rewards, while OPD compresses the extrapolated teacher’s token-level distribution into the current policy (Algorithms 1 and 2). Both phases operate on the same set of rollouts—the responses y ∼ πθn sampled for RLVR are reused for OPD, which therefore adds no additional sampling cost. Reusing pre-RLVR rollouts introduces mild off-policy-ness (contexts come from pre-RLVR), but the distributional loss is well-defined for any input sequence without importance sampling correction. We confirm empirically that this shift has negligible impact (§4.4). Algorithm 1 gives the full pseudocode for logit-space extrapolation, where the teacher distribution is built by extrapolating cached top-K logits from the post-RLVR checkpoint and the anchor without materializing a separate model. Algorithm 2 gives the weight-space variant, where the extrapolated parameters θfuture are materialized explicitly and a forward pass through them produces the teacher distribution. Both variants share the same two-phase structure and differ only in how πfuture is computed. Extrapolation Decay. A fixed β is unsuitable across the full training trajectory. Under the linear-trajectory model (Proposition 3, Appendix A), the safe β range narrows as the policy approaches the optimum: beyond a β-threshold the extrapolated teacher overshoots π ∗ and the suboptimality bound worsens. We therefore apply a monotonically decreasing schedule βn → 1 (e.g., linear: βn = 1 + (β0 − 1)(1 − n/N )). Large βn early in training extrapolates aggressively when the policy is far from optimal; as training progresses, decreasing βn keeps the teacher within the valid range. Anchor Dynamics. The anchor θanchor —the reference checkpoint used to compute the displacement vector—is by default the previous checkpoint θn . Alternatively, an EMA anchor θanchor ← (1 − η)θanchor + η θn+1 both smooths and enlarges the displacement: the anchor lags behind the current policy, averaging the direction over ′ multiple iterations while increasing ∥θn+1 − θanchor ∥. The previous-checkpoint anchor is the special case η = 1. We ablate this choice in §4.4. Role of OPD. Why distill from πfuture rather than adopt it directly as the next policy? The extrapolated point lies beyond the policy’s trust region: adopting it wholesale risks degenerate behavior at large β, and amplifies any noise in the RLVR step. OPD instead acts as a trust-region projection—it moves the policy toward πfuture ’s token-level ′ distribution while the divergence term in LRISE keeps the update anchored near πθn+1 ; we validate this choice and the need for RLVR grounding in §4.3. Importantly, RISE does not introduce new task information beyond what RLVR provides—no external data or model enters the system. Rather, it exploits the model’s own inductive structure to convert a sparse outcome-induced parameter update into a dense token-level target, redistributing the same training signal into a form that enables finer-grained credit assignment. 6

Preprint | Salesforce AI Research

Algorithm 1 RISE (logit-space). The teacher is built once per iteration, before the OPD loop. Require: Initial policy πθ0 , extrapolation scale β0 , top-K, anchor rate η, iterations N 1: Initialize anchor θanchor ← θ0 2: for n = 0, 1, . . . , N − 1 do 3: βn ← 1 + (β0 − 1) · (1 − n/N ) ▷ Extrapolation decay 4: // RLVR phase 5: Sample rollouts D = {(xi , yi )} ∼ πθn , compute rewards R(xi , yi )  ′ 6: θn+1 ← PolicyGradient θn , {(xi , yi , Ri )} 7: // Teacher construction ▷ two forward passes over D ′ ′ 8: At every position of D: set S ← TopK (πθn+1 ) and project πθn+1 , πθanchor onto S ∪ {tail} ′ 9: Cache log πfuture ← log πθanchor + βn · (log πθn+1 − log πθanchor ) ▷ Eq. 8 10: // OPD phase ′ 11: θ ← θn+1 12: for each minibatch B ⊂ D do 13: Retrieve cached πfuture at the positions of B ▷ no teacher forward pass 14: θ ← Optimizer(θ, ∇θ LRISE (B)) ▷ Eq. 9 15: end for 16: θn+1 ← θ 17: Update anchor: θanchor ← (1−η) θanchor + η θn+1 ▷ η = 1: previous ckpt; η < 1: EMA 18: end for 19: return πθN Algorithm 2 RISE (weight-space). The teacher is evaluated inside the OPD loop. Require: Initial policy πθ0 , extrapolation scale β0 , top-K, anchor rate η, iterations N 1: Initialize anchor θanchor ← θ0 2: for n = 0, 1, . . . , N − 1 do 3: βn ← 1 + (β0 − 1) · (1 − n/N ) ▷ Extrapolation decay 4: // RLVR phase 5: Sample rollouts D = {(xi , yi )} ∼ πθn , compute rewards R(xi , yi )  ′ 6: θn+1 ← PolicyGradient θn , {(xi , yi , Ri )} 7: // Teacher construction ′ 8: Materialize θfuture ← θanchor + βn · (θn+1 − θanchor ) ▷ Eq. 5 9: // OPD phase ′ 10: θ ← θn+1 11: for each minibatch B ⊂ D do 12: At every position of B: set S ← TopK (πθ ) from the student forward pass 13: Forward pass through θfuture to get πfuture on S ∪ {tail} ▷ one teacher forward per step 14: θ ← Optimizer(θ, ∇θ LRISE (B)) ▷ Eq. 9 15: end for 16: θn+1 ← θ 17: Update anchor: θanchor ← (1−η) θanchor + η θn+1 18: end for 19: return πθN

4

Experiments

We evaluate RISE across multiple model scales, architectures, and domains to answer five questions: (1) Does RISE improve over RLVR-only training and existing OPSD methods? (2) Does the improvement hold across model scales and families? (3) Does RISE preserve performance on out-of-distribution tasks while improving in-domain accuracy? (4) What mechanisms drive the improvement? (5) Which design choices matter most? 7

Preprint | Salesforce AI Research

4.1

Setup

Models and datasets. We evaluate RISE on four task families. Mathematical reasoning: Qwen3-8B, Qwen31.7B, and Qwen3-1.7B-Base (Team, 2025) on DAPOMath (Yu et al., 2026), and OLMo3-7B-Instruct-SFT (Olmo et al., 2025) on OpenR1-Math-46K (Yan et al., 2026), with GPQA-Diamond, IFEval, and MMLU-Pro as out-ofdistribution checks. Multi-domain (math + STEM): Qwen3-4B-Base on a mixed corpus following Guru (Cheng et al., 2026), using only their STEM data and replacing their math split with DAPOMath (the original problems are low complexity). Code generation: Qwen3-8B-Base on Skywork-OR1-Code (He et al., 2025). Agentic tasks: Qwen2.5-3B-Instruct (Team, 2024) on ALFWorld (Shridhar et al., 2020) and WebShop (Yao et al., 2022) following GIGPO (Feng et al., 2025). Evaluation benchmarks, sample counts, and full training details are in Appendix B. Baselines. We report the base/SFT model to establish the starting point and GRPO (Guo et al., 2025) as the RLVR-only baseline. To isolate the contribution of RISE’s teacher construction, we compare against three methods that also pair GRPO’s sequence-level signal with per-token self-distillation from the same privileged teacher—the student conditioned on a sibling correct solution—but differ in how they integrate it: GRPO+SDPO (Hübotter et al., 2026) adds an auxiliary KL loss, SDAR (Lu et al., 2026) gates the GRPO advantage by the teacher–student probability gap, and RLSD (Yang et al., 2026a) reweights advantage estimates for finer-grained credit assignment. We exclude external-teacher OPD, which requires a separate stronger model, whereas RISE and all OPSD baselines use only the model itself (Appendix B.3). RISE configurations. We report both logit-space and weight-space RISE. Default hyperparameters: β0 = 1.2, linear decay to βN = 1, K = 100 (K = 20 for code, where the output distribution is more peaked); for the anchor we use η = 0.1 (EMA) for Qwen models and η = 1 (previous checkpoint) for OLMo (ablated in §4.4); the distillation loss uses Jensen–Shannon divergence (cf. §3.2). All other training hyperparameters (learning rate, batch size, rollout length) match the GRPO baseline for fair comparison. All methods are trained for one epoch on the respective training set. 4.2

Main Results

Mathematical Reasoning. Table 1 presents results across three model configurations spanning different scales (1.7B–8B), architectures (Qwen, OLMo), and training corpora (DAPOMath, OpenR1). RISE consistently outperforms all baselines. Both RISE variants beat every baseline on in-domain Math Avg in all three settings, and at least one ranks first or second on every individual in-domain benchmark. The gains are most pronounced on competition benchmarks: on OLMo3-7B, RISE (logit) improves AIME’24 from 30.2 to 46.9 (+16.7) and Math Avg from 47.6 to 56.4 (+8.8); on Qwen3-1.7B, RISE (logit) lifts Math Avg from 45.4 to 50.2 (+4.8); and on Qwen3-8B, RISE (weight) lifts it from 60.0 to 62.7 (+2.7). Neither extrapolation space consistently dominates, suggesting the gains stem from the extrapolation principle itself. Multi-seed experiments confirm reproducibility: RISE (weight) Math Avg = 62.4 ± 0.2 vs. GRPO 60.1 ± 0.2 on Qwen3-8B, and RISE (logit) 49.6 ± 0.4 vs. GRPO 45.1 ± 0.4 on Qwen3-1.7B, across three seeds (Appendix C.10). Privileged-conditioning baselines provide limited or negative gains. GRPO+SDPO, which jointly optimizes the GRPO loss and a KL divergence toward the privileged teacher, underperforms GRPO on both Qwen3-8B (55.9 vs. 60.0) and Qwen3-1.7B (43.2 vs. 45.4). SDAR and RLSD fare better but remain within about two points of GRPO in most settings. This pattern is consistent with our analysis in §2: privileged-conditioning teachers are limited by ICL capability, capping the benefit of per-token supervision regardless of integration strategy. Out-of-distribution preservation. RISE maintains or slightly improves OOD performance across all models—Qwen38B OOD Avg rises from 70.6 (GRPO) to 72.0 (RISE weight), OLMo3-7B from 51.2 to 55.5 (RISE logit)—so token-level refinement does not degrade general capabilities. Sample efficiency. RISE also learns faster: Figure 2 plots evaluation accuracy at successive checkpoints across three configurations, and RISE reaches higher accuracy earlier throughout, with the gap widest in early training where the extrapolated teacher supplies dense signal before RL advantage estimates stabilize (further curves in Appendix C.2). Multi-Domain Results. Table 2 tests whether RISE holds when the displacement aggregates improvements across heterogeneous domains. On Qwen3-4B-Base trained with mixed math and STEM data, RISE (weight) achieves the 8

Preprint | Salesforce AI Research

Table 1: Mathematical reasoning results. Accuracy (%) across model configurations. We evaluate on in-domain math benchmarks and out-of-distribution (OOD) tasks to assess preservation. Best in bold, second best underlined. In-Domain Math

Method

Qwen3-8B (DAPOMath) 73.2 27.1 Base 83.8 54.4 GRPO GRPO+SDPO 83.2 42.3 SDAR 84.5 50.2 RLSD 83.0 53.1 RISE (logit) 84.8 56.9 RISE (weight) 84.4 58.1

23.1 42.9 32.7 38.3 40.0 46.7 45.8

66.1 89.4 84.4 89.4 90.6 91.6 91.9

21.1 31.5 30.6 31.2 31.3 32.5 32.6

48.6 57.9 62.5 64.7 60.5 62.4 63.3

43.2 60.0 55.9 59.7 59.8 62.5 62.7

50.8 57.0 55.7 55.5 56.2 58.2 59.0

82.4 81.5 82.8 82.4 82.6 81.3 83.4

68.4 73.3 69.4 70.1 72.6 74.3 73.5

67.2 70.6 69.3 69.3 70.5 71.3 72.0

Qwen3-1.7B (DAPOMath) Base 64.6 11.5 GRPO 75.6 30.0 76.0 21.5 GRPO+SDPO SDAR 75.0 26.0 RLSD 75.3 26.7 RISE (logit) 75.2 36.3 RISE (weight) 77.3 32.9

12.1 26.3 23.1 26.0 25.2 33.3 31.3

41.9 65.6 65.0 71.3 59.4 78.8 75.0

17.6 23.9 23.1 23.6 23.0 24.9 25.7

40.3 51.1 50.3 50.8 49.8 52.9 53.3

31.3 45.4 43.2 45.5 43.2 50.2 49.2

34.9 32.3 34.9 35.3 33.4 34.8 34.8

67.7 68.8 69.0 69.1 67.5 68.2 70.2

48.3 53.4 52.0 53.2 54.1 53.7 56.3

50.3 51.5 52.0 52.5 51.7 52.2 53.8

OLMo3-7B-Instruct-SFT (OpenR1) Base 57.0 6.0 8.1 77.5 30.2 28.3 GRPO GRPO+SDPO 77.7 33.8 28.5 SDAR 79.8 35.2 27.1 RLSD 77.2 28.3 24.0 RISE (logit) 82.9 46.9 36.0 RISE (weight) 82.8 42.5 32.5

43.8 70.6 70.6 73.8 68.1 83.1 76.9

19.6 25.2 25.3 25.3 24.6 27.6 27.8

31.5 53.5 53.4 56.3 52.9 61.9 59.6

27.7 47.6 48.2 49.6 45.9 56.4 53.7

33.0 31.0 37.4 39.0 37.6 40.7 37.1

77.8 76.0 77.1 77.8 78.4 79.3 77.1

35.4 46.6 46.1 44.6 44.1 46.4 46.9

48.7 51.2 53.5 53.8 53.4 55.5 53.7

Qwen3-1.7B

— AIME'24 (avg@16)

Qwen3-8B

35

— AIME'25 (avg@16)

OLMo3-7B — OlympiadBench (avg@8) 60

30 25 GRPO

20

RISE (logit) RISE (weight)

15

SDAR 0

50

100

150

200

40 35 GRPO

30

RISE (logit) RISE (weight)

25

SDAR

20 250

0

50

100

150

Training Steps

Training Steps

(a) 1.7B AIME’24

(b) 8B AIME’25

200

250

Accuracy (%)

45

Accuracy (%)

Accuracy (%)

OOD

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench Avg. GPQA IFEval MMLU Avg.

50 40 30

0

GRPO RISE (logit) RISE (weight) SDAR 50 100 150 200 250 300 350

Training Steps

(c) OLMo OlyBench

Figure 2: RISE reaches higher accuracy in fewer steps across three model scales, demonstrating improved sample efficiency. highest Math Avg (44.8 vs. GRPO’s 40.2) and STEM Avg (47.5 vs. 45.5), improving competition math (AIME’24: 28.8, +5.0) while simultaneously lifting STEM (TheoremQA: 51.5, +3.7)—extrapolation along a multi-domain trajectory does not dilute gains in any single domain. GRPO+SDPO is more competitive here (STEM Avg 47.0) than in the math-only setting, consistent with STEM providing more informative privileged context (Hübotter et al., 2026), yet RISE still outperforms it without privileged information. The sample-efficiency advantage extends to this setting. On AMC’23 (Fig. 3a) both RISE variants stay above GRPO throughout training, with RISE (weight) reaching 70.0% while GRPO rises mid-training but declines to 59.4%; on SuperGPQA (Fig. 3b) RISE (weight) holds a consistent edge, so the signal benefits both math and STEM.

9

Preprint | Salesforce AI Research

Table 2: Multi-domain results (Qwen3-4B-Base, mixed math + STEM training). Accuracy (%) on math reasoning and STEM benchmarks. Best in bold, second best underlined. Math AIME24

AIME25

AMC23

Minerva

OlyBench

Avg.

48.2 69.1 70.5 69.1 68.9 71.9 71.3

9.8 23.8 24.8 22.9 23.1 25.6 28.8

10.4 20.6 20.4 20.0 20.0 22.1 25.4

35.0 12.1 59.4 26.6 68.1 25.9 65.6 23.7 60.0 24.6 70.6 25.9 71.3 26.9

27.5 42.0 43.1 41.2 39.9 43.8 44.9

23.8 14.1 40.2 42.4 42.1 44.7 40.4 40.6 39.4 33.7 43.3 44.1 44.8 44.7

Base GRPO GRPO+SDPO SDAR RLSD RISE (logit) RISE (weight)

— AMC'23 (avg@4)

Qwen3-4B-Base (Mixed)

— SuperGPQA (avg@1)

60 GRPO

50

GRPO+SDPO RISE (logit)

40

RISE (weight) 0

50

100

150

200

250

300

350

Accuracy (%)

Accuracy (%)

70

GPQA

SuperGPQA

8.0 32.5 31.0 30.3 28.8 32.1 32.9

80

25 20

GRPO

15

GRPO+SDPO

10

RISE (weight)

RISE (logit)

0

50

100

150

200

250

300

Training Steps

Training Steps

(a) AMC’23

(b) SuperGPQA

MMLU

70 60 50 40

350

0

20

40

60

Training Steps

GRPO RISE (logit) RISE (weight) 80 100

(a) HumanEval+

ThrmQA

Avg.

6.6 27.0 13.9 59.5 47.8 45.5 61.7 50.5 47.0 58.7 49.1 44.7 55.1 46.5 41.0 61.9 51.3 47.4 60.9 51.5 47.5

Qwen3-8B-Base (Code) — HumanEval+ (avg@4)

30

Accuracy (%)

Qwen3-4B-Base (Mixed)

STEM

MATH500

Qwen3-8B-Base (Code) 65

Accuracy (%)

Method

— MBPP+ (avg@4)

60 55 50 GRPO 45

RISE (logit) RISE (weight)

40 0

20

40

60

80

100

Training Steps

(b) MBPP+

Figure 3: RISE reaches higher accuracy on both math Figure 4: RISE converges faster than GRPO on code and STEM benchmarks. generation. Code Generation. We evaluate Qwen3-8B-Base on code generation (Fig. 4). Both RISE variants converge faster than GRPO early and reach comparable final accuracy—RISE (logit) matches GRPO’s final HumanEval+ accuracy at step 50 versus step 90—so the extrapolation principle transfers beyond mathematical reasoning. Gains are more modest than in math, likely because all methods saturate quickly on these benchmarks. Full curves (avg@4 and pass@4) are in Appendix C.3. Agentic Tasks. Table 3 reports results on two multi-turn agentic bench- Table 3: Success rate on ALFWorld, marks using Qwen2.5-3B-Instruct as the base model, following the GIGPO Score/Acc on WebShop. training setup (Feng et al., 2025). RISE (weight) substantially outperforms Method ALF WS-S WS-A GRPO on both ALFWorld (+9.4) and WebShop (+10.9 Acc), demonstrating Base 21.9 6.7 0.8 that the extrapolation principle extends to sequential decision-making tasks 63.3 GRPO 75.0 79.8 RISE (logit) 78.1 78.8 68.0 where the reward signal is sparse and delayed. RISE (weight) 84.4

4.3

86.3

74.2

Analysis: How Does Extrapolation Help?

Beyond aggregate accuracy, we ask why extrapolation helps. RISE couples two ingredients—an RLVR phase that supplies the displacement direction, and an OPD phase that projects the extrapolated teacher back onto the policy. We first isolate each phase by removing it (Fig. 5), then test the extrapolation premise directly on saved checkpoints (Fig. 6). Extrapolation is only safe when the direction is grounded in verifiable reward. We remove the RLVR phase entirely and extrapolate along the displacement induced by self-distillation alone. Training collapses within 60 steps (Fig. 5a): MATH-500 accuracy drops from the base level to 2.4%, response length explodes from 2K to the 8K context cap, and training reward drops to zero (Fig. C.3). GRPO and RISE show no such behaviour. Without an outcome reward to anchor it, the displacement is purely self-referential—each iteration extrapolates along the model’s prior move—and repeated amplification drives the policy into degenerate non-terminating generation. The gain comes from distillation, not from taking a longer step. We next remove the OPD phase, adopting the extrapolated θfuture directly as the next policy. This variant is stable but does not improve the average: on Qwen3-8B Math Avg moves from 60.0 (GRPO) to 60.3 (w/o OPD), and on Qwen3-1.7B from 45.4 to 45.6— negligible gains compared to RISE’s +2.7 and +4.8 (Table C.3). Individual benchmarks shift in opposite directions 10

8000

60

6000

GRPO RISE (logit) w/o RLVR

40

4000

20 0

accuracy length

0

20

40

60

80

Training Steps

100 120

(a) Removing RLVR (8B)

2000 0

— AIME'25 (avg@16)

Qwen3-1.7B — OlympiadBench (avg@8)

40

Accuracy (%)

80

Qwen3-8B 45

Accuracy (%)

MATH-500 avg@4 (%)

Qwen3-8B — removing the RLVR phase

Response length (tokens)

Preprint | Salesforce AI Research

35 30

GRPO RISE (weight)

25

w/o OPD

20 0

50

100

150

200

50 45

GRPO RISE (logit) w/o OPD 150 200

40 0

250

Training Steps

(b) Removing OPD (8B)

50

100

Training Steps

(c) Removing OPD (1.7B)

Figure 5: Removing either phase hurts. (a) Without RLVR, accuracy collapses (solid) as generation length explodes (dotted). (b, c) Without OPD, training is stable but gains vanish.

The safe extrapolation range narrows over training. We test the extrapolation premise directly by materializing θbase + β(θn − θbase ) from GRPO checkpoints of Qwen3-1.7B and evaluating on AIME’24 (Fig. 6). Two patterns emerge. First, within RISE’s operating range (β ≤ 1.2, shaded), extrapolation is safe throughout: gains are nearneutral early and reach +2–4 points at steps 100–200, while never degrading accuracy. Second, the tolerable range of β contracts sharply as training proceeds: at step 50, even β = 2.0 adds +7.3 points, but by step 100 the same β is catastrophic (−15 points), and by step 200 β = 1.5 already costs 7.7 points. This is the empirical counterpart of Proposition 3 and the direct justification for the decaying schedule: a fixed aggressive β would eventually destroy the teacher, whereas RISE’s conservative, decaying β stays inside the safe region throughout.

Gain over checkpoint (pts)

at different scales (Figs. 5b, 5c), but these redistributions do not translate into consistent improvement. Taking a longer step in weight space shifts the policy without reliably improving it; it is the OPD phase—projecting the extrapolated teacher’s per-token distribution back onto the policy—that converts the extrapolation direction into consistent gains across benchmarks and scales. 1.7B, AIME'24 (avg@16) RISE range

5 0

−5 −10 −15 1.0

step 5

step 50

step 10

step 100

step 20

step 200

1.2

1.4

1.6

1.8

Extrapolation scale β

2.0

Figure 6: Gain on AIME’24 (avg@16) from extrapolating Qwen3-1.7B GRPO checkpoints at varying β, relative to β = 1. Shaded: RISE’s operating range (β0 = 1.2, decaying to 1). RISE broadens solution coverage, not just average quality. A natural concern with distillation is that it sharpens the policy around solutions it already finds, improving mean accuracy at the expense of coverage. We test this by comparing avg@16 with pass@16 (best-of-16) on the AIME benchmarks (Appendix C.4). RISE improves pass@16 over GRPO in nearly every setting, with the effect strongest at 1.7B: RISE (logit) lifts AIME’24 pass@16 by +9.6 points versus +6.3 on avg@16, indicating that the extrapolated teacher expands the set of solvable problems rather than merely sharpening existing solutions. At 8B the coverage gains are more modest, consistent with less headroom when the base policy is already stronger. 4.4

Ablation Studies

We conduct ablations primarily on Qwen3-8B and Qwen3-1.7B(-Base) to isolate the contribution of each design choice. Extrapolation scale β and decay schedule. Table 4 reports sen- Table 4: β0 and schedule ablation (Qwen3-8B, sitivity to β0 and the decay schedule on Qwen3-8B (logit-space Math Avg %). extrapolation). Performance is stable across β0 ∈ [1.2, 1.5], with β0 (linear decay) Schedule (β0 = 1.2) β0 = 1.2 slightly ahead; β0 = 2.0 diverged in preliminary ex1.2 1.3 1.5 cosine linear exp fixed periments, consistent with the narrowing safe range shown in Math Avg 62.5 62.3 61.9 63.0 62.5 62.0 61.8 Fig. 6: although β = 2.0 is benign early in training, it becomes catastrophic once the policy nears optimality, and linear decay from β0 = 2.0 does not reduce β fast enough to avoid this regime. All decay schedules perform comparably, though any form of decay edges out a fixed β—consistent with the prediction that β should shrink as the policy approaches optimality. On Qwen3-1.7B11

Preprint | Salesforce AI Research

Base the pattern is similar: β0 = 1.3 (linear) gives the best average (29.4), while β0 = 1.5 degrades to 28.5 (per-benchmark breakdowns in Appendix C.6).

Math Avg (%)

Anchor dynamics (η). The EMA anchor (η < 1) smooths the displacement 51 Logit direction over multiple iterations, which helps when per-step RLVR updates are 50 Weight noisy. Fig. 7 ablates η on Qwen3-1.7B. In logit space, reducing η monotonically 49 improves Math Avg: from 46.2 (η = 1.0) to 48.4 (η = 0.3) to 50.2 (η = 0.1); in 48 weight space the trend is consistent (+2.8 from η = 1.0 to η = 0.1). However, 47 on OLMo3-7B η = 1.0 outperforms η = 0.1 (56.4 vs. 50.1), likely because 46 GRPO on OLMo already produces large, stable per-step directions, making EMA 0.1 0.3 0.95 1.0 η (anchor update rate) smoothing unnecessary and its lag counterproductive. We therefore default to η = 0.1 for Qwen and η = 1 for OLMo (Appendix C.8). Figure 7: Anchor ablation on Qwen3-1.7B. Lower η improves ′ Resample after RLVR vs. rollout reuse. Generating fresh rollouts from πθn+1 for the OPD phase eliminates the mild off-policy-ness discussed in §3.3, but does Math Avg. not improve results: on Qwen3-8B the resample variant matches rollout reuse exactly (62.5 Math Avg), and on Qwen3-1.7B-Base it slightly underperforms (28.4 vs. 29.2). Rollout reuse is therefore the practical default, halving the per-iteration sampling cost (per-benchmark details in Appendix C.7). Compute-matched comparison. Since RISE applies two gradient phases per iteration (RLVR + OPD), we compare against GRPO that performs a second inner-loop gradient pass on the same rollouts within each iteration (GRPO2×), matching the total gradient budget. The extra pass helps only marginally—+1.9 Math Avg on Qwen3-1.7B (47.3 vs. 45.4) and +0.5 on Qwen3-8B (60.5 vs. 60.0)—while RISE gains +4.8 and +2.7 over GRPO at the same budget, outperforming GRPO-2× by +2.9 and +2.2 respectively (per-benchmark details in Appendix C.9). The gains therefore stem from the quality of extrapolated supervision rather than additional optimization steps. Training cost. RISE reuses the RLVR rollouts for OPD, so the overhead comes from the OPD gradient phase and the anchor forward pass. On Qwen3-8B, RISE runs at ∼1.6× GRPO wall time per iteration; on Qwen3-1.7B, ∼1.3×. The smaller relative overhead at 1.7B reflects that rollout generation—shared between GRPO and RISE—dominates at smaller scale (73% of GRPO wall time vs. 63% at 8B).

5

Conclusion

We introduced RISE, which constructs a per-token teacher by extrapolating the model’s own RLVR training trajectory—in logit space or weight space—then distills the resulting future policy back into the current student. The two phases are complementary: outcome rewards ground the extrapolation in verified improvement, while distillation converts the teacher’s token-level distribution into durable policy gains. Across four task families (math, STEM, code, and agentic tasks) spanning 1.7B–8B parameters, RISE consistently outperforms RLVR and privileged-conditioning baselines, with the largest margins on challenging competition benchmarks. Analysis confirms that both phases are necessary, that the safe extrapolation range narrows over training, and that RISE broadens solution coverage alongside mean accuracy—all at 1.3–1.6× GRPO wall time with no additional sampling cost. More broadly, RISE demonstrates that a model’s own training trajectory contains sufficient structure to serve as a self-improving teacher—converting sparse outcome signals into dense token-level supervision without any external knowledge. We believe this principle extends naturally to other post-training paradigms such as preference optimization and multi-agent settings. Limitations. RISE relies on the training trajectory being sufficiently low-dimensional for linear extrapolation to remain meaningful at modest β. The β-decay schedule and EMA anchor provide empirical robustness, yet principled detection of when extrapolation becomes unreliable remains open. Additionally, RISE inherits any biases in the RLVR reward signal: if the reward is hackable, extrapolation amplifies the spurious direction. Integrating reward model uncertainty or multi-objective rewards to qualify the extrapolation is a promising avenue for future work.

12

Preprint | Salesforce AI Research

References Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. In International Conference on Learning Representations, volume 2024, pp. 21246–21263, 2024. Rishabh Agrawal, Jacob Fein-Ashley, and Paria Rashidinejad. Reinforcement learning from rich feedback with distributional dagger. arXiv preprint arXiv:2606.05152, 2026. Yuchen Cai, Ding Cao, Xin Xu, Zijun Yao, Yuqing Huang, Zhenyu Tan, Benyi Zhang, Guangzhong Sun, Guiquan Liu, and Junfeng Fang. On predictability of reinforcement learning dynamics for large language models. arXiv preprint arXiv:2510.00553, 2025. Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. Theoremqa: A theorem-driven question answering dataset. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7889–7901, 2023. Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335, 2024. Jorge Zhoujun Cheng, Shibo Hao, Tianyang Liu, Fan Zhou, Yutao Xie, Feng Yao, Yuexin Bian, Nilabjo Dey, Yonghao Zhuang, Yuheng Zha, et al. Revisiting reinforcement learning for llm reasoning from a cross-domain perspective. Advances in Neural Information Processing Systems, 38, 2026. Xeron Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, et al. Supergpqa: Scaling llm evaluation across 285 graduate disciplines. Advances in Neural Information Processing Systems, 38, 2026. Wenkai Fang, Shunyu Liu, Yang Zhou, Kongcheng Zhang, Tongya Zheng, Kaixuan Chen, Mingli Song, and Dacheng Tao. Serl: Self-play reinforcement learning for large language models with limited data. Advances in Neural Information Processing Systems, 38:103706–103738, 2026. Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An. Group-in-group policy optimization for llm agent training. arXiv preprint arXiv:2505.10978, 2025. Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. Linear mode connectivity and the lottery ticket hypothesis. In International conference on machine learning, pp. 3259–3269. PMLR, 2020. Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, and Akshay Nambi. Privileged, but biased: How pi-conditioned teachers break self-distillation. arXiv preprint arXiv:2608.04794, 2026. Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, et al. Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. arXiv preprint arXiv:2402.14008, 2024. Jujie He, Jiacai Liu, Chris Yuhao Liu, Rui Yan, Chaojie Wang, Peng Cheng, Xiaoyu Zhang, Fuxiang Zhang, Jiacheng Xu, Wei Shen, Siyuan Li, Liang Zeng, Tianwen Wei, Cheng Cheng, Bo An, Yang Liu, and Yahui Zhou. Skywork open reasoner 1 technical report. arXiv preprint arXiv:2505.22312, 2025.

13

Preprint | Salesforce AI Research

Chengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang, Zongxia Li, Ruosen Li, Jiaxin Huang, Haitao Mi, and Dong Yu. R-zero: Self-evolving reasoning llm from zero data. arXiv preprint arXiv:2508.05004, 2025. Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, et al. Reinforcement learning via self-distillation. arXiv preprint arXiv:2601.20802, 2026. Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089, 2022. Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. In 34th Conference on Uncertainty in Artificial Intelligence, pp. 876–885, 2018. Naman Jain, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. In International Conference on Learning Representations, volume 2025, pp. 58791–58831, 2025. Ijun Jang, Jewon Yeom, Juan Yeo, Hyunggu Lim, and Taesup Kim. Stable on-policy distillation through adaptive target reformulation. arXiv preprint arXiv:2601.07155, 2026. Maxwell Jia. Aime problem set 2024, 2024. Maxwell-Jia/AIME_2024.

URL https://huggingface.co/datasets/

Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, and Zequn Sun. Asymmetric on-policy distillation: Bridging exploitation and imitation at the token level. arXiv preprint arXiv:2605.06387, 2026. Li Jiang, Haoran Xu, Yichuan Ding, and Amy Zhang. arXiv:2606.08432, 2026.

Trajectory-refined distillation.

arXiv preprint

Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, and Kimin Lee. Entropy-aware on-policy distillation of language models. arXiv preprint arXiv:2603.07079, 2026. Seongryong Jung, Suwan Yoon, DongGeon Kim, and Hwanhee Lee. Todi: Token-wise distillation via finegrained divergence control. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 8089–8102, 2025. Sham Kakade and John Langford. Approximately optimal approximate reinforcement learning. In Proceedings of the nineteenth international conference on machine learning, pp. 267–274, 2002. Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, and Yuqing Yang. Why does self-distillation (sometimes) degrade the reasoning capability of llms? arXiv preprint arXiv:2603.24472, 2026. Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen, and Pashmina Cameron. Scaling reasoning efficiently via relaxed on-policy distillation. arXiv preprint arXiv:2603.11137, 2026. Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. Solving quantitative reasoning problems with language models. Advances in neural information processing systems, 35:3843–3857, 2022. Changhao Li, Rushi Qiang, Jiawei Huang, Chenxiao Gao, Chao Zhang, Niao He, and Bo Dai. Revisiting dagger in the era of llm-agents. arXiv preprint arXiv:2605.12913, 2026a. Yang Li, Erik Nijkamp, Semih Yavuz, and Shafiq Rayhan Joty. Learning from language feedback via variational policy distillation. arXiv preprint arXiv:2605.15113, 2026b. 14

Preprint | Salesforce AI Research

Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, et al. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe. arXiv preprint arXiv:2604.13016, 2026c. Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In International Conference on Learning Representations, volume 2024, pp. 39578–39601, 2024. Bo Liu, Leon Guertler, Simon Yu, Zichen Liu, Penghui Qi, Daniel Balcells, Mickel Liu, Cheston Tan, Weiyan Shi, Min Lin, et al. Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning. arXiv preprint arXiv:2506.24119, 2025. Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=1qvx610Cu7. Kaiyuan Liu, Ziyuan Zhuang, Yang Bai, Bing Wang, Rongxiang Weng, and Jieping Ye. Prefix teach, suffix fade: Local teachability collapse in strong-to-weak on-policy distillation. arXiv preprint arXiv:2605.13643, 2026. Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy-distillation. Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang, Jinyang Wu, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, et al. Self-distilled agentic reinforcement learning. arXiv preprint arXiv:2605.15155, 2026. MAA. AMC 10/12 2023, 2023. URL https://www.maa.org/math-competitions/amc-1012. math ai. Aime problem set 2025, 2025. URL https://huggingface.co/datasets/math-ai/aime25. Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi, and Yohan Jo. Kl for a kl: On-policy distillation with control variate baseline. arXiv preprint arXiv:2605.07865, 2026. Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng, Shane Arora, Shashank Gupta, Taira Anderson, Teng Xiao, Tyler Murray, Tyler Romero, Victoria Graf, Akari Asai, Akshita Bhagia, Alexander Wettig, Alisa Liu, Aman Rangapur, Chloe Anastasiades, Costa Huang, Dustin Schwenk, Harsh Trivedi, Ian Magnusson, Jaron Lochner, Jiacheng Liu, Lester James V. Miranda, Maarten Sap, Malia Morgan, Michael Schmitz, Michal Guerquin, Michael Wilson, Regan Huff, Ronan Le Bras, Rui Xin, Rulin Shao, Sam Skjonsberg, Shannon Zejiang Shen, Shuyue Stella Li, Tucker Wilde, Valentina Pyatkin, Will Merrill, Yapei Chang, Yuling Gu, Zhiyuan Zeng, Ashish Sabharwal, Luke Zettlemoyer, Pang Wei Koh, Ali Farhadi, Noah A. Smith, and Hannaneh Hajishirzi. Olmo 3, 2025. URL https://arxiv.org/abs/2512.13961. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022. Miguel Moura Ramos, Duarte M Alves, and André FT Martins. Combining on-policy optimization and distillation for long-context reasoning in large language models. arXiv preprint arXiv:2605.12227, 2026. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022, 2023. Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 627–635. JMLR Workshop and Conference Proceedings, 2011. 15

Preprint | Salesforce AI Research

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong, Zhilin Wang, Yanxu Chen, Rongduo Han, Sunbowen Lee, and Yi R Fung. On the geometry of on-policy distillation. arXiv preprint arXiv:2606.07082, 2026. Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768, 2020. Mingyang Song and Mao Zheng. A survey of on-policy distillation for large language models. arXiv preprint arXiv:2604.00626, 2026. Zhengwei Tao, Ting-En Lin, Xiancai Chen, Hangyu Li, Yuchuan Wu, Yongbin Li, Zhi Jin, Fei Huang, Dacheng Tao, and Jingren Zhou. A survey on self-evolution of large language models. arXiv preprint arXiv:2404.14387, 2024. Qwen Team. Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm.github.io/ blog/qwen2.5/. Qwen Team. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. Tianle Wang, Zhongyuan Wu, Shenghao Jin, Hao Xu, Wei Chen, and Ning Miao. Not all steps are informative: On the linearity of llms’ rlvr training. arXiv preprint arXiv:2601.04537, 2026a. Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, and Ling Yang. Openclaw-rl: Train any agent simply by talking. arXiv preprint arXiv:2603.10165, 2026b. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024. Zhepei Wei, Xinyu Zhu, Wei-Lin Chen, Chengsong Huang, Jiaxin Huang, and Yu Meng. You only need minimal rlvr training: Extrapolating llms via rank-1 trajectories. arXiv preprint arXiv:2605.21468, 2026. Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carlin, Simon Kornblith, et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International conference on machine learning, pp. 23965–23998. PMLR, 2022. LLM-Core Xiaomi. Mimo-v2-flash technical report, 2025. URL https://github.com/XiaomiMiMo/ MiMo-V2-Flash/paper.pdf. Xingrun Xing, Haoqing Wang, Boyan Gao, Ziheng Li, and Yehui Tang. Trust region on-policy distillation. arXiv preprint arXiv:2606.01249, 2026. Wei Xiong, Jiarui Yao, Yuhui Xu, Bo Pang, Lei Wang, Doyen Sahoo, Junnan Li, Nan Jiang, Tong Zhang, Caiming Xiong, et al. A minimalist approach to llm reasoning: from rejection sampling to reinforce. arXiv preprint arXiv:2504.11343, 2025. Hongling Xu, Qi Zhu, Heyuan Deng, Jinpeng Li, Lu Hou, Yasheng Wang, Lifeng Shang, Ruifeng Xu, and Fei Mi. Kdrl: Post-training reasoning llms via unified knowledge distillation and reinforcement learning. arXiv preprint arXiv:2506.02208, 2025. Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, and Alborz Geramifard. Beyond grpo and on-policy distillation: An empirical sparse-to-dense reward principle for language-model post-training. arXiv preprint arXiv:2605.12483, 2026a. 16

Preprint | Salesforce AI Research

Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, and Alborz Geramifard. Tip: Token importance in on-policy distillation. arXiv preprint arXiv:2604.14084, 2026b. Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang, Ganqu Cui, Xiaoye Qu, Yu Cheng, and Yue Zhang. Learning to reason under off-policy guidance. Advances in Neural Information Processing Systems, 38:117157–117186, 2026. Chenxu Yang, Chuanyu Qin, Qingyi Si, Minghui Chen, Naibin Gu, Dingyu Yao, Zheng Lin, Weiping Wang, Jiaqi Wang, and Nan Duan. Self-distilled rlvr. arXiv preprint arXiv:2604.03128, 2026a. Han Yang, Mingyan Wu, Bailan He, Zeyu Cao, Sikuan Yan, Kevin Qinghong Lin, and Zifeng Ding. Reasoning compression with mixed-policy distillation. arXiv preprint arXiv:2605.08776, 2026b. Haoyan Yang, Mario Xerri, Solha Park, Huajian Zhang, Yiyang Feng, Sai Akhil Kogilathota, and Jiawei Zhou. Self-improvement of large language models: A technical overview and future outlook. arXiv preprint arXiv:2603.25681, 2026c. Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation. arXiv preprint arXiv:2602.12125, 2026d. Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757, 2022. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. Advances in Neural Information Processing Systems, 38:113222–113244, 2026. Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488, 2022. Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang, Dhananjay Ram, Shuo Yang, Zhuowen Tu, Wei Xia, and Stefano Soatto. Reinforcement-aware knowledge distillation for llm reasoning. arXiv preprint arXiv:2602.22495, 2026. Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, and Xiaoyu Shen. Decoupling kl and trajectories: A unified perspective for sft, dagger, offline rl, and opd in llm distillation. arXiv preprint arXiv:2605.16826, 2026a. Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and Aditya Grover. Self-distilled reasoner: On-policy self-distillation for large language models. arXiv preprint arXiv:2601.18734, 2026b. Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023. Siqi Zhu, Xuyan Ye, Hongyu Lu, Weiye Shi, and Ge Liu. The many faces of on-policy distillation: Pitfalls, mechanisms, and fixes. arXiv preprint arXiv:2605.11182, 2026.

17

Preprint | Salesforce AI Research

A

Theoretical Analysis

A.1

Suboptimality Decomposition and Optimal Teacher

We first state the performance difference lemma adapted to the autoregressive setting (Kakade & Langford, 2002). Lemma 1 (Performance difference). For two policies π1 , π2 generating responses of length T with terminal reward R(x, y) ∈ [0, 1]: " # T X X J(π1 ) − J(π2 ) = Est ∼ρtπ (π1 (v | st ) − π2 (v | st )) Qπ2 (st , v) , (A.1) 1

v

t=1

where st = (x, y<t ) is the context at position t, ρtπ1 is the distribution over position-t contexts induced by π1 , and Qπ2 (st , v) = Eyt+1:T ∼π2 [R(x, y) | st , yt = v] ∈ [0, 1] is the action-value function under π2 . Proof. The performance difference lemma (Kakade & Langford, 2002), originally stated for the discounted infinitehorizon setting, specializes to the finite-horizon undiscounted case as J(π1 ) − J(π2 ) =

T X

Est ∼Ptπ1 [Eat ∼π1 [Aπ2 (st , at )]],

(A.2)

t=1

where Aπ2 (s, a) = Qπ2 (s, a) − V π2 (s) is the advantage under π2 . In the autoregressive setting, states are st = (x, y<t ) and actions are tokens v. Expanding Aπ2 = Qπ2 − V π2 and noting that V π2 (st ) is constant w.r.t. the sum over v: X X X π1 (v|st ) Aπ2 (st , v) = π1 (v|st ) Qπ2 (st , v) − V π2 (st ) π1 (v|st ) v

v

v

| =

X

π1 (v|st ) Qπ2 (st , v) −

X

v

X

}

π2 (v|st ) Qπ2 (st , v)

v

| =

{z

=1

{z

= V π2 (st )

(π1 (v|st ) − π2 (v|st )) Qπ2 (st , v).

} (A.3)

v

Finally, Qπ2 (st , v) ∈ [0, 1] because it is a conditional expectation of R ∈ [0, 1]. Theorem 2 (Suboptimality decomposition for OPD). Let πθ be a student and πT a teacher, with π ∗ = arg maxπ J(π), J(π) = Ex [Ey∼π [R(x, y)]], R ∈ [0, 1], and response length T . The student’s suboptimality decomposes as: q J(π ∗ ) − J(πθ ) ≤ J(π ∗ ) − J(πT ) + T · L̄OPD /2 , | {z } | {z } teacher gap

(A.4)

distillation error

PT 1

where L̄OPD = T t=1 Est ∼ρtπ [DKL (πθ (· | st ) ∥ πT (· | st ))] is the average per-token reverse KL under the θ student’s own context distribution. The bound identifies two independent error sources: a distillation error that vanishes as OPD converges (πθ → πT ), and a teacher gap that persists regardless of optimization quality. Because the T -dependent prefactor makes the bound a qualitative mechanism decomposition rather than a tight numerical certificate, its practical value is directional: the decomposition shows that, all else equal, a stronger teacher (smaller gap) admits a tighter bound on student suboptimality. Proof. Decompose the gap as J(π ∗ ) − J(πθ ) = [J(π ∗ ) − J(πT )] + [J(πP T ) − J(πθ )]. For P the second term, apply T Lemma 1 with π1 = πθ and π2 = πT , which gives J(πθ ) − J(πT ) = t=1 Est ∼ρtπ [ v (πθ (v | st ) − πT (v | θ st )) QπT (st , v)]. Multiplying by −1 flips both the performance difference and the inner difference: " # T X X πT J(πT ) − J(πθ ) = Est ∼ρtπ (πT (v | st ) − πθ (v | st )) Q (st , v) . (A.5) θ

t=1

v

18

Preprint | Salesforce AI Research

πT Note that this orientation places the expectation under the student’s own context distribution. Since P Q (st , v) ∈ [0, 1] (it is a conditional expectation of R ∈ [0, 1]), the range maxv Q − minv Q ≤ 1. Because v (πT (v|st ) − πθ (v|st )) = 0, we may subtractP any constant from Q inside the sum. Choosing c = (max Q + min Q)/2 and P applying Hölder’s inequality: | v (πT − πθ )Q|P= | v (πT − πθ )(Q − c)| ≤ ∥Q − c∥∞ · ∥πT − πθ ∥1 = max Q−min Q · 2 TV ≤ TV, where TV(p, q) = 21 v |p(v) − q(v)|. Therefore: 2

J(πT ) − J(πθ ) ≤

T X

Est ∼ρtπ [TV(πT (·|st ), πθ (·|st ))] . θ

(A.6)

t=1

p Applying Pinsker’s inequality (TV(p, q) ≤ D in either argument order, since p to each term—valid √KL (p∥q)/2) √ TV is symmetric—then Jensen’s inequality (E[ X] ≤ E[X] since · is concave): T X

Est ∼ρtπ [TVt ] ≤

T q X

θ

θ

t=1

T

Est ∼ρtπ [Dt ]/2 = T ·

t=1

1 Xp E[Dt ]/2, T t=1

(A.7)

q P P √ 1 where Dt = DKL (πθ (·|st )∥πT (·|st )). By Cauchy–Schwarz, T1 t at ≤ t at for at = E[Dt ]/2 ≥ 0. T p P p Therefore t E[Dt ]/2 ≤ T · L̄OPD /2, giving the stated bound. The teacher gap J(π ∗ ) − J(πT ) ≥ 0 by definition of π ∗ , with equality iff πT is itself optimal, i.e., J(πT ) = J(π ∗ ). A.2

Extrapolation Gap and Convergence

Proposition 3 (Safe extrapolation range under a linear trajectory). Let φ : Π → Z be a representation map such that the training trajectory is linear: φ(πθn ) = φ(πθ0 ) + αn · d for a unit direction d and monotonically increasing coefficients 0 = α0 < αn < α∗ , where φ(π ∗ ) = φ(πθ0 ) + α∗ · d. The extrapolated teacher φ(πfuture ) = φ(πθ0 ) + β · αn · d is closer to π ∗ than the student is, i.e. ∥φ(πfuture ) − φ(π ∗ )∥ < ∥φ(πθn ) − φ(π ∗ )∥,

(A.8)

exactly when 1 < β < 2α∗ /αn − 1. This is an immediate consequence of the linear-trajectory assumption—it defines the safe β range rather than derives it from weaker premises—and its value lies in making two properties explicit: (i) any β ∈ (1, α∗ /αn ) places the teacher strictly between the student and π ∗ without overshooting, and (ii) the upper limit 2α∗ /αn − 1 shrinks as αn → α∗ , so the safe range narrows as training converges. Proof. The gap between the extrapolated teacher and π ∗ is ∥φ(πfuture ) − φ(π ∗ )∥ = |βαn − α∗ |. The gap between the current student and π ∗ is α∗ − αn . The extrapolated teacher is closer whenever |βαn − α∗ | < α∗ − αn , which holds iff 1 < β < 2α∗ /αn − 1. For β ∈ (1, α∗ /αn ), we have βαn < α∗ , so the teacher lies strictly between the student and π ∗ . For β ∈ [α∗ /αn , 2α∗ /αn − 1), the teacher reaches or overshoots π ∗ (with equality at β = α∗ /αn ) but remains closer to it than the student is. ′ Corollary 4 (RISE contracts the suboptimality bound). Consider one RISE iteration: RLVR updates θn → θn+1 ′ ∗ (advancing the trajectory coefficient αn → αn+1 < α ), then OPD distills the extrapolated teacher πfuture into ′ ′ θn+1 to yield θn+1 . Under the assumptions of Proposition 3 (applied with πθn+1 in the role of the current policy) and L-Lipschitz continuity of J in φ-space, the Lipschitz suboptimality bound contracts: q ′ J(π ∗ ) − J(πθn+1 ) ≤ γ(β) · L(α∗ − αn+1 ) + T L̄OPD /2, (A.9) |{z} {z } | | {z } post-OPD

<1

≥ J(π ∗ )−J(πθ′

n+1

) (pre-OPD)

where the contraction factor is governed by the teacher’s distance to π ∗ , γ(β) =

′ α∗ − βαn+1 , ′ ∗ α − αn+1

19

(A.10)

Preprint | Salesforce AI Research

′ and γ(β) < 1 exactly when β lies in the range (1, 2α∗ /αn+1 − 1) established by Proposition 3. In the idealized limit L̄OPD → 0 the residual term vanishes and the bound contracts by γ(β) < 1. In practice, OPD runs for only a few gradient steps and deliberately does not converge to πfuture —it acts as a trust-region projection (§3) that moves ′ the policy toward the extrapolated teacher while remaining anchored near πθn+1 . The corollary’s role is therefore directional: it identifies the extrapolated teacher as a descent direction for the bound rather than prescribing full convergence, which §4.3 confirms is unnecessary and even harmful (the w/o-OPD variant that adopts πfuture directly shows negligible gains). ′ ′ Proof. Before OPD, the post-RLVR student πθn+1 lies at φ-distance α∗ − αn+1 from π ∗ , so L-Lipschitz continuity of J in φ-space gives ′ ′ J(π ∗ ) − J(πθn+1 ) ≤ L(α∗ − αn+1 ). (A.11) ′ The extrapolated teacher sits at φ-distance |α∗ − βαn+1 | from π ∗ . Applying L-Lipschitz continuity again and factoring out the pre-OPD bound: ′ J(π ∗ ) − J(πfuture ) ≤ L α∗ − βαn+1 =

′ α∗ − βαn+1 ′ ′ · L(α∗ − αn+1 ) = γ(β) · L(α∗ − αn+1 ). ′ α∗ − αn+1

(A.12)

Lipschitz continuity therefore constrains J(πfuture ) through the teacher’s distance to π ∗ alone, irrespective of which side of π ∗ the teacher falls on; this is what makes Eq. A.10 valid for overshooting as well as non-overshooting β. We then apply Theorem 2 with teacher πT = πfuture and student πθ = πθn+1 , the checkpoint produced by OPD. Substituting the teacher gap bounded above gives q J(π ∗ ) − J(πθn+1 ) ≤ J(π ∗ ) − J(πfuture ) + T L̄OPD /2, (A.13) | {z } ≤ γ(β)L(α∗ −α′n+1 )

′ ′ which is Eq. A.9. Finally, γ(β) < 1 holds iff |α∗ − βαn+1 | < α∗ − αn+1 , which by Proposition 3 is exactly the ∗ ′ condition 1 < β < 2α /αn+1 − 1. Therefore, once the distillation error vanishes, the post-OPD bound is γ(β) times the pre-OPD bound, hence strictly smaller.

Remark 2 (Limitations of the linear assumption). Proposition 3 and Corollary 4 assume a perfectly linear trajectory in φ-space. In practice, linearity holds only approximately and locally: empirical studies of task vectors and linear mode connectivity (Ilharco et al., 2022; Frankle et al., 2020) show that interpolation and modest extrapolation preserve performance, but large extrapolation factors β can exit the linear regime and degrade the teacher quality. The bounds above should therefore be understood as characterizing the mechanism by which extrapolation helps—a teacher closer to π ∗ yields a tighter suboptimality bound—rather than as providing exact convergence rates. In our experiments, we find that moderate values of β consistently improve over self-distillation, consistent with this local-linearity picture. A.3

How Low-Dimensional Are Our Trajectories?

We quantify the gap between the linear assumption and the actual training dynamics, refining the low-rank picture reported in prior work (Cai et al., 2025; Wang et al., 2026a) with a direct measurement on our own runs. For GRPO checkpoints of Qwen3-1.7B and Qwen3-8B on DAPOMath, we form the displacement δt = θt − θ0 at each saved step and compute the Gram matrix ⟨δi , δj ⟩ over all parameter tensors; its eigenspectrum gives the principal components of the trajectory (Fig. A.1). An exactly linear trajectory would be rank one, placing all variance in a single component. The leading component accounts for 68.2% at both scales, and the top three for 86.6% and 88.5% respectively. That the two agree to within 0.1% on the leading component despite a 5× difference in parameter count suggests this is a property of the training dynamics rather than of a particular model. A single direction therefore does not capture the trajectory exactly: the residual 31.8% lies orthogonal to the leading component. What the measurement does show is that the trajectory is confined to a genuinely low-dimensional subspace: three components capture ∼87% of the variance at both scales, out of a parameter space with billions of 20

Variance explained (%)

Preprint | Salesforce AI Research

100

90%

80

Qwen3-1.7B Qwen3-8B cumulative

60 40 20 0

1

2

3

4

5

Principal component

6

Figure A.1: The RLVR weight trajectory is low-dimensional. Principal spectrum of the displacements δt = θt − θ0 from GRPO checkpoints on DAPOMath. Bars show variance explained per component; lines show the cumulative total. A perfectly linear trajectory would place 100% on the first component; the observed 68.2% reflects curvature along the path, while three components capture ∼87% at both scales, confining the trajectory to a low-dimensional subspace. degrees of freedom. This confinement bounds the degree of non-linearity—within a rank-3 subspace the leading direction accounts for ∼78% of the subspace variance, limiting the magnitude of orthogonal curvature. For RISE’s operating range (β ≤ 1.2, i.e. extending 20% beyond the current displacement), the error introduced by the linear approximation is therefore controlled by the small residual variance in the orthogonal components. Fig. 6 provides empirical confirmation: extrapolation gains persist out to β ≈ 1.3 and the safe range contracts as the policy converges, consistent with Proposition 3’s prediction even though its linear premise holds only approximately. A.4

Top-K Approximation Bias

′ As described in §3.2, we approximate the logit-space extrapolated teacher by selecting S = TopK (πθn+1 ) and ′ projecting the post-RLVR checkpoint πθn+1 and the anchor πθn onto the same (K + 1)-simplex before applying the extrapolation (Eq. 8). This differs from the unbiased procedure, which would extrapolate over the full vocabulary first and then select S ∗ = TopK (πfuture ). The approximation error is captured by the set difference ′ ′ S△S ∗ : tokens outside πθn+1 ’s top-K but with a large log-ratio log πθn+1 /πθn can be promoted into S ∗ by extrapolation, yet are absent from S. ′ In practice this discrepancy is negligible. With K = 100, the top-K tokens of πθn+1 account for > 99% of the probaP ′ bility mass under typical autoregressive LLM distributions, leaving a residual tail mass ε = 1 − v∈S πθn+1 (v) < 0.01. This does not formally bound |S△S ∗ |, since extrapolation with β > 1 can in principle amplify tail tokens with large log-ratios. However, for the moderate β values used in our experiments, almost all of πfuture ’s mass concentrates on tokens already in S, so the induced bias is negligible in practice. For weight-space extrapolation, no such approximation arises: the forward pass through θfuture directly produces πfuture , and top-K is applied to the exact extrapolated distribution.

B

Experiment Setup

This appendix supplements the setup summary in §4 with full training hyperparameters (§B.1), evaluation protocol details (§B.2), and a description of the OPSD baselines and their shared privileged-teacher construction (§B.3). B.1

Training Setup

All runs use the VeRL framework on a single node with 8 GPUs, training for one epoch over the respective training set. We use AdamW with a constant learning rate and no warmup, GRPO advantage estimation without 21

Preprint | Salesforce AI Research

standard-deviation normalization, and token-level importance-sampling correction (clip threshold 2.0). Prompts are truncated to 2,048 tokens and responses to 8,192 tokens. All methods—GRPO, the three OPSD baselines, and both RISE variants—share the hyperparameters in Table B.1 within each configuration, so performance differences are attributable to the teacher construction rather than training dynamics. Table B.1: Training hyperparameters. n denotes the number of rollouts sampled per prompt. Train batch size is the number of prompts sampled per iteration for rollout generation; mini-batch size is the number of prompts per gradient update within each iteration (the train batch is split into train batch size/mini-batch size gradient steps). The OPD phase uses the same mini-batch size as the RLVR phase. Qwen2.5-3B agentic runs follow GIGPO’s setup (Feng et al., 2025); see their paper for full details. Qwen3-1.7B Qwen3-1.7B-Base Qwen3-8B Qwen3-4B-Base Qwen3-8B-Base OLMo3-7B-Instruct-SFT (DAPOMath) (DAPOMath) (DAPOMath) (multi-domain) (code) (OpenR1-Math)

B.2

Learning rate Train batch size Mini-batch size Rollouts n Epochs Rollout temperature KL coefficient

1e-6 64 32 8 1 1.0 0

1e-6 64 32 8 1 1.0 0

1e-6 64 32 8 1 1.0 0

1e-6 64 32 8 1 1.0 0

1e-6 64 32 8 1 1.0 0

1e-6 128 64 8 1 1.0 0

RISE β0 RISE βN RISE η RISE K

1.2 1.0 0.1 100

1.2 1.0 0.1 100

1.2 1.0 0.1 100

1.2 1.0 0.1 100

1.2 1.0 0.1 20

1.2 1.0 1.0 100

Evaluation Setup

Benchmarks. Mathematical reasoning is evaluated on MATH-500 (Lightman et al., 2024), AIME 2024/2025 (Jia, 2024; math ai, 2025), AMC 2023 (MAA, 2023), Minerva (Lewkowycz et al., 2022), and OlympiadBench (He et al., 2024), with GPQA-Diamond (Rein et al., 2023), IFEval (Zhou et al., 2023), and MMLU-Pro (Wang et al., 2024) as out-of-distribution checks. The multi-domain suite adds SuperGPQA (Du et al., 2026) and TheoremQA (Chen et al., 2023). Code generation is evaluated on HumanEval+, MBPP+ (Liu et al., 2023), and LiveCodeBench v6 (Jain et al., 2025). Agentic tasks use ALFWorld (Shridhar et al., 2020) (success rate: fraction of episodes completed) and WebShop (Yao et al., 2022), which reports two metrics: Score (100× average reward, where reward is a product of attribute recall, option selection, price satisfaction, and category match) and Acc (success rate: fraction of episodes achieving a perfect reward of 1.0). Sampling protocol. Evaluation samples are drawn stochastically with temperature 1.0, top-p = 1.0, and top-k = −1 (i.e., unrestricted), matching the rollout distribution used during training. The number of samples per prompt is set per benchmark according to its size and variance: 16 for AIME 2024/2025, GPQA-Diamond, and Minerva, 8 for OlympiadBench, and 4 for MATH-500 and AMC 2023. Deterministic benchmarks that admit a single correct completion—IFEval, MMLU-Pro, SuperGPQA, and TheoremQA—use a single sample. All Qwen3 models are evaluated in non-thinking mode (thinking is disabled). For code generation we use 4 samples on HumanEval+, MBPP+, and LiveCodeBench. We report two complementary metrics. avg@N is the mean accuracy across N sampled completions, measuring expected single-sample quality. pass@N estimates the probability that at least one of N independent draws is correct: for each problem we draw 1,000 bootstrap subsets of size N (with replacement) from the N sampled completions, record whether each subset contains a correct solution, and average; this per-problem estimate is then averaged across prompts. Reporting both separates per-sample reliability from exploration breadth: a method can improve avg@N while leaving pass@N unchanged (sharpening the distribution) or improve both (genuinely expanding coverage). B.3

Baselines

All three OPSD baselines—GRPO+SDPO (Hübotter et al., 2026), SDAR (Lu et al., 2026), and RLSD (Yang et al., 2026a)—share the same privileged teacher and differ only in how its signal enters the update. 22

Preprint | Salesforce AI Research

Privileged teacher construction. For each prompt, the model generates n rollouts. If the group contains at least one correct solution, the teacher is the model re-prompted with that correct solution as privileged context, and its token distribution serves as the distillation target for all responses to that prompt. If no rollout in the group succeeds, the prompt contributes only the RLVR loss—no distillation signal is available. The teacher is maintained as an exponential moving average of the student (rate 0.05), following the recommended hyperparameters in Hübotter et al. (2026). Averaged over training, the teacher is active on ∼85% of prompts for Qwen3-8B and ∼70% for Qwen3-1.7B, so the baselines’ limited gains are not attributable to low coverage. Integration strategies. GRPO+SDPO adds a KL divergence toward the teacher’s token distribution as an auxiliary loss alongside the policy gradient. SDAR uses the teacher only as a gate: the teacher–student probability gap scales the GRPO advantage per token, upweighting positions where the teacher is more confident than the student. RLSD similarly uses the teacher’s per-token signal to reweight advantage estimates, providing finer-grained credit assignment than a single sequence-level reward. Contrast with RISE. All three baselines require privileged information—a correct solution the student did not produce—so their teacher is available only on problems where such a solution exists in the rollout group. RISE derives its teacher from the training trajectory itself and therefore applies uniformly to every prompt, regardless of whether any rollout succeeded.

C

Experiment Results

This appendix provides supplementary results for the experiments in §4: Qwen3-1.7B-Base results (§C.1), permetric convergence curves (§C.2), code generation convergence (§C.3), grounding and OPD ablation details (§C.5), extrapolation strength and schedule ablations (§C.6–§C.8), a compute-matched comparison (§C.9), and a multi-seed reproducibility analysis (§C.10). C.1

Qwen3-1.7B-Base Results

Table C.1 reports results for Qwen3-1.7B-Base on DAPOMath. RISE (logit) achieves the best in-domain average (29.2) and competitive OOD performance, while RISE (weight) leads on MATH-500 (62.4) and Minerva (20.7). The margin over GRPO is the narrowest of any configuration (+1.2 Math Avg), consistent with the capability ceiling of a base model that has not been instruction-tuned. Table C.1: Qwen3-1.7B-Base (DAPOMath). Accuracy (%) on math and OOD benchmarks. Best in bold, second best underlined. Method Base GRPO GRPO+SDPO SDAR RLSD RISE (logit) RISE (weight)

C.2

In-Domain Math

OOD

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench Avg. GPQA IFEval MMLU Avg.

40.0 60.6 61.3 55.6 57.2 59.6 62.4

4.4 9.4 8.3 9.6 11.9 12.1 10.4

2.9 6.9 6.7 8.3 7.5 8.1 5.8

27.5 42.5 41.9 43.1 43.1 45.6 38.8

11.6 19.3 20.1 17.7 17.8 20.2 20.7

19.8 29.2 32.3 27.1 28.2 29.4 30.8

17.7 28.0 28.4 26.9 27.6 29.2 28.1

12.7 26.5 28.7 23.2 23.0 25.4 25.7

14.4 31.1 30.5 28.5 29.0 31.1 29.9

8.0 39.6 40.5 42.5 42.4 42.9 41.0

Per-Metric Convergence Curves

Figure C.1 shows per-metric training curves for all model configurations discussed in §4.2.

23

11.7 32.4 33.2 31.4 31.5 33.1 32.2

Preprint | Salesforce AI Research

Qwen3-1.7B

— AIME'24 (avg@16)

Qwen3-1.7B

— AIME'25 (avg@16)

Qwen3-1.7B

80

— AMC'23 (avg@4)

Qwen3-1.7B

25 GRPO

20

RISE (logit) RISE (weight)

15

SDAR 0

50

100

150

25 20

GRPO RISE (logit)

15

RISE (weight)

200

250

70 60

GRPO RISE (logit)

50

RISE (weight)

SDAR

10 0

50

Training Steps

100

150

200

Accuracy (%)

30

Accuracy (%)

30

Accuracy (%)

Accuracy (%)

35

SDAR 40

250

0

50

100

Training Steps

150

200

— Minerva (avg@16)

24 22 GRPO

20

RISE (logit) RISE (weight)

18

250

SDAR 0

50

100

Training Steps

150

200

250

Training Steps

(a) Qwen3-1.7B (DAPOMath)

— AIME'24 (avg@16)

Qwen3-8B

— AIME'25 (avg@16)

40

GRPO RISE (logit) RISE (weight)

30

SDAR 0

50

100

150

40 35 GRPO

30

RISE (logit) RISE (weight)

25

250

80 GRPO RISE (logit)

70

RISE (weight)

SDAR

20

200

Qwen3-8B 32.5

90

Accuracy (%)

Accuracy (%)

Accuracy (%)

50

— AMC'23 (avg@4)

Qwen3-8B

45

0

50

Training Steps

100

150

200

Accuracy (%)

Qwen3-8B

60

30.0 27.5

0

50

100

Training Steps

150

200

GRPO

25.0

RISE (logit) RISE (weight)

22.5

SDAR 250

— Minerva (avg@16)

SDAR

20.0

250

0

50

100

Training Steps

150

200

250

Training Steps

(b) Qwen3-8B (DAPOMath)

— MATH-500 (avg@4)

OLMo3-7B

— AIME'24 (avg@16)

OLMo3-7B

— AMC'23 (avg@4)

OLMo3-7B

28

— Minerva (avg@16)

70

GRPO RISE (logit)

65

RISE (weight)

60 0

50

100

150

200

250

30 GRPO 20

RISE (logit) RISE (weight)

10

SDAR 300

Accuracy (%)

75

40

350

SDAR 0

50

100

Training Steps

150

200

250

300

70 60

GRPO RISE (logit)

50

RISE (weight) SDAR

40 350

Accuracy (%)

80

Accuracy (%)

Accuracy (%)

OLMo3-7B 80

0

50

100

Training Steps

150

200

250

300

26 24 GRPO

22

RISE (logit) RISE (weight)

20

350

SDAR 0

Training Steps

50

100

150

200

250

300

350

Training Steps

(c) OLMo3-7B-Instruct-SFT (OpenR1Math) — AIME'24 (avg@16)

Qwen3-4B-Base (Mixed)

— AIME'25 (avg@16)

Qwen3-4B-Base (Mixed) 70

GRPO GRPO+SDPO

15

RISE (logit)

20 15

GRPO GRPO+SDPO

10

RISE (logit)

RISE (weight)

10 0

50

100

150

200

250

Training Steps

300

350

Accuracy (%)

20

Accuracy (%)

Accuracy (%)

25 25

(Mixed) — OlympiadBench (avg@8) — AMC'23 (avg@4) Qwen3-4B-Base 45

60 GRPO

50

GRPO+SDPO RISE (logit)

40

RISE (weight)

RISE (weight) 0

50

100

150

200

250

300

0

350

50

100

150

200

250

300

Training Steps

Training Steps

350

Accuracy (%)

Qwen3-4B-Base (Mixed)

40 35 30 0

50

GRPO GRPO+SDPO RISE (logit) RISE (weight) 100 150 200 250 300 350

Training Steps

(d) Qwen3-4B-Base (mixed math + STEM)

Figure C.1: Per-metric convergence curves for all model configurations. Each row shows four representative benchmarks. RISE (both variants) consistently reaches higher accuracy in fewer training steps. C.3

Code Generation Convergence Curves

Figure C.2 shows avg@4 and pass@4 convergence curves for Qwen3-8B-Base trained on Skywork-OR1-Code, evaluated on HumanEval+, MBPP+, and LiveCodeBench. Both RISE variants converge faster than GRPO in early training across all three benchmarks. C.4

Average vs. Best-of-N Accuracy

Table C.2 reports avg@16 and pass@16 on the AIME benchmarks, isolating whether RISE’s gains come from sharpening the policy around solutions it already finds or from expanding the set of solvable problems. If distillation merely concentrated probability mass on existing solutions, avg@16 would improve while pass@16 stagnated or declined. Instead, RISE improves both, and at 1.7B the pass@16 gains exceed the avg@16 gains on AIME’24 (+9.6 vs. +6.3 for RISE (logit)), indicating genuinely broadened coverage. At 8B the coverage margins are smaller, consistent with a stronger base policy leaving less headroom; the one regression is RISE (logit) on AIME’25 pass@16 (−0.6), where the weight-space variant instead gains +3.9.

24

Preprint | Salesforce AI Research

60 50 40

0

20

40

60

Training Steps

GRPO RISE (logit) RISE (weight) 80 100

60 55 50 GRPO 45

RISE (logit) RISE (weight)

40 0

20

(a) HumanEval+ (avg@4)

80 70 60

0

20

40

60

Training Steps

GRPO RISE (logit) RISE (weight) 80 100

40

60

80

— LiveCodeBench (avg@4)

50 45 40 GRPO 35

RISE (logit) RISE (weight)

30

100

0

20

40

60

(b) MBPP+ (avg@4)

(c) LiveCodeBench (avg@4)

— MBPP+ (pass@4)

70.0 67.5 65.0 GRPO

62.5

RISE (logit)

Qwen3-8B-Base (Code)

— LiveCodeBench (pass@4)

60 55 GRPO

50

RISE (logit) RISE (weight)

RISE (weight)

60.0

(d) HumanEval+ (pass@4)

100

Training Steps

72.5

0

80

Training Steps

Qwen3-8B-Base (Code) Accuracy (%)

Accuracy (%)

Qwen3-8B-Base (Code) — HumanEval+ (pass@4) 90

Qwen3-8B-Base (Code) Accuracy (%)

65

Accuracy (%)

Accuracy (%)

70

— MBPP+ (avg@4)

Qwen3-8B-Base (Code)

80

Accuracy (%)

Qwen3-8B-Base (Code) — HumanEval+ (avg@4)

20

40

60

80

100

0

20

40

60

80

100

Training Steps

Training Steps

(e) MBPP+ (pass@4)

(f) LiveCodeBench (pass@4)

Figure C.2: Code generation convergence curves (Qwen3-8B-Base). Both RISE variants consistently converge faster than GRPO across all three benchmarks under both avg@4 and pass@4 metrics. Table C.2: Average vs. best-of-16 accuracy (%) on AIME benchmarks. ∆ denotes improvement over GRPO. RISE improves pass@16 (coverage) alongside avg@16 (mean quality), with the largest coverage gains at 1.7B. AIME’24 AIME’25 avg@16 pass@16 avg@16 pass@16

Model

Method

Qwen3-8B

GRPO 54.4 78.7 42.9 63.3 RISE (logit) 56.9 +2.5 79.0 +0.3 46.7 +3.8 62.7 -0.6 RISE (weight) 58.1 +3.7 81.5 +2.8 45.8 +2.9 67.2 +3.9

GRPO 30.0 54.4 26.3 45.8 Qwen3-1.7B RISE (logit) 36.3 +6.3 64.0 +9.6 33.3 +7.0 53.9 +8.1 RISE (weight) 32.9 +2.9 57.0 +2.6 31.3 +5.0 49.2 +3.4

C.5

Grounding and OPD Ablation

Figure C.3 plots training reward and policy entropy for the grounding ablation (§4.3). Without RLVR, mean training reward collapses from 0.56 to 0.00 by step 60, tracking the accuracy cliff in Fig. 5a. Policy entropy spikes to 0.78 at step 49—just before the final crash—then drops to near zero, consistent with the policy degenerating into repetitive non-terminating generation. GRPO and RISE show stable entropy throughout. Table C.3 reports the full without-OPD ablation results discussed in §4.3. Across all three model scales, removing the OPD phase (i.e., directly adopting θfuture without distillation) leaves Math Avg essentially unchanged relative to GRPO (60.0→60.3 at 8B, 45.4→45.6 at 1.7B, 28.0→27.7 at 1.7B-Base), forfeiting the gains that RISE achieves. Per-benchmark accuracy shifts—e.g., the no-OPD arm gains on OlympiadBench at 8B but loses on the AIME pair, while the pattern reverses at 1.7B—but these redistributions do not translate into consistent improvement without the distillation phase. C.6

Extrapolation Strength and Decay Schedule

Table C.4 reports per-benchmark results for the β0 and decay schedule ablations summarised in Table 4. All runs use logit-space extrapolation.

25

Training Reward

Policy Entropy GRPO

Policy Entropy

Mean Training Reward

Preprint | Salesforce AI Research

0.6 GRPO

0.4

RISE (logit) w/o RLVR

0.2

RISE (logit)

0.6

w/o RLVR

0.4 0.2

0.0 0

20

40

60

80

100

120

0

20

Training Steps

40

60

80

100

120

Training Steps

Figure C.3: Grounding ablation: training reward and policy entropy. Same three Qwen3-8B runs as Fig. 5a. Without RLVR (grey), training reward collapses to zero (left) while policy entropy spikes then crashes (right), confirming the phase-transition nature of the collapse. Table C.3: Without-OPD ablation. Accuracy (%) on math benchmarks. The w/o OPD variant adopts θfuture directly without distillation. All runs are configuration-matched to their RISE counterpart.

C.7

Model

Method

Qwen3-8B

GRPO RISE (logit) RISE (weight) w/o OPD

83.8 84.8 84.4 84.6

54.4 56.9 58.1 51.7

42.9 46.7 45.8 39.8

89.4 91.6 91.9 90.6

31.5 32.5 32.6 31.2

57.9 62.4 63.3 64.1

60.0 62.5 62.7 60.3

GRPO RISE (logit) Qwen3-1.7B RISE (weight) w/o OPD

75.6 75.2 77.3 72.0

30.0 36.3 32.9 32.3

26.3 33.3 31.3 27.3

65.6 78.8 75.0 70.0

23.9 24.9 25.7 24.6

51.1 52.9 53.3 47.4

45.4 50.2 49.2 45.6

GRPO Qwen3-1.7B RISE (logit) -Base RISE (weight) w/o OPD

60.6 59.6 62.4 58.8

9.4 12.1 10.4 10.6

6.9 8.1 5.8 5.8

42.5 45.6 38.8 43.1

19.3 20.2 20.7 19.4

29.2 29.4 30.8 28.6

28.0 29.2 28.1 27.7

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench

Avg.

Resample vs. Rollout Reuse

Table C.5 reports per-benchmark results for the resample ablation discussed in §4.4. All runs use logit-space extrapolation with β0 = 1.2 and linear decay. C.8

Anchor Dynamics (η) Ablation

Table C.6 reports per-benchmark results for the anchor ablation discussed in §4.4. All runs use β0 = 1.2. On OLMo3-7B-Instruct-SFT, the drop from EMA smoothing is substantial (−6.3 logit, −2.7 weight), likely because GRPO on OLMo produces large, stable per-step displacements that do not benefit from smoothing. C.9

Compute-Matched Comparison

Table C.7 reports per-benchmark results for the compute-matched comparison discussed in §4.4. GRPO-2× performs a second inner-loop gradient pass on the same rollouts within each iteration, matching the total gradient budget of RISE’s RLVR + OPD phases.

26

Preprint | Salesforce AI Research

Table C.4: Detailed β0 and decay schedule ablation results. Accuracy (%) on individual math benchmarks. The top block varies β0 with linear decay; the middle block varies the schedule with β0 = 1.2. Both blocks use Qwen3-8B. The bottom block varies β0 on Qwen3-1.7B-Base (schedule noted in parentheses). Configuration

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench

Avg.

Qwen3-8B — β0 sensitivity (linear decay) 84.8 56.9 46.7 β0 = 1.2 β0 = 1.3 83.3 56.9 47.3 β0 = 1.5 83.5 55.2 47.9

91.6 92.5 91.2

32.5 32.1 31.9

62.4 61.9 61.6

62.5 62.3 61.9

Qwen3-8B — decay schedule (β0 = 1.2) Linear 84.8 56.9 46.7 84.3 58.5 47.7 Cosine Exponential 82.4 57.5 47.1 83.8 56.2 44.4 Fixed (β = β0 )

91.6 91.8 91.2 90.6

32.5 32.0 31.7 31.6

62.4 63.8 61.8 63.9

62.5 63.0 62.0 61.8

Qwen3-1.7B-Base β0 = 1.2 (linear) β0 = 1.2 (cosine) β0 = 1.3 (linear) β0 = 1.5 (linear)

45.6 42.5 46.9 45.6

20.2 19.8 19.1 19.5

29.4 29.7 29.4 28.6

29.2 28.3 29.4 28.5

59.6 58.6 61.0 59.2

12.1 10.8 11.7 11.0

8.1 8.3 8.3 6.9

Table C.5: Resample vs. rollout reuse. Accuracy (%) on individual math benchmarks. “Resample” generates fresh ′ rollouts from πθn+1 for OPD; “Rollout reuse” reuses RLVR rollouts.

C.10

Model

Variant

Qwen3-8B

Rollout reuse Resample

84.8 83.7

56.9 58.5

46.7 47.9

91.6 90.6

32.5 31.3

62.4 63.1

62.5 62.5

Qwen3-1.7B Rollout reuse -Base Resample

59.6 59.6

12.1 11.9

8.1 7.5

45.6 43.8

20.2 18.8

29.4 29.0

29.2 28.4

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench

Avg.

Reproducibility Across Random Seeds

Table C.8 reports mean ± standard deviation across three random seeds for Qwen3-1.7B and Qwen3-8B on DAPOMath. Table 1 reports a single representative seed; the multi-seed statistics confirm that RISE’s gains are well outside seed-to-seed variance at both scales. At 1.7B, RISE (logit) improves Math Avg by +4.5 over GRPO (49.6 vs. 45.1), roughly 10× the per-method standard deviation (≤ 0.6). At 8B, RISE (weight) improves Math Avg by +2.2 over GRPO (62.4 vs. 60.1) and by +2.8 over SDAR (59.6), again far exceeding the per-method standard deviation (≤ 0.6). Individual competition benchmarks (AIME, AMC) exhibit larger variance due to small problem sets, but the aggregate Math Avg is stable across all methods.

27

Preprint | Salesforce AI Research

Table C.6: Anchor ablation. Accuracy (%) on individual math benchmarks. η = 1 uses the previous checkpoint as anchor; η < 1 uses an EMA anchor. EMA helps on Qwen3-1.7B but degrades on OLMo3-7B-Instruct-SFT. Configuration

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench

Avg.

Qwen3-1.7B — Logit-space η = 1.0 77.4 29.4 75.4 32.7 η = 0.95 η = 0.3 74.6 33.3 75.2 36.3 η = 0.1

26.0 29.8 29.4 33.3

71.3 73.1 76.9 78.8

24.8 25.0 25.5 24.9

48.6 50.2 50.6 52.9

46.2 47.7 48.4 50.2

Qwen3-1.7B — Weight-space η = 1.0 75.8 31.3 η = 0.3 76.3 31.3 77.3 32.9 η = 0.1

29.2 29.0 31.3

67.5 73.1 75.0

24.5 24.8 25.7

50.0 52.5 53.3

46.4 47.8 49.2

OLMo3-7B-Instruct-SFT — Logit-space η = 1.0 82.9 46.9 36.0 83.1 η = 0.1 78.4 31.3 32.3 76.9

27.6 25.4

61.9 56.2

56.4 50.1

OLMo3-7B-Instruct-SFT — Weight-space 82.8 42.5 32.5 76.9 η = 1.0 80.6 35.0 31.0 75.0 η = 0.1

27.8 26.6

59.6 57.9

53.7 51.0

Table C.7: Compute-matched comparison: RISE vs. GRPO-2×. Accuracy (%) on math benchmarks. GRPO-2× performs a second inner-loop gradient pass on the same rollouts within each iteration, matching RISE’s total gradient budget at equal sampling cost. Model

Method

Qwen3-8B

GRPO GRPO-2× RISE (logit) RISE (weight)

83.8 84.5 84.8 84.4

54.4 53.5 56.9 58.1

42.9 42.3 46.7 45.8

89.4 92.5 91.6 91.9

31.5 31.4 32.5 32.6

57.9 58.9 62.4 63.3

60.0 60.5 62.5 62.7

GRPO GRPO-2× Qwen3-1.7B RISE (logit) RISE (weight)

75.6 74.5 75.2 77.3

30.0 32.9 36.3 32.9

26.3 30.6 33.3 31.3

65.6 70.6 78.8 75.0

23.9 25.6 24.9 25.7

51.1 49.7 52.9 53.3

45.4 47.3 50.2 49.2

MATH500 AIME24 AIME25 AMC23 Minerva OlyBench

Avg.

Table C.8: Multi-seed reproducibility (DAPOMath): mean ± std across 3 random seeds. Standard deviations are shown in gray. RISE’s Math Avg gains over GRPO far exceed seed-to-seed variance at both scales. Method

In-Domain Math MATH500

AIME24

AIME25

AMC23

OOD

Minerva

OlyBench

Avg.

GPQA

IFEval

MMLU

Avg.

Qwen3-8B (DAPOMath) GRPO 84.0 ±0.5 53.1 ±1.2 42.5 ±0.3 88.5 ±2.1 31.4 ±0.4 61.3 ±2.7 60.1 ±0.2 56.3 ±0.5 82.2 ±0.5 71.6 ±1.2 70.0 ±0.4 SDAR 84.7 ±0.4 48.9 ±1.0 39.1 ±1.2 90.0 ±1.3 30.6 ±0.6 64.2 ±0.7 59.6 ±0.6 55.3 ±0.3 83.2 ±0.6 71.4 ±1.5 69.9 ±0.6 RISE (weight) 84.9 ±0.5 55.8 ±1.7 44.1 ±1.2 92.8 ±0.6 32.3 ±0.2 64.4 ±0.9 62.4 ±0.2 58.1 ±0.7 84.1 ±0.5 72.3 ±1.0 71.5 ±0.4 Qwen3-1.7B (DAPOMath) GRPO 75.4 ±0.1 28.2 ±1.3 26.2 ±0.3 66.5 ±1.6 24.1 ±0.5 50.4 ±0.6 45.1 ±0.4 31.2 ±0.8 68.7 ±0.2 54.4 ±0.7 51.4 ±0.1 RISE (logit) 75.3 ±0.2 34.5 ±1.5 31.7 ±1.5 78.8 ±2.0 25.0 ±0.1 52.4 ±0.7 49.6 ±0.4 34.3 ±0.5 69.1 ±1.0 55.5 ±1.5 53.0 ±0.7 RISE (weight) 76.5 ±0.6 32.3 ±0.8 30.2 ±0.8 74.8 ±0.5 25.1 ±0.5 51.7 ±1.3 48.4 ±0.6 35.9 ±1.1 68.7 ±1.4 55.3 ±1.5 53.3 ±1.0

28

Record · ID 660854 · SHA-256 d9ce6bd6239f73c1
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.