Conceptio › Archive › arXiv CS
arXiv CSopen access

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Preprint.

N EGATIVE S ELF -D ISTILLATION : L EARNING TO R EASON BY AVOIDING F LAWS Rongcan Pei1 , Zhepei Wei1 , Shuyao Xu2 , Xinyu Zhu1 , Wei-Lin Chen1 , and Yu Meng1 Department of Computer Science, University of Virginia 2 Stanford University {peirongcan,zhepei.wei,xinyuzhu,wlchen,yumeng5}@virginia.edu [email protected] 1

arXiv:2609.11699v1 [cs.CL] 10 Sep 2026

GitHub

Hugging Face

A BSTRACT On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (e.g., acting as a “careless reasoner”) and pushes the student’s distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model’s foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model’s linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines. Across seven mathematical reasoning benchmarks (AIME 24/25/26, HMMT, AMC, OlympiadBench, and MATH), NSD achieves average gains of 2.3%, 7.5%, and 6.0% for 1.7B, 4B, and 8B models, respectively. Further analyses show that NSD achieves higher training efficiency while preserving the self-correction behaviors crucial for complex reasoning.

“Please reason carelessly by wrongly calculating the prime factorization of …”

Student Question

LLM

Negative Teacher

Probabilities of key tokens

Student

Negative Condition

▲

(Question Specific)

▲

Negative Teacher

Decrease this token's probability

▲

… The prime factorization is

Before

After

Figure 1: Overview of the NSD framework. (Left) We construct a negative teacher from the same base model via self-generated negative conditioning. (Right) The student model is optimized to diverge its distribution from that of the negative teacher.

1

Preprint.

1

I NTRODUCTION

Reinforcement Learning with Verifiable Rewards (RLVR) (Shao et al., 2024; Yu et al., 2025; Lambert et al., 2025) has emerged as an effective paradigm for enhancing the reasoning capabilities of large language models (LLMs). However, RLVR is often bottlenecked by computational inefficiency and training signal sparsity. These challenges arise because (1) sampling multiple rollouts per query is expensive, and rollouts within a group frequently receive identical rewards on exceptionally easy or difficult problems, leading to advantage collapse and vanishing gradients (Liao et al., 2026; Xu et al., 2026a; Zhang et al., 2025b); and (2) outcome-based rewards are applied uniformly across the entire generated sequence, which obscures fine-grained, token-level credit assignment. To mitigate these limitations, On-Policy Distillation (OPD) (Agarwal et al., 2024; Lu & Lab, 2025; Song & Zheng, 2026) utilizes a stronger, external teacher model to provide dense token-level supervision over the student model’s self-sampled reasoning trajectories. While this approach successfully yields richer feedback, it introduces a practical constraint: obtaining a strictly superior external teacher that is both sufficiently capable of providing accurate dense supervision and compatible with the student’s tokenizer is often impractical. To circumvent the reliance on external teacher models, On-Policy Self-Distillation (OPSD) (Zhao et al., 2026a; Shenfeld et al., 2026; Hübotter et al., 2026) has been proposed as a scalable alternative. In OPSD, the model acts as its own teacher by utilizing privileged information (e.g., ground-truth answers) to generate dense supervision signals for the student’s self-sampled trajectories. However, because the OPSD teacher inherently knows the ground-truth solution, it tends to produce artificially confident and highly linear reasoning trajectories (Kim et al., 2026b; Harne et al., 2026). Consequently, forcing the student to minimize the divergence from this teacher distribution inadvertently suppresses high-entropy exploration, expressions of uncertainty, and self-corrective behaviors, which are essential for complex problem-solving. Motivated by the observation that imitating a synthetically confident oracle can degrade natural reasoning processes, we explore an alternative training paradigm: optimizing the model to explicitly avoid flawed reasoning patterns. We introduce Negative Self-Distillation (NSD), a fully selfbootstrapped framework that operates without external privileged data. Instead of utilizing a teacher conditioned on the correct answer, the model is prompted to generate a question-specific negative condition (e.g., acting as a “careless reasoner”) to instantiate a negative teacher. The student is then optimized to move its token distribution away from the negative teacher, encouraging it to avoid premature conclusions and other flawed reasoning patterns. Importantly, the negative signal is generated from the model itself and does not require ground-truth solutions or external annotations. A central challenge, however, is that not every token assigned high likelihood by the negatively conditioned teacher corresponds to a reasoning error. A naive divergence or unlikelihood objective (Welleck et al., 2020) can also penalize ordinary linguistic tokens, degrading the model’s pretrained linguistic priors. NSD therefore introduces a dynamic token-level gating mechanism that compares the negative teacher with a benign reference model and activates the negative objective only when the negative condition increases the likelihood of the sampled token. We further stabilize these updates with a bounded unlikelihood formulation and a KL-based regularization term, preventing excessive updates on high-confidence structural tokens while retaining targeted supervision on reasoning-critical tokens. This design yields a training signal that is both selective and computationally efficient: NSD requires only a single student rollout per sample, avoids full-vocabulary logit alignment, and can parallelize the reference and negative-teacher computations. Beyond accuracy, our analysis shows that NSD preserves and strengthens reflective self-correction behavior rather than encouraging overly confident, linear reasoning. Our main contributions are summarized as follows: • We propose Negative Self-Distillation (NSD), a label-free, fully self-bootstrapped framework that enhances reasoning capabilities by optimizing the model to diverge from self-generated flawed trajectories, eliminating the need for ground-truth solutions or an external teacher. • We introduce a token-level gating mechanism together with a bounded unlikelihood objective, enabling targeted divergence from flawed reasoning while preserving foundational language priors. • We demonstrate that NSD consistently outperforms existing training paradigms (i.e., OPSD (Zhao et al., 2026a), Intuitor (Zhao et al., 2026b) and TTRL (Zuo et al., 2025)) across 1.7B, 4B, and 8B model sizes on seven reasoning tasks. Furthermore, NSD achieves superior training efficiency, mitigates overconfidence, and preserves the model’s intrinsic reflection capabilities. 2

Preprint.

2

NSD: N EGATIVE S ELF -D ISTILLATION

Negative Self-Distillation Dataset {𝒙i } Student 𝜋𝜃 Self-Generate

𝑥i : question

D= 𝑛i : negative condition

Per-Token Gating Function 𝑦0

𝑦1 … 𝑦t−1

Per-Token NSD Objective

𝑦t …

Case 1: 𝐆𝐭 >0

𝐺𝑡 = 𝑚𝑎𝑥 0, 𝜋neg,t − 𝜋ref,t 𝑦0

𝑦1 … 𝑦t−1 𝑦t

Retained

…

Not sensitive to negative condition, filtered Negative Teacher 𝜋neg

Reference Model 𝜋ref

This token's probability is suppressed

Probability on 𝑦t

𝜋ref,t 𝜋neg,t

𝜋θ,t

𝓛(t) = 𝑫KL +𝑮t ⋅ 𝝈[𝐔𝐧𝐥𝐢𝐤𝐞𝐥𝐢𝐡𝐨𝐨𝐝(𝜋θ,t )] Case 2: 𝐆𝐭 =0 𝓛(t) = 𝑫KL

Maintain the original probability

Figure 2: Overview of Negative Self-Distillation. The student model generates negative conditions from the unlabeled training data (Left). We then compare the token distributions between benign and negative contexts, isolating the tokens whose probabilities are abnormally boosted by the negative condition (Mid). Finally, the model is penalized to suppress the probabilities of these isolated tokens, while the filtered benign tokens are regularized only by KL divergence (Right). We consider a label-free training dataset denoted as Draw = {(xi )}N i=1 , where xi represents the problem statement. Our method consists of two core components: self negative conditioning and NSD training. We first prompt the student model πθ to generate a negative condition prompt ni for each problem. By conditioning the model on this prompt ni , we construct a negative teacher πneg . We then penalize the student’s alignment with the teacher under a simple gating mechanism to avoid applying penalty to reasoning-irrelevant tokens. The overview of NSD is shown in Figure 2. 2.1

N EGATIVE C ONDITION P ROMPT G ENERATION

The objective of this module is to allocate a negative instruction ni to each training sample designed to induce flawed reasoning patterns, thereby augmenting the original Draw into a full negativeconditioned dataset D = {(xi , ni )}N i=1 . The negative condition generation strategy should follow the self-generation or easy-to-get principle, without utilizing any gold answer. By default, we adopt an online generation strategy: For a given training problem x, we first sample an initial solution yinit from the student model πθ . Conditioned on both the problem and this initial response, we then prompt the student model to generate an adaptive negative condition n based on its existing reasoning trace (the complete prompt is provided in Appendix C.3.): yinit ∼ πθ (· | x),

n ∼ πθ (· | x, yinit )

(1)

Our framework can naturally accommodate alternative negative condition generation strategies (discussed in Section 4.3). We default to generating negative conditions on the fly during training as it provides the most stable and effective supervision signal. 2.2

BACKGROUND AND C HALLENGES IN U NLIKELIHOOD T RAINING

Our motivation of the NSD training objective is to move the student’s logits distribution away from the negative teacher model’s flawed reasoning behaviors through dense token-level supervision. A natural approach to achieve this is standard unlikelihood training (Welleck et al. (2020)), which minimizes the following objective to suppress the probability of undesirable tokens: Lunlikelihood = − log(1 − πθ (yt | xi , y<t )) However, directly optimizing the objective to distance the student model from the negative teacher’s distribution presents two critical challenges and research questions (RQs): (1) Indiscriminately treating every highly probable token under the negatively conditioned teacher as a flaw and applying the unlikelihood training is problematic, as ordinary grammatical tokens can 3

Preprint.

appear in both normal and flawed reasoning; unlearning them could easily lead to the degradation of fundamental reasoning capabilities. RQ1: How to identify the tokens that represent genuine reasoning flaws? (2) This unbounded unlikelihood objective is catastrophic for highly confident, trivial tokens (e.g., punctuation or spaces) — it triggers loss explosions and overly strong gradient that destabilize training and destroy the model’s inherent logic. Specifically, as πθ → 1, the Lunlikelihood approaches ∞ and the gradient approaches the maximum (as detailed in Appendix A). As a considerable number of tokens have a relatively high probability, this unbounded penalty triggers gradient explosions, also forcing the student to unlearn fixed fundamental linguistic priors (e.g., how to use punctuations) and rapidly update the model parameters in an unstable direction. RQ2: How to formulate a penalty to avoid gradient and loss explosions for training stability? 2.3

T HE NSD T RAINING O BJECTIVE

Token-level adaptive gating. To address RQ1, we propose the gating mechanism to filter out grammatical tokens. During the training phase, the student model generates reasoning trajectories y = (y1 , . . . , yT ) ∼ πθ . To construct the gating signals, we instantiate two frozen teacher models based on the same initial student model: • Reference model (πref ): Conditioned only on the original problem xi , predicting the nominal probability πref (yt | xi , y<t ). • Negative teacher (πneg ): Conditioned on both the problem and the generated negative prompt ni , predicting the negatively biased probability πneg (yt | xi , ni , y<t ). Note that πneg shares the same model weights as πref , differing only by the negative context. We introduce a simple gating function that compares the probabilities of both models to filter out ordinary linguistic tokens and identify the tokens sensitive to the negative injection. For a given student-generated token yt ∼ πθ , the gate is defined as the adjusted positive divergence between the probability of negative and reference model on this token:   Gt = max 0, πneg (yt | xi , ni , y<t ) − πref (yt | xi , y<t ) (2) The gate Gt ∈ [0, 1] acts as an automatic noise filter. If πref ≥ πneg , the token is not activated by a negative condition and naturally exempt from penalization, preserving the model’s original generative distribution. Conversely, if πneg > πref , it indicates that the negative prompt has boosted the token’s likelihood, marking it as a critical target for suppression. Crucially, the penalty weight scales proportionally to this positive gap: a larger divergence directly translates to a heavier penalization. Gated unlikelihood penalty. To formulate a mathematically sound penalty (i.e., the second challenge) and address RQ2, after filtering structural noise via the dynamic gate Gt , we introduce a Sigmoid-bounded unlikelihood penalty: We squash the penalty using a Sigmoid function, yielding 1 2−πθ (yt |xi ,y<t ) . The Gated Unlikelihood (GU) penalty is formulated as:   (t) LGU = Gt · σ − log 1 − πθ (yt | xi , y<t ) = Gt ·

1 2 − πθ (yt | xi , y<t )

(3)

This bounded formulation actively repels the student from negative flaws while safely preserving essential structural tokens. As shown in Figure 3, our sigmoid formulation allocates the strongest unlearning signals to low-to-mid confidence tokens, thereby avoiding gradient explosion on highprobability tokens. We further discuss the GU objective in detail in Section 5.3. Regularization and overall objective. Let πθ (yt | xi , y<t ) denote the current student model being optimized. Our goal is to push the student’s distribution away from the identified vulnerabilities without destroying its fundamental linguistic priors. To further regularize the objective, we introduce a point-wise forward KL penalty evaluated on the sampled token yt . Instead of computing the fullvocabulary KL divergence, which is computationally heavy during rollouts, we apply an empirical 4

Preprint.

10 0

Wait , maybe I should write it as a fraction to be precise . 2 2 . 5 is 4 5 / 2 . again . Let me check that

GU value

0.10

Sigmoid: gate × σ(−log(1 − πs ))

Wait , maybe I should write it as a fraction to be precise . 2 2 . 5 is 4 5 / 2 . again . Let me check that unactivated

0.15

0.05

GU value (log scale)

Raw: gate × (−log(1 − πs ))

0.00

activated

10 −1 10 −2 10 −3 10 −4

Raw (non-sigmoid): gate × (−log(1 − πs )) Sigmoid: gate × σ(−log(1 − πs )) Upper bound (raw, smoothed) Upper bound (sigmoid, smoothed)

0.2

0.4

0.6

πθ

0.8

1.0

Figure 3: Left: After applying the sigmoid function, the gated unlikelihood (GU) values are reduced for basic tokens (e.g., punctuations), preventing gradient explosion. Right: The LGU value distribution over 4,096 tokens from 100 training samples, showing that the sigmoid objective avoids penalization spikes on high-probability tokens, redistributing the LGU weights toward tokens with low-to-mid probabilities in the student model. reference-weighted anchor: (t)

LKL = πref (yt | xi , y<t ) · log

πref (yt | xi , y<t ) πθ (yt | xi , y<t )

(4)

This is a single-sample importance-weighted estimator of DKL (πref ∥πθ ) evaluated on the sampled token yt . We finally formulate the NSD loss for a single token yt as a composite objective: (t)

(t)

(t)

LNSD = LGU + α · LKL

(5)

where α is a hyperparameter. The full NSD algorithm is shown in Algorithm 1. For every component in the LNSD , we validate its necessity and effectiveness through ablation studies in Section 5. The overall objective is calculated by aggregating the token-level losses across the dataset:   |y| X (t) J (θ) = E(x,n)∼D,y∼πθ  LNSD  (6) t=1

Algorithm 1 Negative Self-Distillation (NSD) Training Require: Unlabeled dataset Draw = {xi }N i=1 , Initial model πθ0 , KL weight α Ensure: Optimized student model πθ 1: Initialize student πθ , and frozen teachers πref , πneg ← πθ0 2: for xi ∈ Draw do Sample reasoning trajectory y = (y1 , . . . , yT ) ∼ πθ (· | xi ) and the negative prompt ni ∼ 3: πθ (· | xi , y) 4: Initialize loss Ji ← 0 5: for t = 1, . . . , T do 6: pref ← πref (yt | xi , y<t ) 7: pneg ← πneg (yt | xi , ni , y<t ) 8: pθ ← πθ (yt | xi , y<t ) 9: Gt ← max(0, pneg − pref ) (t) Gt 10: LNSD ← 2−p + αpref log pprefθ θ (t)

11: Ji ← Ji + LNSD 12: end for 13: Update θ using gradient ∇θ Ji 14: end for 15: return πθ

5

Preprint.

3

E XPERIMENTAL S ETUP

Training setup. We use the MATH (Hendrycks et al., 2021) dataset as training dataset (for NSD, Intuitor and TTRL training, we discard the gold labels). We conduct training on the following models: Qwen3-1.7B, Qwen3-4B, and Qwen3-8B (Team, 2025). All models are trained for a total of 2 epochs, which is enough to plateau in all baselines. We set α = 0.01, top-k = 32, batch size = 32. For NSD, we set the max generation length to 4096. Evaluation. We evaluate the math reasoning ability of all models on the following benchmarks: AIME 2024, AIME 2025, AIME 2026, HMMT 2025 (Dekoninck et al., 2026), MATH-500, AMC 2023 and OlympiadBench (He et al., 2024). For OlympiadBench, we exclude the proof problems. By default, we set hyperparameters according to the recommended setting in Qwen3 report (Team, 2025): temperature = 0.6; top-p = 0.95; top-k = 20. The output length is set to 32K. Baselines. We compare with the following methods representing three different training paradigms: OPSD (Zhao et al., 2026a): A standard distillation framework that minimizes the full-vocabulary KL divergence between the student and a teacher conditioned on the gold solution. Intuitor (Zhao et al., 2026b): A representative RLIF (Reinforcement Learning from Internal Feedback) implementation, which is a variant of GRPO and utilizes average confidence (self-certainty) as the intrinsic reward. TTRL (Zuo et al., 2025): A variant of GRPO that utilizes the majority-voting consensus as pseudogold labels. While vanilla TTRL typically optimizes directly on the test set, we apply it to the training dataset to ensure a fair comparison with other baseline methods. A conceptual comparison of NSD with existing related methods, along with their implementation details and prompt templates, is provided in Appendix B and C.

4

E VALUATION R ESULTS

In this section, we first present the main experimental results across multiple mathematical reasoning benchmarks (§4.1). Then we empirically demonstrate NSD achieves better training efficiency and promotes reflection abilities compared to other baselines (§4.2, §4.4). Finally, we show that employing simpler negative conditioning strategies in NSD can also yield comparable effectiveness (§4.3). 4.1

M AIN R ESULTS

The main results are shown in Table 1. We highlight the following key observations: NSD achieves the overall best performance on the models with different sizes. As shown in Table 1, NSD consistently achieves the highest average improvements across all model scales, yielding ∆ Avg gains of +2.3%, +7.5%, and +6.0% on the three models respectively. While baselines like OPSD† and RL excel narrowly on AIME 2024, our 4B and 8B models achieve broader generalization across diverse math tasks, maintaining peak AIME accuracies of 35.8% and 39.6%. Notably, unlike other baselines where small improvements possibly partly stem from randomness, NSD guarantees stable performance gains, supported by a significantly low p-value. NSD is more promising on larger model sizes due to self-generated negative conditions. An observation from Table 1 is that NSD exhibits stronger performance gains on larger models compared to the smaller 1.7B variant. This scaling behavior is tied to our online negative condition generation mechanism: NSD uses on the model itself to generate solution-specific negative conditions. By optimizing against these higher-quality conditions, larger models receive a stronger contrastive training signal, which translates into substantial improvements on challenging reasoning tasks. Why does NSD outperform other baselines? Compared to OPSD, label-free training of NSD without the privileged information prevents bias (e.g., reinforcing reasoning shortcuts due to the gold solution) and reflection collapse caused by overconfidence (Kim et al., 2026b). We provide additional analysis in Section 4.2 that further confirms NSD better promotes reflection behaviors than other methods. Case studies in Appendix D.4 also illustrate how NSD-trained models abandon the wrong reasoning trajectory and switch to the right one. Compared to Intuitor and TTRL which use model 6

Preprint.

Table 1: Main evaluation results on mathematical reasoning benchmarks. We report the Avg@8 (%) performance under non-thinking mode (we report the performance under thinking mode in Appendix D.2). ∆ Avg is the average absolute improvement over the same-size baseline across all 7 benchmarks. We report the best checkpoint on the validation set within 2 training epochs. Bold marks the best result in each model-size group; underline marks the second best. † denotes methods that require ground-truth labels. The last two columns report the 95% CI and one-sided p-value (which measures the probability of observing an improvement at least as large as the really observed one if the improvement were due to chance) for ∆ Avg@8, respectively. OlympiadBench is evaluated on the 675 open-ended math problems (excluding proof problems) using the official judger with symbolic comparison. AIME 2024

AIME 2025

AIME 2026

HMMT 2025 Feb

AMC 2023

OlympiadBench

MATH500

∆ Avg

95% CI

p

1.7B Models Qwen3-1.7B OPSD† Intuitor TTRL NSD

9.6 15.0 13.8 11.3 14.2

10.0 14.2 8.3 11.3 17.9

9.6 8.8 8.3 9.6 10.0

7.1 5.8 6.7 8.3 7.1

44.1 44.1 43.4 41.6 45.9

37.1 37.2 35.4 36.8 38.7

62.5 62.5 60.6 63.4 62.6

— +1.1 −0.5 +0.3 +2.3

— [−0.3, +2.4] [−1.8, +0.8] [−1.0, +1.5] [+0.7, +4.0]

— 0.06 0.29 0.29 0.001

4B Models Qwen3-4B OPSD† Intuitor TTRL NSD

23.8 25.4 24.6 25.8 35.8

20.4 22.5 25.8 19.6 31.3

17.9 15.8 18.3 18.3 29.2

10.8 15.8 13.8 11.7 16.3

68.8 68.8 70.0 68.1 76.3

47.8 47.6 47.7 47.1 51.0

71.2 71.7 69.8 71.7 73.1

— +1.0 +1.3 +0.2 +7.5

— [−0.3, +2.4] [−0.5, +3.1] [−1.2, +1.7] [+5.4, +9.5]

— 0.10 0.05 0.33 < 10−4

8B Models Qwen3-8B OPSD† Intuitor TTRL NSD

28.8 30.0 34.6 29.2 39.6

19.2 21.3 20.4 18.3 26.3

18.3 17.1 18.3 17.1 25.0

11.7 12.1 13.8 10.8 17.9

67.2 66.9 70.9 69.1 75.6

48.9 48.2 49.5 49.1 50.6

73.1 73.5 72.7 73.0 74.1

— +0.3 +1.9 −0.1 +6.0

— [−1.3, +1.9] [+0.2, +3.4] [−1.4, +1.3] [+4.0, +7.9]

— 0.33 0.02 0.57 < 10−4

Method

confidence or majority voting to generate training signals, NSD removes the reliance on the model’s self-judgement ability, which leads to possible incorrect training signals. For example, weaker models hardly gain improvement from Intuitor (-0.5% on Qwen3-1.7B) because their high confidence does not necessarily equate to high accuracy. Furthermore, those confidence-based bootstrapping methods also degrade the reflection ability, as shown in Section 4.2. 4.2

NSD I NSPIRES R EFLECTION

NSD prevents over-confidence and preserves exploratory reflection. We evaluate model reflection capabilities by measuring the average frequency of reflection tokens (e.g., “Wait”) across AIME and HMMT benchmarks (Table 2). The detailed definition of reflection tokens is shown in Appendix E. We observe that OPSD and Intuitor severely suppress reflective behavior (dropping to 2.18 and 0.75 per response, respectively), as training on ground-truth or unverified positive rollouts encourages overly direct, non-verifying reasoning trajectories. Table 2: The average reflection token frequency per response on Qwen3-4B. Method

AIME 2024

AIME 2025

HMMT 2025

Average

Baseline OPSD Intuitor NSD

6.8 2.6 0.6 6.9

2.2 2.1 1.0 7.5

1.7 1.8 0.7 8.1

3.6 2.2 0.8 7.5

Conversely, NSD substantially enhances reflection frequency (yielding up to 7.5 per response). By penalizing flawed reasoning paths, NSD avoids over-confidence and enables the model to autonomously re-evaluate potential errors during complex inference, which is also demonstrated by our case study in Appendix D.4. 7

Preprint.

4.3

N EGATIVE C ONDITION VARIANTS S TUDY

In our main experiments, we default to an online self negative condition generation strategy (denoted as online strategy briefly). While intuitively well-motivated, this approach incurs computational overhead from online rollouts. To explore more efficient alternatives, we investigate the impact of simpler conditioning strategies, selecting the variants based on effective LLM negative conditioning paradigms identified in (Chatziveroglou et al., 2025). In this section, we discuss the following offline generation strategies while our primary evaluations in the previous sections are conducted using the online paradigm: Strategy 1: Solution-aware negative conditioning. Besides inputting the question, we let the student model rollout first, then prompt it to generate a negative condition based on the question and rollout. Strategy 2: Question-only negative conditioning. Input the training sample question to the frozen initial student and prompt it to generate a possible negative condition based on it. Strategy 3: Noise conditioning. Simply add irrelevant Wikipedia articles as noise (denoted as wiki-irr strategy; “irr” stands for irrelevant). 4B solution-aware (online, default)

4B solution-aware (offline)

4B question-only (offline)

4B wiki-irr (offline)

100

74.1 72.9 73.6 74.5

Score (%)

80 60 40

35.8 32.9 36.7 33.3

31.3

20 0

AIME 2024

22.5 26.2 27.9 AIME 2025

16.3 15.8 18.8 16.7

7.8 4.5 7.3 6.6

HMMT Feb 2025

MATH-500

∆ Avg

Figure 4: Evaluation results of NSD across different conditioning strategies. ∆ denotes the average absolute improvement over the base model across these 4 datasets. The dashed line represents the reference ∆ achieved by the default online strategy. We evaluate the three NSD conditioning strategies on the Qwen3-4B model. The experimental result of different conditioning strategies is shown in Figure 4. Notably, the question-only negative conditioning strategy achieves a 7.3% average improvement, comparable to 7.8% using our default online solution-aware approach. Furthermore, even the most lightweight offline strategy (wiki-irr) also performs competitively with our default approach, highlighting NSD’s broad scalability to diverse and efficient negative conditions. In contrast, the offline solution-aware strategy exhibits relatively lower performance, primarily driven by its reliance on outdated offline-generated solutions during conditioning. E FFICIENCY OF NSD

187s

200

NSD exhibits superior training efficiency compared to OPSD and RLIF. We focus our detailed latency analysis on the computational overhead of the rollout phase, which is the dominant source of training time discrepancy across different algorithms. In the rollout stage, the student model samples batch size × n rollouts, where n denotes the number of samples per prompt. While GRPO-based baselines (Intuitor and TTRL) demand n = 8, both NSD and OPSD require only n = 1. Subsequently, OPSD and NSD perform additional forward passes on the generated sequences: OPSD prefills each concatenated prompt-response pair to extract top-k logprobabilities, where k = 32 in NSD and 128 in OPSD; online NSD generates an online negative condition based on the student’s solution before running two forward passes to compute πref and πneg . 8

Time (s)

4.4

150 100 50 0

105s 68s

54s

NSD NSD (Online) (Wiki-Irr)

OPSD

Intuitor

Figure 5: Average wall-clock time per training step with 6 or 8 A100 GPUs. For NSD and OPSD, the student model occupies 4 GPUs and the teacher occupies 2 GPUs. For Intuitor, the generation stage is executed across all 8 GPUs. Notably, the wiki-irr strategy effectively reduces the latency compared to using the default online rollout in NSD.

Preprint.

The time consumption is shown in Figure 5. Overall, NSD achieves superior training efficiency through three primary factors: (1) Minimal Rollout Overhead: Unlike multi-sample GRPOstyle baselines, NSD requires only a single rollout per sample, reducing rollout time by ∼60%. Furthermore, static negative condition generation strategies (e.g., wiki-irr) save online rollout time entirely, lowering the overall latency from 68s to 54s. (2) Parallelized Forward Prefilling: Although computing πref and πneg involves two distinct prompts, both share the same model weights and can be prefilled concurrently in parallel. (3) Scalar-Only Loss Computation: NSD requires only three scalar token probabilities, avoiding full-vocabulary logit projections. The result shows that NSD trains faster overall than OPSD, proving that our parallelized prefilling costs substantially less than OPSD’s Top-k logit alignment. 4.5

M ORE E VALUATIONS AND A NALYSES

We conduct several supplementary evaluations provided in Appendix D. First, we evaluate our models under the Pass@8 metric in Appendix D.1. Second, we report the performance under thinking mode in Appendix D.2. NSD continues to outperform all baselines under these settings. Third, in Appendix D.3, we investigate an alternative objective formulation that treats the negative of the loss as an advantage signal for policy-gradient optimization, demonstrating that the NSD framework is scalable to policy-gradient-style training paradigms.

5

U NDERSTANDING NSD T RAINING O BJECTIVE

5.1

A DAPTIVE G ATING A NALYSIS

The adaptive gating function is designed to filter out trivial tokens while retaining essential ones. RLCSD (Pan et al., 2026) rigorously conceptualizes this by categorizing tokens into style and task tokens, treating the former as noise. A detailed definition is shown in Appendix E. To evaluate how effectively the NSD gating function and existing weighting methods filter out style tokens, we sample a subset of 100 training queries and analyze the logit distributions across the initial 4,096 tokens and calculate the style-task ratio (T denotes the set of task tokens and S denotes the set of style tokens): # " # 1 X 1 X wt / wt R= |S| |T | "

t∈S

(7)

t∈T

We compare NSD adaptive gating with initial ratio, entropy-based OPSD weighting (Wang et al., 2026b), and vanilla OPSD loss (Zhao et al., 2026a):  P NSD: wt = max 0, πneg (yt | x, a, y<t ) − πref (yt | x, y<t ) ; Entropy-OPSD: wt = − v πθ (v | P θ (v) x, y<t ) log πθ (v | x, y<t ); OPSD: wt = v πθ (v) log ππgold (v) . A lower R inherently signifies a better approach (Pan et al., 2026); it implies the method prioritizes task tokens, suppressing gradient generation on meaningless tokens — previous work (Pan et al., 2026; Zhao et al., 2026a) shows that in OPSD, the training signal might be dominated by style tokens, causing the student to imitate styles rather than learning reasoning ability. As shown in Table 3, the NSD gating mechanism alone filters style tokens more effectively than both entropy-based and OPSD-loss-based weighting methods. Table 3: Comparison of style-task ratio across different methods. A lower value indicates a better approach (Pan et al., 2026). Method style-task ratio

NSD gate (wiki)

NSD gate (solution-aware)

NSD gate (question-only)

EntropyOPSD

OPSD

2.6×

3.4×

3.5×

3.9×

5.4×

9

Preprint.

5.2

KL A BLATION

We investigate the necessity of the KL divergence constraint within the NSD framework. Figure 6 illustrates this via an ablation study comparing the standard NSD against a variant without the KL anchor (NSD-noKL). KL Penalty Mean

Removing the KL constraint leads to a midtraining collapse. As shown in Figure 6 (left), NSD-noKL drastically shifts the student’s distribution, inducing a cycle of learning and forgetting, evidenced by sharp oscillations in the curve. Furthermore, Figure 6 (right) demonstrates that the gate activation ratio in NSD-noKL initially increases but drops precipitously around the 120th step, coinciding precisely with the KL collapse. We also observe that the mean LGU value decreases significantly, indicating that the gating mechanism activates spuriously and loses its effectiveness—a direct result of the model drifting excessively from the reference without KL regularization. 5.3

1.2

Gate Active Ratio

w/o KL

0.5

w/o KL

w/ KL

w/ KL

0.8

0.4

0.4

0.3 0.2

0.0

0.1 −0.4 0

60

120

180

240

300

0.0

0

60

120

180

240

300

Figure 6: Training log of NSD w/ and w/o KL constraint on Qwen3-4B. Left: Forward KL between the reference model and student model per training step; Right: The average activated gating (Gt ) per training step.

D ISCUSSION ON G ATED U NLIKELIHOOD

A natural inherent consequence of the gating formulation is its sensitivity to minor probability fluctuations in highly predictable tokens (where πneg ≈ πref → 1, usually trivial tokens like punctuations). Occasionally, inherent variance may cause the negative teacher model to assign a marginally higher probability than the reference model, bypassing the filter (e.g., probabilities for space tokens often fluctuate slightly around 99%). Nevertheless, this artifact is controlled: the minuscule divergence yields a near-zero gate value Gt , ensuring that the overall gating on these high-confidence tokens remains negligible. To further avoid distancing from these tokens, we bound the pure unlikelihood penalty via Sigmoid function (Equation 3), so that the model enjoys implicit gradient attenuation. The Sigmoid penalty serves as a structural failsafe against gating imperfections. As shown in Figure 3, LGU value on high-probability tokens is significantly lower than the value of pure unlikelihood objectives. Gradient analysis is mathematically discussed in Appendix A.

6

R ELATED W ORK

On policy distillation. The original OPD (Agarwal et al., 2024; Lu & Lab, 2025; Song & Zheng, 2026) relies on external reward models. Self distillation (Zhao et al., 2026a; Hübotter et al., 2026; Shenfeld et al., 2026) removes external teachers by using ground-truth solutions as hints, but suffers from solution bias and overconfidence (Kim et al., 2026b; Harne et al., 2026; Wang et al., 2026a); Recent studies have increasingly optimized the distillation method across various dimensions, mainly including weak supervision (He et al., 2026; Li et al., 2026), credit assignment (Pan et al., 2026; Wang et al., 2026d; Xu et al., 2026c; Wang et al., 2026c), agentic scenarios (Wu et al., 2026; Lu et al., 2026), and other better learning objectives (Yang et al., 2026; Heo et al., 2026; Jiang et al., 2026; Shen et al., 2026; Kim et al., 2026a). Label-free reinforcement learning. Existing label-free training methods primarily rely on substituting rewards with self-generated ones (usually based on confidence or entropy) within RLVR frameworks (Zhao et al., 2026b; Li et al., 2025; Yuan et al., 2025; Prabhudesai et al., 2026; Huang et al., 2026b;a) or OPD frameworks (Gkountouras et al., 2026; Li et al., 2026), as well as generating gold labels by the model itself (Zhang et al., 2025a; Zuo et al., 2025). Training with negative signals. Unlikelihood objective (Welleck et al., 2020; Li et al., 2020) has been proposed to train earlier small language models. Several studies incorporate both positive and negative trajectories into distillation or RLVR frameworks (Xu et al., 2026b; Hamdan & Yuret, 2025; Yang et al., 2024). Notably, NSR (Zhu et al., 2026) explores RLVR training driven exclusively by 10

Preprint.

negative signals, demonstrating that it can preserve high-confidence priors while mitigating overfitting. Furthermore, in the context of self-distillation, recent works introduce negative signals to alleviate student overconfidence (Shen et al., 2026; Kim et al., 2026a).

7

C ONCLUSION

In this work, we propose Negative Self-Distillation (NSD), a label-free training framework comprising negative conditioning and gated unlikelihood training. Empirical results demonstrate that NSD consistently outperforms existing baselines across seven mathematical reasoning benchmarks under various model sizes. Comprehensive analyses and ablation studies show that our adaptive gating mechanism effectively isolates genuinely flawed tokens, while the sigmoid unlikelihood objective ensures smoother reasoning gradients. Furthermore, NSD accommodates diverse negative conditioning strategies, establishing it as a highly scalable framework. Its training efficiency is enhanced by bypassing full-vocabulary computations and leveraging parallelized forward passes for the negative teacher and reference model. Importantly, NSD inherently preserves and stimulates the model’s capacity for self-reflection, highlighting its potential as a promising post-training method for enhancing the reasoning capabilities of LLMs.

L IMITATIONS NSD relies on the student model’s inherent capacity to generate negative conditions. Consequently, this approach may be less effective for extremely small or weak models that struggle to produce meaningful negative contrasts for optimization. However, given the rapid capability scaling of modern foundational models, this capacity bottleneck is expected to diminish naturally in future architectures or in stronger models. Under the online strategy, NSD requires negative-condition generation and two forward passes through the two same frozen models. Nevertheless, we explored alternative conditioning strategies, including efficient generation-free methods like the wiki-irr strategy, which can mitigate the rollout costs while maintaining competitive performance. Moreover, executing the two forward passes in parallel at each training step effectively minimizes overall wall-clock latency.

ACKNOWLEDGMENTS This research is partially funded by the NVIDIA Academic Grant and Amazon Research Award. We thank Xinyu Wang and Yu Gu for their valuable feedback and suggestions, particularly for the experimental design.

R EFERENCES Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=3zKtaqxLhW. Giannis Chatziveroglou, Richard Yun, and Maura Kelleher. Exploring LLM reasoning through controlled prompt variations, 2025. URL https://arxiv.org/abs/2504.02111. Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, and Martin Vechev. Beyond benchmarks: MathArena as an evaluation platform for mathematics with LLMs, 2026. URL https://arxiv.org/abs/2605.00674. Mukesh Ghimire, Aosong Feng, Liwen You, Youzhi Luo, Fang Liu, and Xuan Zhu. PRISM: A unified framework for post-training LLMs without verifiable rewards, 2026. URL https: //arxiv.org/abs/2601.04700. John Gkountouras, Josip Jukić, and Ivan Titov. Consensus as privileged context for label-free self-distillation, 2026. URL https://arxiv.org/abs/2607.13643. 11

Preprint.

Shadi Hamdan and Deniz Yuret. How much do LLMs learn from negative examples?, 2025. URL https://arxiv.org/abs/2503.14391. Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, and Akshay Nambi. Privileged, but biased: How pi-conditioned teachers break self-distillation, 2026. URL https://arxiv. org/abs/2608.04794. Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems, 2024. URL https://arxiv.org/abs/2402.14008. Yinghui He, Simran Kaur, Adithya Bhaskar, Yongjin Yang, Jiarui Liu, Narutatsu Ri, Liam Fowl, Abhishek Panigrahi, Danqi Chen, and Sanjeev Arora. Self-Distillation Zero: Self-revision turns binary rewards into dense supervision, 2026. URL https://arxiv.org/abs/2604.12002. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset, 2021. URL https://arxiv.org/abs/2103.03874. Byeongho Heo, Jaehui Hwang, Sangdoo Yun, and Dongyoon Han. On-policy delta distillation, 2026. URL https://arxiv.org/abs/2607.15161. Chengsong Huang, Haolin Liu, Tong Zheng, Runpeng Dai, Langlin Huang, Jinyuan Li, Zongxia Li, Zhepei Wei, Yu Meng, and Jiaxin Huang. G-Zero: Self-play for open-ended generation from zero data. arXiv preprint arXiv:2605.09959, 2026a. Chengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang, Zongxia Li, Ruosen Li, Jiaxin Huang, Haitao Mi, and Dong Yu. R-Zero: Self-evolving reasoning LLM from zero data. In The Fourteenth International Conference on Learning Representations, 2026b. URL https: //openreview.net/forum?id=96apU6YzSO. Jonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, and Andreas Krause. Reinforcement learning via self-distillation, 2026. URL https://arxiv.org/abs/2601. 20802. Li Jiang, Haoran Xu, Yichuan Ding, and Amy Zhang. Trajectory-refined distillation, 2026. URL https://arxiv.org/abs/2606.08432. Jeonghye Kim, Jiwon Jeon, Dongsheng Li, and Yuqing Yang. Rebellious student: Reversing teacher signals for reasoning exploration with self-distilled RLVR, 2026a. URL https://arxiv.org/ abs/2605.10781. Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, and Yuqing Yang. Why does self-distillation (sometimes) degrade the reasoning capability of LLMs?, 2026b. URL https://arxiv.org/abs/2603.24472. Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luca Soldaini, Noah A. Smith, Yizhong Wang, Pradeep Dasigi, and Hannaneh Hajishirzi. Tulu 3: Pushing frontiers in open language model post-training, 2025. URL https://arxiv.org/ abs/2411.15124. Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, and Jason Weston. Don’t say that! making inconsistent dialogue unlikely with unlikelihood training. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4715–4728, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.428. URL https://aclanthology.org/2020.acl-main.428/. 12

Preprint.

Pengyi Li, Matvey Skripkin, Alexander Zubrey, Andrey Kuznetsov, and Ivan Oseledets. Confidence is all you need: Few-shot RL fine-tuning of language models, 2025. URL https://arxiv. org/abs/2506.06395. Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, and Nuno Vasconcelos. On-policy selfdistillation without any supervision, 2026. URL https://arxiv.org/abs/2608.06296. Baohao Liao, Hanze Dong, Xinxing Xu, Christof Monz, and Jiang Bian. Self-hinting language models enhance reinforcement learning, 2026. URL https://arxiv.org/abs/2602.03143. Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy-distillation. Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang, Jinyang Wu, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, and Yongliang Shen. Self-distilled agentic reinforcement learning, 2026. URL https://arxiv.org/abs/2605.15155. Leyi Pan, Shuchang Tao, Yunpeng Zhai, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, and Lijie Wen. RLCSD: Reinforcement learning with contrastive on-policy self-distillation, 2026. URL https://arxiv.org/abs/2606.11709. Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak. Maximizing confidence alone improves reasoning, 2026. URL https://openreview.net/ forum?id=Qhg479eBmo. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models, 2024. URL https://arxiv.org/abs/ 2402.03300. Guobin Shen, Xiang Cheng, Chenxiao Zhao, Lei Huang, Jindong Li, Dongcheng Zhao, and Xing Yu. Anti-self-distillation for reasoning RL via pointwise mutual information, 2026. URL https: //arxiv.org/abs/2605.11609. Idan Shenfeld, Mehul Damani, Jonas Hübotter, and Pulkit Agrawal. Self-distillation enables continual learning, 2026. URL https://arxiv.org/abs/2601.19897. Mingyang Song and Mao Zheng. A survey of on-policy distillation for large language models. arXiv preprint arXiv:2604.00626, 2026. Qwen Team. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. Meng Wang, Haohan Zhao, Wenzhuo Liu, Lu Yang, Geng Liu, Haiyang Guo, Guo-Sen Xie, Gaofeng Meng, Hongbin Liu, and Fei Zhu. Denser ̸= better: Limits of on-policy self-distillation for continual post-training, 2026a. URL https://arxiv.org/abs/2607.01763. Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, Kai Dang, Xiong-Hui Chen, Jianxin Yang, Zhenru Zhang, Yuqiong Liu, An Yang, Andrew Zhao, Yang Yue, Shiji Song, Bowen Yu, Gao Huang, and Junyang Lin. Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026b. URL https://openreview.net/forum? id=yfcpdY4gMP. Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, and Hongxia Yang. Not all disagreement is learnable: Token teachability in on-policy distillation, 2026c. URL https://arxiv.org/abs/2605.26844. Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, and Yujiu Yang. AgentOPSD: Recursive self-distillation for agentic reinforcement learning, 2026d. URL https://arxiv.org/abs/ 2608.05987. 13

Preprint.

Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. Neural text generation with unlikelihood training. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SJeYe0NtvH. Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, and Jianhua Tao. SEED: Self-evolving on-policy distillation for agentic reinforcement learning, 2026. URL https://arxiv.org/abs/2607.14777. Haobo Xu, Sirui Chen, Ruizhong Qiu, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, and Hanghang Tong. Prune as you generate: Online rollout pruning for faster and better RLVR. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens (eds.), Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13876–13893, San Diego, California, United States, July 2026a. Association for Computational Linguistics. ISBN 979-8-89176-390-6. doi: 10.18653/v1/2026.acl-long.632. URL https://aclanthology.org/2026.acl-long.632/. Shuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu, Wei Chu, and Yuan Qi. Harnessing negative signals: Reinforcement distillation from teacher data for LLM reasoning. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens (eds.), Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1618–1639, San Diego, California, United States, July 2026b. Association for Computational Linguistics. ISBN 979-8-89176-390-6. doi: 10.18653/v1/2026.acl-long.74. URL https://aclanthology. org/2026.acl-long.74/. Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, and Alborz Geramifard. TIP: Token importance in on-policy distillation, 2026c. URL https://arxiv.org/abs/2604.14084. Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian. RLCD: Reinforcement learning from contrastive distillation for LM alignment. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id= v3XXtxWKi6. Shenzhi Yang, Guangcheng Zhu, Bowen Song, Haobo Wang, Mingxuan Xia, Xing Zheng, Yingfan Ma, Zhongqi Chen, Weiqiang Wang, Junbo Zhao, and Gang Chen. OPRD: On-policy representation distillation, 2026. URL https://arxiv.org/abs/2606.06021. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, WeiYing Ma, Ya-Qin Zhang, Lin Yan, Mu Qiao, Yonghui Wu, and Mingxuan Wang. DAPO: An open-source LLM reinforcement learning system at scale, 2025. URL https://arxiv.org/ abs/2503.14476. Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. Self-rewarding language models, 2025. URL https://arxiv.org/abs/ 2401.10020. Kongcheng Zhang, QI YAO, Shunyu Liu, Yingjie Wang, Baisheng Lai, Jieping Ye, Mingli Song, and Dacheng Tao. Consistent paths lead to truth: Self-rewarding reinforcement learning for LLM reasoning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025a. URL https://openreview.net/forum?id=ckW70ls93V. Xingjian Zhang, Siwei Wen, Wenjun Wu, and Lei Huang. EDGE-GRPO: Entropy-driven GRPO with guided error correction for advantage diversity, 2025b. URL https://arxiv.org/abs/ 2507.21848. Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, and Aditya Grover. Self-Distilled Reasoner: On-policy self-distillation for large language models, 2026a. URL https://arxiv.org/abs/2601.18734. 14

Preprint.

Xuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song. Learning to reason without external rewards. In The Fourteenth International Conference on Learning Representations, 2026b. URL https://openreview.net/forum?id=OU9nFEYR2M. Xinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen, Danqi Chen, and Yu Meng. The surprising effectiveness of negative reinforcement in LLM reasoning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. URL https://openreview.net/ forum?id=ftVlLG9cks. Yuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Yuchen Zhang, Xinwei Long, Ermo Hua, Biqing Qi, Youbang Sun, Zhiyuan Ma, Lifan Yuan, Ning Ding, and Bowen Zhou. TTRL: Test-time reinforcement learning, 2025. URL https://arxiv.org/ abs/2504.16084.

15

Preprint.

A

A NALYSIS OF C ANDIDATE G ATED U NLIKELIHOOD G RADIENTS IN NSD AND OPSD O BJECTIVES

The NSD loss function is defined as L = α · DKL (πref ∥ πθ ) + LGU . Specifically, we focus on isolating and analyzing the LGU term, which penalizes the student model on tokens vulnerable to negative conditions. Let πc = πθ (c) denote the student’s predicted probability for the target sampled token, and G = max(0, πneg − πref ) serve as the adaptive gate. To demonstrate the necessity of our LGU item with Sigmoid design, we compare two distinct candidates for this penalty: Standard Unlikelihood: GUstd = G · [− log(1 − πc )] (8) 1 Ours: GUsig = G · σ(− log(1 − πc )) = G · (9) 2 − πc The parameter update magnitude is driven by the gradient of the loss with respect to the pre-softmax ∂πc logit, ∂GU ∂zc . Given the logit-probability Jacobian ∂zc = πc (1 − πc ), we evaluate the optimization behavior of both formulations below. 0.15

1.0 → 1.0 as πc → 1

10 0 max = 0.1250 at πc = 0.67

10 −2

0.6 0.4

0.10

10 −4

zc

GU sig

zc

GU std

0.8

GU sig : π(2c (1−−πcπ)c2)

0.05

10 −6

0.2 0.0 0.00

0.25

0.50 πc

0.75

| L/ zc | (log scale)

GU std : πc

1.00

0.00 0.00

0.25

0.50 πc

0.75

1.00

10 −8

OPSD: |πθ − πg | GU sig : Gπ(2θ (1−−πθπ)θ2) GU std : Gπθ Upper bound (OPSD) Upper bound (GU sig ) Upper bound (GU std )

0.2

0.4

0.6

πθ

0.8

1.0

Figure 7: (Left and Mid) Comparison of the variations of two types of gradients by probability. (Right) The real gradient distribution in 100 training samples. OPSD tends to assign larger gradients to high-probability tokens, making the model more prone to drastic updates. In contrast, compared to the vanilla unlikelihood loss, our GU objective further suppresses the gradients on high-probability tokens. A.1

G RADIENT H AZARD IN S TANDARD U NLIKELIHOOD

Applying the chain rule to the standard unbounded logarithmic penalty (Equation 8), the gradient with respect to the logit is: ∂GUstd 1 =G· · πc (1 − πc ) = G · πc (10) ∂zc 1 − πc Mathematically, the gradient in standard unlikelihood scales strictly linearly with the student’s confidence πc (as shown in the Figure 7). This creates an optimization hazard: In causal language modeling, tokens with extreme confidence (πc > 0.9) are probably trivial structural tokens—such as fixed collocations, prepositions, and punctuation. Under this formulation, whenever the gate G is triggered when πc → 1, the optimizer delivers its almost maximum update magnitude to these hyper-confident function words. This aggressively penalizes the model’s fundamental linguistic priors, leading to a degradation in generation fluency, especially when the high-probability token ratio is high per rollout. A.2

I MPLICIT G RADIENT ATTENUATION IN S IGMOID -S QUASHED GU

To construct a noise-resilient supervision signal, our method utilizes the Sigmoid-squashed penalty (Equation 9). Deriving the logit gradient for this formulation yields: ∂GUsig 1 =G· · πc (1 − πc ) ∂zc (2 − πc )2 πc (1 − πc ) =G· (11) (2 − πc )2 16

Preprint.

This formulation introduces an elegant, parameter-free implicit gradient attenuation mechanism. The presence of the (1 − πc ) term in the numerator fundamentally alters the gradient landscape. As the student model’s probability approaches 1, the gradient magnitude decays toward zero: lim

πc →1

∂GUsig =0 ∂zc

(12)

Since πc > 0.9 predominantly corresponds to uninformative syntactic tokens, the Sigmoid function inherently protects the model’s structural fluency by silencing huge gradient on these tokens. Instead, as shown in Figure 7 it naturally concentrates the highest gradient magnitude on mid-confidence tokens (πc ≈ 0.6), which are more likely to be the ambiguous, reasoning-critical tokens where the student model requires the strongest corrective supervision. Consequently, our squashed formulation guarantees that dense supervision remains targeted and stable. A.3

OPSD G RADIENT A NALYSIS

OPSD objective can be described by:   (t) LOPSD = DKL πθ (· | xi , si , y<t ) ∥ πθ (· | xi , y<t )

(13)

si denotes the gold solution in i-th training sample. The gradient visualization is shown in Figure 7. This figure shows that, compared with the NSD objective, whose gradient generally decreases as the token probability increases, the OPSD objective exhibits an increasing trend. This indicates that OPSD encourages the model to learn more from high-probability tokens, which may lead to certain forms of reward hacking, such as overlearning style tokens. In contrast, the candidate objectives in NSD exhibit relatively stable gradient patterns, while the sigmoid-based GU can more effectively suppress gradients on high-probability tokens.

B

C OMPARISON OF NSD WITH OTHER M ETHODS

Table 4: Comparison of NSD with other methods. We conceptually compare them across the following dimensions: Sampling denotes whether the training relies on trajectories generated by the model itself; Source of reward signal indicates the core component driving the training loss function; Teacher specifies whether the approach depends on an external teacher model; Gold label refers to whether ground-truth answers are required; and Monitor signal quality represents whether the method actively filters training signals (e.g., unconsciously or intentionally) rather than indiscriminately optimizing over all tokens. Note that our NSD is a label-free approach, which is not directly comparable to baselines that rely on additional or external supervision. Consequently, our main experiments mostly focus on comparable methods, with OPSD as a representative labeldependent method for reference. Method SFT/Off-Policy Distillation RLVR (GRPO) OPD OPSD/SDPO SD-Zero RLIF TTRL/U-OPSD NSD

Sampling

Source of reward signal

/ off-policy

external teacher

, on-policy , on-policy , on-policy , on-policy , on-policy

external reward gold label trained reviser internal metric majority-voting

, on-policy

, on-policy

Teacher

Gold label

Monitor signal quality

/ external

, no

/ no

/ external , self , self , no , no

, no / needed / needed , no , no

, no

gold label

negative condition

17

, self

/ needed

, no

, noise gradients are , counteracted / no / no / no / no / no , noise is filtered by gating

Preprint.

C

E XPERIMENT D ETAILS

C.1

H YPERPARAMETERS

Table 5 lists the training hyperparameters for all methods. All experiments are conducted on a single node equipped with 8 NVIDIA A100 (80GB) GPUs. Unless otherwise specified, we adopt the default hyperparameters from the respective official implementations, with the following controlled adjustments for fair comparison: For OPSD, we evaluate configurations both with and without LoRA and report the best-performing variant (where LoRA achieves superior results on the 1.7B and 4B models). For Intuitor, we standardize the training batch size to 128, deviating from their scale-dependent defaults (64 for smaller models and 128 for larger models). For TTRL, as majority voting relies on complete final solutions, we extend the maximum generation length to 8192 tokens to prevent output truncation. Table 5: Training hyperparameters for all methods. “—” means not applicable. Hyperparameter

NSD (Online, Solution-aware)

OPSD

Intuitor

TTRL

GPUs Train batch size PPO mini-batch size Max prompt length Max response length Actor learning rate LR warmup ratio Rollout per sample n Top-k logits KL coefficient Total epochs

4 for actor + 2 for teacher 32 32 512 4096 1 × 10−6 0.1 1 32 0.01 2

4 for actor + 2 for teacher 32 32 512 4096 5 × 10−6 0.1 1 -1 — 2

8 128 128 512 3072 3 × 10−6 0.1 8 — 0.005 2

8 8 1 512 8192 5 × 10−7 0.03 8 — 0.00 2

OPSD LoRA target modules: all-linear, with LoRA rank = 64 and alpha = 128.

C.2

T EMPLATES

We use the Qwen3 instruct chat template throughout. All training are conducted in non-thinking mode: the chat template is invoked with enable_thinking=False, which causes the model to emit an empty <think> block and proceed directly to the answer. This applies uniformly to the student rollout, the teacher log-probability computation, and all downstream evaluations. The template for a single-turn exchange takes the following form: Qwen3 Chat Template (non-thinking, enable_thinking=False) <|im_start|>user {user message} <|im_end|> <|im_start|>assistant <think> </think> {model response}

The empty <think>...</think> block is prepended automatically by the template when enable_thinking=False and add_generation_prompt=True. The model then generates its response after the second blank line. C.3

P ROMPTS

The student always receives the plain problem prompt below. During NSD training the teacher receives either the same prompt (reference pass) or a negative prompt (negative pass), depending on the variant. All prompts are wrapped in the chat template described in Appendix C.2. 18

Preprint.

Student / reference teacher prompt (all methods). Student Prompt Problem:

{problem}

Let’s think step by step and output the final answer within \boxed{}.

NSD negative condition prompt generator (Question-only). The following meta-prompt is sent to a helper LLM to produce the per-sample negative condition prompt ni used in the question-only offline variant. The generated prompt replaces the system context seen by the teacher model. Meta-Prompt:

Question-only Negative Condition Generation

You are an expert Math Educator and AI Prompt Engineer. Your task is to analyze the following math problem and generate a “Generalized Attack Prompt” that will force an LLM to make a highly plausible, human-like cognitive error. Anatomy of a Universal Attack Prompt: 1. Persona: Must start exactly with “You are a student who. . . ”. Describe a specific bad habit relevant to this problem. 2. Trigger: Abstract the problem’s mathematical class. Never use specific numbers or variables from the current problem. 3. Flawed Execution: Instruct a naive heuristic or impulsive shortcut that would give a wrong answer. 4. Fatal Omission: Explicitly forbid the critical verification step. Now, perform this task for the following problem: Problem: {problem} Output only the “Generalized Attack Prompt”. Start your response with “You are a student who. . . ”. Keep it concise (2–3 sentences).

NSD negative condition prompt generator (Solution-aware). When the model’s own rollout is available, the meta-prompt is augmented with the student’s solution to produce a more targeted negative condition. Meta-Prompt:

Solution-aware Negative Condition Generation

You are an expert Math Educator and AI Prompt Engineer. Your task is to analyze the following math problem and a student’s existing solution, then generate a “Targeted Attack Prompt” that exploits the exact reasoning steps the student used to cause a highly plausible cognitive error. The student’s solution reveals how they solved this problem—use that to craft an attack targeting their specific reasoning steps. Anatomy of a Targeted Attack Prompt: 1. Persona: Must start exactly with “You are a student who. . . ”. Describe a specific bad habit that would corrupt the exact step where this student’s reasoning is most fragile. 2. Trigger: Reference the type of reasoning the student used (not specific numbers or variables from this problem). 3. Flawed Execution: Instruct a shortcut that mirrors the student’s approach but introduces a subtle error. 4. Fatal Omission: Forbid the specific verification the student performed correctly. Problem: {problem} Student’s Existing Solution: {solution} Output only the “Targeted Attack Prompt”. Start your response with “You are a student who. . . ”. Keep it concise (2–3 sentences).

NSD teacher prompt (wiki-irr variant). In the wiki-irr variant no meta-prompt generator is used. Instead, each training sample is paired with a randomly sampled Wikipedia passage that is 19

Preprint.

concatenated as spurious “context”. The teacher sees the following prompt while the student still receives the plain student prompt above. Teacher Prompt: Problem:

Wiki Irrelevant Negative Condition

{problem}

Below is some context you may find useful to answering the question above: {wikipedia_passage} Let’s think step by step and output the final answer within \boxed{}.

NSD teacher prompt. For the question-only and solution-aware offline variants, the teacher receives the following prompt, where {negative_condition} is the output of the meta-prompt generator above. Teacher Prompt: Problem:

Question-only / Solution-aware Negative Condition

{problem}

{negative_condition} Now solve the problem following this instruction: Let’s think step by step and output the final answer within \boxed{}.

D

A DDITIONAL E XPERIMENTAL R ESULTS

D.1

PASS @8 P ERFORMANCE

We report the performance of pass@8 in Table 6. Table 6: Main evaluation results reported as pass@8 (%): at least one of 8 sampled solutions is correct. Same evaluation setting as Table 1. ∆ Avg is the average absolute improvement over the same-size baseline across all 7 benchmarks. Bold marks the best result in each model-size group; underline marks the second best. † denotes methods that require ground-truth labels. AIME 2024

AIME 2025

AIME 2026

HMMT 2025 Feb

AMC 2023

OlympiadBench

MATH500

∆ Avg

1.7B Models Qwen3-1.7B OPSD† Intuitor TTRL NSD

16.7 40.0 30.0 30.0 33.3

23.3 23.3 23.3 26.7 36.7

13.3 13.3 16.7 23.3 23.3

16.7 16.7 13.3 13.3 16.7

70.0 72.5 75.0 77.5 77.5

57.8 59.6 57.0 58.7 60.9

77.6 77.6 76.2 77.0 77.6

— +3.9 +2.3 +4.4 +7.2

4B Models Qwen3-4B OPSD† Intuitor TTRL NSD

50.0 40.0 50.0 63.3 60.0

40.0 46.7 53.3 40.0 63.3

40.0 36.7 36.7 46.7 53.3

20.0 30.0 26.7 23.3 30.0

95.0 92.5 95.0 90.0 95.0

67.0 66.4 66.1 64.3 68.3

81.6 81.6 80.0 81.6 82.0

— +0.0 +2.0 +2.2 +8.3

8B Models Qwen3-8B OPSD† Intuitor TTRL NSD

56.7 46.7 60.0 56.7 70.0

30.0 43.3 40.0 33.3 46.7

43.3 40.0 36.7 36.7 60.0

23.3 23.3 26.7 20.0 40.0

92.5 87.5 95.0 92.5 95.0

66.8 67.6 69.0 66.8 69.0

82.0 81.8 81.8 82.0 81.8

— −0.6 +2.1 −1.0 +9.7

Method

20

Preprint.

D.2

P ERFORMANCE ON T HINKING M ODE

We evaluate the performance of all methods under the thinking mode on the Qwen3-4B model, reporting the results of the best-performing checkpoints evaluated under the thinking mode. As shown in Table 7, NSD also outperforms all other baselines overall. Notably, both OPSD and NSD achieve more improvements on challenging datasets such as AIME and Olympiad Bench, aligning with the observations from the non-thinking setting. On datasets with limited headroom for improvement (e.g., AMC), all methods perform comparably to the base model. Furthermore, we observe that the Intuitor-trained model tends to over-think, causing many responses to exceed the maximum generation length limit (even after extending it to 38k tokens), which leads to a severe degradation in accuracy. This phenomenon is also discussed in the previous works (Zhao et al., 2026b; Ghimire et al., 2026). Table 7: Thinking mode evaluation results on 4B models reported as avg@8 (%). Same evaluation setting as Table 1.

D.3

Model

AIME 2024

AIME 2025

AIME 2026

HMMT 2025

AMC 2023

MATH500

Olympiad Bench

∆ Avg

Qwen3-4B OPSD Intuitor TTRL NSD

75.8 76.2 52.9 72.5 77.3

69.1 69.8 45.8 64.3 73.3

67.5 67.2 51.2 65.1 67.7

46.0 46.2 39.6 46.0 48.4

97.2 96.6 90.9 96.6 97.8

79.8 80.0 78.1 79.2 79.9

45.9 46.7 43.9 44.4 57.9

+0.2 -11.3 -1.9 +3.0

A LTERNATIVE O BJECTIVE : P OLICY G RADIENT O PTIMIZATION (t)

While the NSD loss LNSD can be directly backpropagated as a supervised objective, we find it also fits a sampled-token advantage policy-gradient framework, following the spirit of Zhao et al. (2026a) and Lu & Lab (2025). For each token yt in a student rollout y ∼ πθ (· | xi ), we define a token-level advantage as the NSD loss: (t) At = −LNSD (14) Intuitively, a token with high NSD loss receives a strongly negative advantage, signaling the policy to reduce its probability. Conversely, tokens with low NSD loss receive near-zero or positive advantage, leaving their probabilities unchanged. We treat At as a constant with respect to θ and optimize the student via the standard policy gradient surrogate objective: " # X At log πθ (yt | xi , y<t ) (15) JPG (θ) = E(x,a)∼D,y∼πθ t

We evaluate the NSD based on the alternative objective under the same setting as our main experiment. The result is shown in Table 8. Compared to models optimized with the J objective (Eq. 6), NSD trained under JPG (Eq. 15) achieves superior performance on the 1.7B model (+5.0% on average). However, on the 4B and 8B models, the J -objective NSD yields better overall results. Notably, the wiki-irr strategy consistently performs best under the policy gradient setting, while the online solution-aware strategy emerges as the second best. This discrepancy arises because the gradients are truncated by the advantage function, decreasing the capture of richer gradient signals. In contrast, wiki-irr utilizes noise to introduce more generalized interference (causing an overall degradation of the model’s reasoning capabilities in long contexts), which ultimately makes it a more effective strategy in this regime. D.4

C ASE S TUDY

A case study is shown in Table 9. These results indicate that NSD-trained models more readily explore novel and correct solutions that are entirely absent from the outputs of both the base and OPSD-trained models. Furthermore, by prompting an external LLM (Sonnet) to analyze the reasoning traces of each response, we observe that the NSD-trained model engages in several reflection steps, successfully circumventing erroneous trajectories that commonly trap the base model. 21

Preprint.

Table 8: Evaluation results of NSD with different negative conditioning strategies on mathematical reasoning benchmarks. The models are trained based on the NSD policy gradient objective in Eq. 15. AIME 2024

AIME 2025

HMMT Feb 2025

MATH-500

∆ Avg

1.7B Models NSD (Solution-aware, Offline) NSD (Solution-aware, Online) NSD (Wiki-irr, Offline)

15.8 15.8 20.0

14.2 12.5 15.0

5.4 7.5 9.6

63.2 63.3 64.4

+2.4 +2.5 +5.0

4B Models NSD (Question-only, Offline) NSD (Solution-aware, Offline) NSD (Solution-aware, Online) NSD (Wiki-irr, Offline)

30.4 33.8 31.3 33.8

24.6 22.5 25.0 25.4

16.2 14.6 16.3 17.1

72.7 72.1 74.0 73.5

+4.4 +4.2 +5.1 +5.9

8B Models NSD (Solution-aware, Offline) NSD (Solution-aware, Online) NSD (Wiki-irr, Offline)

30.8 35.0 34.2

23.3 21.7 28.8

12.5 13.8 19.2

73.6 73.6 73.9

+1.9 +2.8 +5.8

Method / Variant

Table 9: Case study on AIME 2025 II #12 comparing Qwen3-4B baseline, OPSD, and NSD. The table shows a summary from Sonnet. Numbers denote correct samples out of 8 independent draws. ✓ correct; ✗ incorrect. Problem

Baseline

OPSD

NSD (Ours)

AIME 2025 II #12 — Geometry. Let A1 A2 · · · A11 be a non-convex simple 11-gon satisfying: √(1) [Ai A1 Ai+1 ] = 1 for 2 ≤ i ≤ 10; m n−p (2) cos(∠Ai A1 Ai+1 ) = 12 (n squarefree, no prime divides 13 for 2 ≤ i ≤ 10; (3) perimeter = 20. Express A1 A2 + A1 A11 = q all of m, p, q); find m + n + p + q. Pass@8 Key reasoning

0/8

✗

0/8 12 13

From cos θ = derives sin θ = 5 , hence |A1 Ai | · |A1 Ai+1 | = 26 . 13 5 The product constraint gives an alternating sequence a2 = x, a3 = 26 , a4 = x, . . . Attempts to 5x use the perimeter but conflates the sum of radii from A1 with the polygon perimeter, obtaining 5x + 26 = 20 which has no clean closed x form. After extensive numerical trials, guesses the symmetric solution 13 x = √ (i.e. a2 = a10 ), giving 5 √ 13 A1 A2 + A1 A11 = √ +2 5 = 5 √

23 5 , so m = 23, n = 5, p = 5 0, q = 5. Final: 33 ✗

Reflection

No self-correction. After the perimeter approach yields no clean form, the model commits to a guess (x = 13 √ ) without checking whether it sat5 isfies the original constraints, and submits the result directly.

✗

Same product relation and alternating sequence. Applies Law of Cosines: since xi xi+1 = 26 , each 5 inner-polygon side satisfies d2 = x2i + x2i+1 − 48 , so all 9 inner 5 sides are equal. Correctly writes the 26 perimeter equation a+9d+ 5a = 20 26 and sets S = a + 5a . Instead of solving for S, minimises Sq via AM–  26 GM: min a + 5a = 2 26 = 5 √

2 130 , and incorrectly treats this 5 minimum as the answer, √concluding A1 A2 + A1 A11 = 2 5130 , so m = 2, n = 130, p = 0, q = 5. Final: 137 ✗

No self-correction. The model sets up the correct equation structure but replaces the constraint with its relaxation: once the AM–GM bound is computed it is treated as the solution, with no attempt to verify that the minimum is actually attained.

22

4/8

✓

Same product relation and alternating sequence. Applies Law of Cosines; all 9 inner sides equal d. Writes the perimeter equation a + 26 = 20. Then attempts a 9d + 5a symmetric-guess approach: tests p x = 26/5, x = 2, x = 13/5, √ x = 13/ 5 in turn, each time verifying numerically that the two expressions for d2 do not agree. Sets 26 S = a + 5a , so the perimeter equation gives d = 20−S . Rewrites 9 2 d via Law of Cosines: d2 = 2 676 48 26 2 a + 25a2 − 5 = a + 5a − 2 52 48 − = S − 20. Substitut5 5  2 ing d = 20−S yields 20−S = 9 9 S 2 − 20, which expands to 4S 2 + 2S − 101√= 0. Quadratic formula: √ S = −2± 8 1620 = 9 45−1 (positive root), so m = 9, n = 5, p = 1, q = 4. Final: 19 ✓ After the guessing strategy p fails on multiple candidates (x = 26/5, 13 13 2, 5 , √5 ), the model explicitly abandons the approach (“Hmm. Maybe my approach is not working. Alternative idea:”) and reframes the problem around the aggregate 26 variable S = a + 5a , turning an intractable system into a single 2 quadratic 4S + 2S − 101 = 0.

Preprint.

E

D EFINITION OF TASK , S TYLE , AND R EFLECTION T OKENS

Considering the similar task setting with RLCSD (Pan et al., 2026), we follow theie definition of both task and style tokens: (1) empty or whitespace-only → style;

(2) matches any of the math regexes (a digit \d; an arithmetic operator in +− = ∗/ <> ×÷ ≤≱=; a LaTeX command \[A-Za-z]+; a double backslash; or one of $ ^ _) → task; (3) the normalized form is in the math wordlist {mod, prime, factor, gcd, lcm, log, ln, sin, cos, tan, exp, integral, sqrt, boxed, frac, sum, prod, pi, alpha, beta, gamma, theta, delta, lambda, mu, sigma, infty, leq, geq, neq, cdot, times, div} → task; (4) pure punctuation or a literal newline token (\n, \\n) → style;

(5) the normalized form is in the discourse wordlist (connectives therefore, so, thus, hence, then, because, since; hedges wait, maybe, perhaps, seems, okay, ok, well, now, first, next, finally, actually, alternatively, however; scaffolding step, answer, let, lets; closedclass function words is, are, us, we, the, a, an, of, to, for, in, on, by, at, as, and, or, but, if, yes, no, this, that, these, those, it, its, be, been, being, have, has, had, do, does, did, will, would, should, could, can, may) → style;

(6) otherwise → neutral.

The following tokens are considered as reflection tokens: wait, actually, hmm, let me reconsider, let me rethink, i made an error, i made a mistake, that’s wrong, that is wrong, incorrect, reconsider, rethink, re-examine, let me check, let me verify, double check, double-check, going back, revisit, on second thought

23

Record · ID 673516 · SHA-256 56dfb2f60efa90ba
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.