Conceptio › Archive › arXiv CS
arXiv CSopen access

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

arXiv:2609.05198v1 [cs.AI] 4 Sep 2026

Zhinan Hou∗ Tsinghua University

Jiaqi Zhang† Meituan

Xunliang Cai Meituan

Keyou You† Tsinghua University

Abstract On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large language models in reasoning domains. However, the data-centric mechanisms in OPD remain relatively underexplored. This paper presents a empirical study of data efficiency and data selection in OPD. We begin by investigating an extreme setting: training OPD on only one example, namely 1-shot OPD. Surprisingly, we find that 1-shot OPD is consistently effective across all sampled training examples and harder examples often yield superior performance gain. We next investigate what actually drives the student model’s improvement in the training data. Our analysis reveals that the improvement is not driven by high token entropy, but the longer CoT paths which hard problem naturally generate. Training on longer CoT can help maintain closer alignment with the teacher over a long reasoning horizon, and learn critical thinking patterns usually missing in short CoTs, such as reflection (e.g., "Alternatively"). Based on these insights, we propose a simple data selection method that selects only hard examples for training, where even ”unsolvable” examples that completely exceed the teacher’s capability can be successfully used. Our experiments conducted on four models ranging from 1.5B to 7B show that training the student model on only 8 selected hard examples matches the performance of the 17K dataset baseline.

1

Introduction

On-policy distillation (OPD) has rapidly emerged as a core technique for large language model (LLM) post-training (Lu & Lab, 2025). OPD allows a smaller student model to align its policy with a stronger teacher model by learning directly from on-policy rollouts, enabling the student to effectively inherit complex reasoning capabilitie. Consequently, recent pioneering industry efforts, including Qwen3 (Yang et al., 2025), MiMo (Xiao et al., 2026), and GLM-5 (Zeng et al., 2026), have integrated OPD into their post-training pipelines, establishing it as an indispensable complement to supervised fine-tuning (SFT) and reinforcement learning with verified reward (RLVR). However, while researchers have focused heavily on designing better training algorithms (e.g. EOPD (Jin et al., 2026) and REOPOLD (Ko et al., 2026)), the data side of OPD remain relatively underexplored. Specifically, How much data is truly necessary? What data is most effective? And what actually drives the student model’s improvement in the training data? Answering these questions is of significant practical value because in many real-world scenarios, high-quality data are extremely scarce or expensive to collect, particularly in highly specialized domains such as medicine and engineering. This bottleneck makes a systematic study of data necessity highly important. To answer these questions, we first investigate a special setting where we train OPD on only one example, namely 1-shot OPD. Our empirical evaluations demonstrate that 1-shot OPD is consistently effective across all sampled training examples, which is also observed in a concurrent work ∗ This work was done during Zhinan’s internship at Meituan. † Correspondence to: [email protected], [email protected]

Preprint.

(Fu et al., 2026). Furthermore, we show that its training is exceptionally stable and the validation performance remains remarkably stable up to even 2,000 steps without overfitting or catastrophic policy collapse. We further analyze why 1-shot OPD works so well and find that the student policy aligns closest with the teacher on structural reasoning tokens (e.g. “Alternative”, “Wait”) after distillation. This indicates that the student can successfully learn and generalize the teacher’s reasoning patterns by distilling on only a single training example. Then, we investigate what kind of data makes OPD most effective. We discover a clear trend: training on harder examples often yields better performance than training on easy or medium ones. Interestingly, even when these examples are so difficult that they completely exceed the capabilities of both the teacher and student models, keeping the training accuracy strictly at 0%, training on these harder examples still allows the validation accuracy to continuously improve. We further analyze why the example difficulty can influence the OPD and what actually drives the student model’s improvement in the training data. We find that this improvement is not driven by high-entropy tokens, but is instead determined by the longer CoT. Specifically, we observe that training with longer CoTs helps the student maintain a closer alignment with the teacher across long horizons. Moreover, these longer paths help the model acquire thinking patterns that are rarely learned from short CoTs, such as self-reflection (e.g., ”Alternatively”). Based on these findings, we propose a simple, difficulty-driven data selection method that selects only hard examples for training. We then study how many of these examples are sufficient under this selection. Surprisingly, we find that training a 1.5B student model on only 8 selected hard examples matches the validation accuracy of the full 17K dataset baseline (53.6% vs 53.7%). Furthermore, we show that this extreme data efficiency generalizes robustly across diverse model architectures and scales, ranging from 1.5B to 7B parameters.

2

Preliminaries

On-Policy Distillation. OPD computes supervision on trajectories sampled from the current student policy πθ . Given a prompt x ∼ Dx , the student samples a response sequence ŷ = (ŷ1 , . . . , ŷN ) ∼ πθ (· | x). Both the student model πθ and the teacher model πT are evaluated on the student-generated prefixes ŷ<t , yielding two next-token probability distributions at each step t: pt (v) ≜ πθ (v | x, ŷ<t ) and qt (v) ≜ πT (v | x, ŷ<t ) over the vocabulary v ∈ V. A standard formulation of OPD minimizes the sequence-level reverse Kullback-Leibler (KL) divergence over the trajectories generated by the student: h i LOPD (θ) = Ex∼Dx DKL (πθ (· | x) ∥ πT (· | x)) . (1) Using autoregressive factorization, this sequence-level objective can be decomposed into an exact token-level formulation: " T # X LOPD (θ) = Ex∼Dx , ŷ∼πθ (·|x) DKL (pt ∥ qt ) . (2) t=1

In practice, existing OPD formulations vary in their supervision granularity for computing this token-level KL, generally categorizing into sampled-token, full-vocabulary, and top-k OPD. Among these, top-k OPD restricts the divergence to a subset of high-probability tokens, significantly reducing memory overhead while preserving dense, multi-token supervision in the student’s active generation region. Our experiments are primarily based on top-k OPD. Top-k OPD. Top-k OPD restricts the divergence computation to a subset of tokens St ⊆ V at each step. Specifically, we adopt the student top-k variant, which selects the k tokens assigned the highest probabilities under the student distribution, defined as St = TopK(pt , k). To compute the loss on this subset, we renormalize the student and teacher distributions over St : pt (v)1[v ∈ St ] (S ) p̄t t (v) = P , u∈St pt (u)

qt (v)1[v ∈ St ] (S ) q̄t t (v) = P . u∈St qt (u) 2

(3)

(S )

(S )

Distillation is then performed by minimizing the subset KL divergence DKL (p̄t t ∥ q̄t t ) over the trajectory, yielding the final objective: " T # X (St ) (St ) top-k DKL (p̄t ∥ q̄t ) . LOPD (θ) = Ex∼Dx , ŷ∼πθ (·|x) (4) t=1

3

Pilot Experiments: 1-Shot OPD

We first study an extreme setting: training a student model on only a single example, termed 1shot OPD. This special setting serves as a clean baseline to test whether OPD can achieve policy alignment without the influence of data variety. The observation of 1-shot OPD also provides the guidance for our subsequent analyses on what kind of data is suitable and how much data is sufficient. 3.1

Experimental Setup

Models. We employ DeepSeek-R1-Distill-Qwen-1.5B (Guo et al., 2025) as our student model, and its GRPO-trained counterpart JustRL-DeepSeek-1.5B (He et al., 2025) as the teacher model. By default, these two models are used for our primary experiments. We also conduct experiments across other model architectures and scales in Section 5.2 to validate the generalizability of our findings. Dataset. Due to computational resource limits, we randomly sample a subset of 1,000 examples from the DAPO-Math-17K dataset (Yu et al., 2026) to construct a example pool. To ensure reproducibility and consistent identification across all experiments. we simply rank the training examples by difficulty. Specifically, for each example in our pool, we compute the accuracy of both the student and teacher models across 16 rollouts per problem. Let Si and Ti denote the rollout accuracies of the student and teacher models for the i-th problem, respectively. The average accuracy is computed as Ai = (Si + Ti )/2. The training examples are sorted primarily in descending order of Ai , and secondarily by their original DAPO-Math-17K index to guarantee a deterministic sequence. The resulting ordered sequence is denoted as {πi }1000 i=1 . Based on Ai , we further categorize the training examples into three categories: (1) Easy (Ai > 0.9), corresponding to indices π1 to π199 (199 examples); (2) Medium (0.1 ≤ Ai ≤ 0.9), corresponding to indices π200 to π825 (626 examples); and (3) Hard (Ai < 0.1), corresponding to indices π826 to π1000 (175 examples). Training. OPD are implemented by using the open-source verl (Sheng et al., 2024) framework. Following Li et al. (2026a), the training batch size and mini-batch size are 64, and we sample 8 responses for each prompt. For response generation, we utilize the vLLM engine with a rollout temperature of 1.0 and maximum response length of 7,168 tokens. The per-token reverse KL divergence is computed against only the top-16 tokens. We training the OPD in the DAPO-Math-17K for 1 epoch as full dataset baseline (Full-Set OPD), yields 279 training steps. For fair comparsions, we train our 1-shot OPD under 279 training step by default. All experiment are executed on an 8 × NVIDIA H800 80GB GPU cluster. More training details are included in Appendix A. Evaluation. We evaluate the models across six standard mathematical reasoning benchmarks: AIME 2024 (Art of Problem Solving, a), AIME 2025 (Art of Problem Solving, a), AMC 2023 (Art of Problem Solving, b), MATH500 (Hendrycks et al., 2021), Minerva Math (Lewkowycz et al., 2022), and OlympiadBench (He et al., 2024). For all benchmarks, we sample with a temperature of 0.7 and a maximum output length of 16,384 tokens. Due to the limited sample sizes of AIME 2024, AIME 2025, and AMC 2023, we perform 16 rollouts per problem and report the average accuracy (mean@16) to ensure evaluation stability. For the remaining datasets, we report the average accuracy over 4 rollouts (mean@4). 3.2

1-Shot OPD is Effective for Many Examples

First, we randomly sample 8 training examples from each of the Easy, Medium, and Hard categories and train on them individually. The specific indices of sampled examples are listed in Table 5. During training, we utilize AMC 2023, AIME 2024, and AIME 2025 as the validation set to monitor performance changes over steps. Consistent with our evaluation, we sample 16 solutions 3

π105

π316

π954

0.60

1.00 0.60

1.00 0.60

1.00

0.54

0.95 0.54

0.95 0.54

0.95

0.48

0.90

0.42

0.85

0.36

0.48

0.90

0.42

0.85

0.36 100

200

0.90

0.42

0.85

0.36

0.80 0

0.48

0.80

279

0

100

π178

200

0.80

279

0

100

π794

200

279

π973

0.60

1.00 0.60

1.00 0.60

1.00

0.54

0.95 0.54

0.95 0.54

0.95

0.48

0.48

0.48

0.90

0.42

0.85

0.36

0.90

0.42

0.85

0.36 100 200 Training Step

279

0.85

0.36

0.80 0

0.90

0.42

0.80 0

100 200 Training Step

Average Validation Accuracy

279

Overlap Ratio

0.80 0

100 200 Training Step

279

Full Set OPD

Figure 1: Learning Dynamics of 1-shot OPD. We visualize the training trajectories of six representative examples across three difficulty levels: two easy problems ({π105 } and {π178 }, left), two medium problems ({π316 } and {π794 }, middle), and two hard problems ({π954 } and {π973 }, right). The solid blue curves denote the Average Validation Accuracy (computed across AMC 2023, AIME 2024, and AIME 2025). The solid orange curves with markers track the Overlap Ratio, indicating progressive token-level policy alignment with the teacher model. The horizontal dashed lines represent the full-set baseline (Full-Set OPD) for comparison. per problem with a temperature of 0.7 and a maximum response length of 16,384 tokens, reporting the average accuracy over the 16 samples as our validation metric. Additionally, we track the overlap ratio (Li et al., 2026a) to measure policy alignment between the student and the teacher. Specifically, we randomly select 100 examples from the DAPO-Math-17K as a held-out set. At each training checkpoint, we perform 16 independent rollouts for each example. Then we report the average proportion of tokens that appear simultaneously in the top-16 vocabulary distributions of both the student and the teacher as the overlap ratio. ¼178

¼794

Perhaps Given Wait wait Let What hold The Looking Because Yes Now ?\n\n Sup maybe Since

¼954

Perhaps Given What Wait Let Looking The wait Alternatively hold Because Since This Yes In Sup 0

0.1

0.2 Mean KL diff

0.3

0.4

Perhaps Given What Wait Looking The Alternatively Let wait hold Sup getting In This Because Hmm 0

0.1

0.2 0.3 Mean KL diff

0.4

0

0.1

0.2 0.3 Mean KL diff

0.4

Figure 2: Per-Token KL Divergence Reduction post 1-Shot OPD. We report the top 16 tokens with the highest expected KL reduction after distillation for models trained respectively on {π178 } (left), {π794 } (middle), and {π954 } (right). Highly optimized tokens are concentrated in structural reasoning tokens including reflection (e.g., “Alternatively”), transitions (e.g., “Wait”), and deductive reasoning (e.g.,“Because”). We present the evaluation results across benchmarks in Table 1. Moreover, we select some examples to visualize their training trajectories in Figure 1. As illustrated in Figure 1, the overlap ratio 4

continuously increases, indicating that the student model is progressively aligning with the teacher model. Moreover, the performance improvement demonstrates the effectiveness of 1-shot OPD. Furthermore, we perform a prolonged training run of 2,000 steps to observe whether 1-shot OPD suffers from overfitting. As shown in Figure 6, the validation performance plateaus after 300 steps and remains stable up to 2,000 steps without any performance decay or policy collapse. To understand why 1-shot OPD is so effective, we analyze which tokens shift closest to the teacher after distillation. Specifically, we use the held-out set (100 problems) to generate 16 rollouts per problem. Then we compute the reduction in token-level KL divergence between the student and teacher distributions post-distillation and average this KL reduction over all token IDs. For statistical significance, we discard the rare tokens with an overall occurrence frequency of less than 0.01%. In Figure 2, we report the top 16 tokens with the highest KL reduction for models trained respectively on {π178 }, {π794 }, and {π954 }. The results reveal that the tokens aligning most significantly with the teacher after OPD are predominantly structural reasoning, including reflection (e.g., “Alternatively”), transitions (e.g., “Wait”, “Perhaps”), and deductive reasoning (e.g., “Because”, “Since”). This indicates that the student can successfully learn the teacher’s reasoning patterns by distilling on only a single training example. Table 1: 1-Shot OPD Performance Across Mathematical Reasoning Benchmarks. We compare the student model DeepSeek-R1-Distill-Qwen-1.5B (Base), the teacher model JustRL-DeepSeek1.5B (Teacher), and the full-dataset distillation baseline (Full-Set) against our 1-shot OPD models. Notably, 1-shot models trained on hard problems consistently outperform those trained on easy problems, with the best-performing example ({π973 }) achieving 51.7% average accuracy, nearly matching the Full-Set baseline (53.7%).

Model

AIME24

AIME25

AMC23

MATH500

Olympiad

Minerva

Avg

Base Teacher Full-Set

31.7 55.8 51.9

23.0 35.8 34.0

60.7 83.4 78.9

82.6 87.4 86.7

34.1 43.4 44.3

22.4 29.8 26.6

42.4 55.9 53.7

Easy Problems {π5 } 39.2 {π20 } 40.3 {π21 } 41.8 {π63 } 42.2 {π84 } 41.0 {π105 } 42.5 {π178 } 41.1 {π192 } 39.8

28.8 30.6 29.0 30.8 30.1 27.7 30.4 30.2

69.9 69.1 71.1 71.0 72.6 70.3 71.2 72.0

85.3 86.2 85.9 86.2 85.6 86.0 86.7 85.7

39.3 40.8 39.7 39.7 40.6 41.1 40.4 40.1

25.3 25.7 26.2 25.0 25.8 25.0 25.9 25.7

48.0 48.8 48.9 49.1 49.3 48.8 49.3 48.9

Medium Problems {π316 } 39.8 {π487 } 41.2 {π530 } 42.1 {π543 } 40.4 {π661 } 43.1 {π763 } 41.5 {π794 } 44.2 {π820 } 41.5

30.0 29.8 30.7 29.4 31.2 28.1 30.8 29.1

73.3 72.7 73.1 73.0 73.4 75.6 74.2 73.6

86.4 85.7 86.5 86.4 86.9 86.7 86.5 84.9

40.6 40.4 40.2 41.3 41.3 41.2 41.6 40.7

25.7 25.9 26.8 26.9 26.4 26.8 25.7 25.9

49.3 49.3 49.9 49.6 50.4 50.0 50.5 49.3

Hard Problems {π874 } 43.3 {π890 } 45.2 {π948 } 45.0 {π954 } 44.8 {π961 } 44.4 {π973 } 47.5 {π987 } 44.8 {π997 } 44.0

33.1 29.6 29.8 28.5 30.6 31.7 27.7 30.2

74.3 76.0 75.3 76.9 73.2 75.1 73.6 75.8

86.6 85.6 86.0 86.9 87.2 87.7 87.5 85.6

41.0 42.4 41.8 42.6 41.4 41.5 41.2 42.1

27.1 26.5 27.0 26.7 26.7 26.9 26.4 26.8

50.9 50.9 50.8 51.1 50.6 51.7 50.2 50.8

5

4

Data Selection for OPD

4.1

Harder Problems Often Lead to Better Performance

Average Validation Accuracy

0.54

0.8

0.48

0.6 0.4

0.42 0.36

Average Training Accuracy

1.0

0.2 0

100 Training Step

200

Easy

279

0.0

0

Medium

100 Training Step

200

279

Hard

Figure 3: Learning Trajectories across Difficulty Levels. Left: Average Validation Accuracy over training steps. Right: Average Training Accuracy on the selected training example over training steps. Solid lines indicate the mean and shaded areas represent the standard deviation across the examples in each category. We further investigate whether different data behave differently in 1-shot OPD. For each category, we record both the average training accuracy and validation accuracy every 50 training steps, and present these trajectories in Figure 3. Surprisingly, we find that harder training examples often yield better performance. Moreover, we observe that for easy and hard problems, even after their training accuracies saturate at 100% (for easy tasks) or remain flat at 0% (for hard tasks), their validation accuracies continuously improve. This behavior contrasts sharply with RL (e.g., GRPOstyle algorithms) where training on easy and hard samples is highly ineffective because the computed advantage or gradient updates tend to be zero. 4.2

Deeper Analysis

Then we investigate whether the advantage of difficult problems is caused by their higher token entropy or by the longer CoT reasoning sequences they naturally generate. Setup. To separate the effects of token entropy and CoT length, we set up a length-controlled experiment. By adjusting the maximum rollout length, we expect that almost all rollouts generated for some examples will exceed this limit and get cut off. Consequently, training the 1-shot OPD uses almost the same number of tokens for those examples because the length of every response is forced to be the maximum rollout length. Specifically, we conduct three sets of experiments. In each set, we select two problems with different difficulty or entropy, and set an appropriate maximum rollout length such that empirically over 99.5% of the rollouts for both problems exceed this threshold. The three problem pairs are: 1) Easy problem π178 and Hard problem π973 limited to 1024 tokens, 2) Medium problem π316 and Hard problem π954 limited to 2048 tokens, and 3) two Hard problems with highly different token entropy, π890 and π948 limited to 4096 tokens. Higher Token Entropy Alone Does Not Improve OPD As illustrated in Figure 4, under lengthcontrolled settings, we observe no significant difference in validation performance between the compared pairs, despite their highly divergent token entropy. Specifically, once rollout lengths are restricted to match, the performance gap between the pairs collapses to 0.5% (for {π973 } vs. {π178 }) and −0.9% (for {π954 } vs. {π316 }). However, in the default setting (7,168 maximum rollout length) where the student model can leverage the longer CoT sequences generated by harder examples, the harder problems demonstrate a notable advantage, outperforming {π178 } and {π316 } by 3.9% and 2.4% in validation accuracy, respectively. This collapse under length limits indicates that higher token entropy alone does not translate to better distillation performance. Additionally, we perform an auxiliary experiment where we sample responses for the Easy problem π178 under higher temperatures to artificially increase token entropy. As shown in Figure 7, the resulting model shows no 6

Entropy

0.8

OPD under 2K tokens

OPD under 4K tokens

π178 π973

π316 π954

π948 π954

1.0

0.6

1.0 0.8

0.4

0.2 Average Validation Accuracy

OPD under 1K tokens

0.8

0.6

0

100

200

279

0.4

0

100

200

279

0.6

0.54

0.54

0.54

0.48

0.48

0.48

0.42

0.42

0.42

0.36

0.36

0.36

0

100 200 Training Step

279

0

100 200 Training Step

279

0

100

200

279

0

100 200 Training Step

279

Figure 4: Training Dynamics in Length-Constrained Setting Training dynamics are presented for problem pairs grouped by difficulty: Easy vs. Hard (1K token limit, left), Medium vs. Hard (2K token limit, middle), and Hard vs. Hard (4K token limit, right). The top plots monitor token entropy, while the bottom plots track validation accuracy.

accuracy gains, indicating that the performance improvement of difficult tasks is primarily driven by their longer CoT paths rather than token entropy. Distillation Benefits from Longer CoTs We further KL loss compare student models distilled on the hard problem {π954 } under maximum rollout length constraints of 0.07 2048, 4096, and 7168 tokens (referred to as the 2K, 4K, and 7K models, respectively). For evaluation, we gen- 0.05 erate 16 rollout completions per problem across the 100 held-out problems. At each generation step, we compute the per-token KL divergence between the distilled student 0.03 2K 4K and the teacher policies and plot the average KL loss over 7K positions in Figure 5. The empirical results in Figure 5 0.01 0 4000 8000 12000 16000 reveal that the 7K model consistently maintains a lower Position KL divergence across the entire generation horizon compared to the 2K and 4K models. Crucially, while the Figure 5: Token-level KL Loss over alignment of the 4K model is nearly identical to that of Generation Position. For visual clarthe 7K model within the 0 to 4, 000 token range, the KL ity, we partition the token positions into gap between them widens significantly as the generation non-overlapping intervals of 100 tokens horizon expands beyond 4, 000 tokens. This divergence and plot the binned average KL loss demonstrates that training on longer CoT sequences di- over positions. rectly enhances the model’s long horizon reasoning capabilities, enabling it to sustain policy alignment with the teacher over extended context lengths. Harder Problems can Trigger More Reasoning Patterns. Furthermore, we observe that training on harder problems does not simply scale token volume, but instead learns more reasoning patterns, such as self-reflection and backtracking. As shown in Figure 2, for the model training on easy problem ({π178 }) where trajectories are naturally short, the self-reflection token "Alternatively" is absent from the top-16 KL reduction list. However, it emerges in the medium model ({π794 }), and ranks higher in the hard model ({π954 }). This finding suggests that longer CoT trajectories allows the student model to acquire some sophisticated reasoning patterns that are inherently absent in shorter, simpler reasoning paths. 7

5

How Much Data is Sufficient for OPD?

5.1

Few-Shot OPD Can Match Full-Set Performance

Based on Section 4, we simply select hard examples as training set and investigate how scaling the training set size affects performance. Specifically, we scale the number of Hard training examples N ∈ {1, 4, 8, 16, 64}. To construct our training sets, for N = 1, 4, and 8, we directly use the hard examples evaluated in Table 5 (i.e., {π973 } for 1-shot, {π874 , π890 , π954 , π973 } for 4-shot, and {π874 , . . . , π997 } for 8-shot). For the larger sizes of N = 16 and 64, in addition to these 8 examples, we randomly sample the remaining examples from our hard example pool. For comparison, we also evaluate Easy and Medium baselines of size N = 8 using the eight corresponding examples (i.e., {π5 , π20 , . . . , π192 } for Easy, and {π316 , π487 , . . . , π820 } for Medium). Table 2 reports these results. From the results in Table 2, we find that: 1) increasing the training set of Hard examples from N = 1 to N = 8 steadily improves performance, with 8 examples already matching the full-dataset (17K) distillation baseline, while further increasing the size to 16 or 64 examples yields no additional gains; and 2) consistent with our 1-shot findings, under a fixed size of 8 examples, training on Hard examples yields the best performance compared to Medium and Easy ones. Table 2: Few-Shot OPD Performance across Different Dataset Sizes. The baselines include the full-dataset distillation (DAPO-Math-17K) and a randomly sampled 1K subset (DAPO-Subset). Distilling on just 8 hard examples achieves 53.6% average accuracy, nearly matching the Full-Set 17K baseline (53.7%). Dataset

Size

AIME24

AIME25

AMC23

MATH500

Olympiad

Minerva

Avg

DAPO-Math-17K DAPO-Subset

17K 1K

51.9 50.2

34.0 34.8

78.9 79.3

86.7 87.1

44.3 42.8

26.6 27.3

53.7 53.6

{π5 , π20 , . . . , π192 } {π316 , π487 , . . . , π820 }

8 8

45.8 49.0

30.6 32.7

76.1 76.8

87.1 86.8

41.6 42.1

26.2 26.0

51.2 52.2

{π973 } {π874 , π890 , π954 , π973 } {π874 , . . . , π997 } {π874 , . . . , π997 , . . . } {π874 , . . . , π997 , . . . }

1 4 8 16 64

47.5 50.8 51.7 49.4 50.2

31.7 31.5 34.2 33.3 36.7

75.1 78.5 79.2 78.5 79.1

87.7 86.7 87.2 87.7 87.8

41.5 41.5 42.7 43.0 42.7

26.9 28.2 26.4 27.7 27.1

51.7 52.9 53.6 53.3 53.9

5.2

Few-Shot OPD on Other Models

To investigate whether this extreme data efficiency generalize across other models, we additionally evaluate three teacher–student pairs: 1) Skywork-OR1-Math-7B → DeepSeek-R1-Distill-Qwen7B; 2) Qwen3-4B → Qwen3-1.7B-Base; 3) Qwen3-4B → Qwen3-4B-Base. The setting vary scale, capability, and backbone architecture. For Qwen3-4B, we disable thinking mode. Same as Section 5.1, we train the student models on the selected hard subsets of sizes 1, 4, and 8. Table 3 shows the performance of the student model, teacher model, and distilled models. Across all three evaluated pairings, few-shot OPD on hard examples consistently triggers substantial validation performance gains. Specifically, on the Qwen3-1.7B-Base and Qwen3-4B-Base student models, 8-shot training ({π874 , . . . , π997 }) achieves average accuracies of 21.3% and 29.2%, approaching the performance achieved by the 17K full-set baseline (22.5% and 30.8%, respectively). Similarly, for the DeepSeek-7B student, 1-shot OPD on {π973 } already yields a highly competitive average accuracy of 58.4%, which further scales to 59.5% at 8-shot, matching the full-set performance of 59.6%.

6

Related Work

On-Policy Distillation. Distillation for LLMs is broadly categorized into off-policy and on-policy settings. While off-policy distillation (e.g., SFT) suffers from exposure bias due to the mismatch between static teacher-generated data and student-generated rollouts at inference (Agarwal et al., 2024), OPD mitigates this issue by training on trajectories sampled directly from the student’s current policy under a reverse KL objective (Gu et al., 2024). Recently, OPD has since appeared in 8

Table 3: 1(few)-shot OPD is effective for different models. We report benchmark accuracies across diverse model pairings and training sizes. Model

Size

AIME24

AIME25

AMC23

MATH500

Olympiad

Minerva

Avg

Skywork-OR1-Math-7B → DeepSeek-R1-Distill-Qwen-7B DeepSeek-R1-Distill-Qwen-7B Skywork-OR1-Math-7B

– –

51.9 62.1

38.8 45.6

81.0 84.0

89.6 91.9

43.6 47.7

32.6 34.1

56.3 60.9

Full-Set {π973 } {π874 , π890 , π954 , π973 } {π874 , . . . , π997 }

17K 1 4 8

58.3 56.9 57.1 58.3

44.6 42.7 42.9 43.9

83.4 81.0 82.2 83.1

91.1 90.6 91.1 91.3

46.0 45.4 46.0 46.3

34.2 33.5 33.0 34.0

59.6 58.4 58.7 59.5

Qwen3-1.7B-Base Qwen3-4B

– –

1.4 22.9

2.5 17.1

12.3 62.0

13.3 81.1

4.8 45.4

3.7 26.1

6.3 42.4

Full-Set {π973 } {π874 , π890 , π954 , π973 } {π874 , . . . , π997 }

17K 1 4 8

10.2 4.8 6.7 9.4

4.0 3.5 3.8 4.4

28.6 23.7 25.5 26.8

53.9 45.5 49.6 51.0

24.4 18.3 21.7 23.7

14.1 10.7 11.9 12.7

22.5 17.8 19.9 21.3

Qwen3-4B-Base Qwen3-4B

– –

7.3 22.9

5.0 17.1

24.6 62.0

22.3 81.1

11.9 45.4

4.8 26.1

12.7 42.4

Full-Set {π973 } {π874 , π890 , π954 , π973 } {π874 , . . . , π997 }

17K 1 4 8

14.4 12.3 11.2 13.1

14.4 10.4 11.2 12.5

40.5 35.6 36.8 38.6

65.6 57.7 61.6 64.1

33.3 26.7 30.2 31.3

16.5 14.7 15.2 15.6

30.8 26.2 27.7 29.2

Qwen3-4B → Qwen3-1.7B-Base

Qwen3-4B → Qwen3-4B-Base

some reasoning and post-training recipes, such as Qwen3 (Yang et al., 2025), Deepseek-V4 (Xu et al., 2026) and GLM-5 (Zeng et al., 2026). Building on these successes, a growing body of work has investigated various algorithmic variants of OPD (Oh et al., 2026; Yang et al., 2026), optimized its practical training recipes (Li et al., 2026b; Ko et al., 2026), and integrated it into large-scale posttraining pipelines (Ma et al., 2026). Moreover, several studies have further analyzed the underlying mechanisms of OPD (Wang et al., 2026b; Li et al., 2026a). Data Selection for LLM Post-Training. Selecting a small yet effective subset of training data has become a important problem for improving the efficiency of LLM post-training, with most efforts focusing on data selection for SFT (Ivison et al., 2025) and RLVR (Wu et al., 2026). Existing approaches include LLM-based quality assessment (Chen et al., 2024), gradient-based selection (Xia et al., 2024) and more. However, data selection for OPD remains underexplored. Our work is inspired by recent advances in RLVR, which show that RLVR can achieve significant reasoning gains with as few as a single training example (Wang et al., 2026a). We extend this exploration to the OPD paradigm, discovering different and interesting phenomena. A concurrent work by Fu et al. (2026) also investigates 1-shot OPD and analyzes its efficacy from a state-space coverage perspective. Differently, we analyze 1-shot OPD from the perspective of CoT trajectory and reasoning patterns. Then we propose a simple and effective data selection method that selects only hard examples for training, where even “unsolvable” examples completely exceeding the teacher’s capability can be successfully used.

7

Conclusions

In this work, we show that 1-shot OPD is sufficient to trigger substantial improvements in mathematical reasoning tasks, with as few as 8 selected examples nearly matching the performance of distilling on the full DAPO-Math-17K dataset. Moreover, we observe that distilling on harder tasks consistently outperforms easy ones, and even “unsolvable” tasks where both the teacher and student models fail to produce correct answers still yield strong performance gains. Our analysis reveal that this improvement is fundamentally driven by the longer CoT trajectories rather than high tokenlevel entropy. Future work includes: 1) designing effective dynamic batching strategies algorithms of OPD based on our findings; and 3) developing novel OPD algorithms that can learn from longer CoT reasoning trajectories, where the teacher’s dense reward may lose local exploitability. 9

References Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from selfgenerated mistakes. In International Conference on Learning Representations, volume 2024, pp. 21246–21263, 2024. Art of Problem Solving. Aime problems and solutions. https://artofproblemsolving.com/ wiki/index.php/AIME_Problems_and_Solutions, a. Accessed: 2025-04-20. Art of Problem Solving. Amc problems and solutions. https://artofproblemsolving.com/ wiki/index.php?title=AMC_Problems_and_Solutions, b. Accessed: 2025-04-20. Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, et al. Alpagasus: Training a better alpaca with fewer data. In International Conference on Learning Representations, volume 2024, pp. 34767–34797, 2024. Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, et al. Rethinking on-policy distillation of large language models ii: One training example. arXiv preprint arXiv:2609.04172, 2026. Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. Minillm: Knowledge distillation of large language models. In International Conference on Learning Representations, volume 2024, pp. 32694–32717, 2024. Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. Bingxiang He, Zekai Qu, Zeyuan Liu, Yinghao Chen, Yuxin Zuo, Cheng Qian, Kaiyan Zhang, Weize Chen, Chaojun Xiao, Ganqu Cui, et al. Justrl: Scaling a 1.5 b llm with a simple rl recipe. arXiv preprint arXiv:2512.16649, 2025. Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, et al. Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3828–3850, 2024. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. Hamish Ivison, Muru Zhang, Faeze Brahman, Pang Wei Koh, and Pradeep Dasigi. Large-scale data selection for instruction tuning. arXiv preprint arXiv:2503.01807, 2025. Woogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei, Yi Zhou, Swanand Ravindra Kadhe, Nathalie Baracaldo, and Kimin Lee. Entropy-aware on-policy distillation of language models. arXiv preprint arXiv:2603.07079, 2026. Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen, and Pashmina Cameron. Scaling reasoning efficiently via relaxed on-policy distillation. arXiv preprint arXiv:2603.11137, 2026. Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. Solving quantitative reasoning problems with language models. Advances in neural information processing systems, 35:3843–3857, 2022. Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huanang Gao, Wenkai Yang, Zhiyuan Liu, et al. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe. arXiv preprint arXiv:2604.13016, 2026a. 10

Yuying Li, Leqi Zheng, Yongzi Yu, Wenrui Zhou, Xuchang Zhong, Xing Hu, Jing Jin, Hangjie Yuan, and Tao Feng. Filter, then reweight: Rethinking optimization granularity in on-policy distillation. arXiv preprint arXiv:2606.02684, 2026b. Kevin Lu and Thinking Machines Lab. On-policy distillation. Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy-distillation. Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, et al. Mopd: Multi-teacher on-policy distillation for capability integration in llm post-training. arXiv preprint arXiv:2606.30406, 2026. Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi, and Yohan Jo. Kl for a kl: On-policy distillation with control variate baseline. arXiv preprint arXiv:2605.07865, 2026. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework. arXiv preprint arXiv: 2409.19256, 2024. Yiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren, Liyuan Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, Jianfeng Gao, et al. Reinforcement learning for reasoning in large language models with one training example. Advances in Neural Information Processing Systems, 38: 122721–122764, 2026a. Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, and Hongxia Yang. Not all disagreement is learnable: Token teachability in on-policy distillation. arXiv preprint arXiv:2605.26844, 2026b. Jianghao Wu, Jianfei Cai, Weiqiang Wang, Jin Ye, Daniel F Schmidt, and Yasmeen George. Single-rollout hidden-state dynamics for training-free rlvr data selection. arXiv preprint arXiv:2605.28631, 2026. Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. Less: Selecting influential data for targeted instruction tuning. arXiv preprint arXiv:2402.04333, 2024. Bangjun Xiao, Bingquan Xia, Bo Yang, Bofei Gao, Bowen Shen, Chen Zhang, Chenhong He, Chiheng Lou, Fuli Luo, Gang Wang, et al. Mimo-v2-flash technical report. arXiv preprint arXiv:2601.02780, 2026. Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, et al. Deepseek-v4: Towards highly efficient milliontoken context intelligence. arXiv preprint arXiv:2606.19348, 2026. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. Zhicheng Yang, Zhijiang Guo, Yifan Song, Minrui Xu, Yongxin Wang, Yiwei Wang, Xiaodan Liang, and Jing Tang. Prune-opd: Efficient and reliable on-policy distillation for long-horizon reasoning. arXiv preprint arXiv:2605.07804, 2026. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. Advances in Neural Information Processing Systems, 38:113222–113244, 2026. Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026.

11

A

More Training Details

For every input prompt during the OPD training, we generate n = 8 responses. We constrain the maximum prompt length to 1,024 tokens, while the maximum response length is limited to 7,168 tokens. The optimization process spans a single epoch across 8 H800 80G GPUs, utilizing a learning rate of 1 × 10−6 . We set both the student sampling temperature and the teacher temperature to 1.0, disable KL regularization, and adopt token-mean loss aggregation. Unless otherwise noted, all experiments use the default OPD hyperparameters listed in Table 4. Table 4: Default hyperparameters for OPD. Item

Value

Training temperature Global batch size Mini batch size Rollout number Temperature LogProb top-K Top-K strategy Top-p Max prompt length Max response length Learning rate Epoch KL Coefficient

B

More Experiments

B.1

1-shot OPD over Extended Training

1.0 64 64 8 1.0 16 Student Top-K 1.0 1024 7168 1e-6 1 0.0

We further investigate whether 1-shot OPD suffers from overfitting. To do so, we perform a prolonged training experiment by extending the optimization duration to 2,000 steps for 1-shot OPD training on {π105 }, {π794 }, and {π973 }, respectively. As shown in Figure 6, we observe that the validation accuracy across all three models rises steadily during the first 300 steps and subsequently remains remarkably stable up to 2,000 steps, showing no signs of performance decay or catastrophic policy collapse. This behavior is fundamentally different from SFT and RL in single-example settings (Wang et al., 2026a), which often suffer from severe overfitting or policy degradation under extended optimization. One potential reason is that the sequence-level KL divergence between the teacher and student models naturally decreases as training progresses, leading to loss convergence and preventing policy collapse even under prolonged training.

Accuracy

8

π794 0.60

8

π973 0.60

0.54

6 0.54

6 0.54

0.48

0.48

0.48

4

0.42

4

0.42 2

0.36 0

500

1000 1500 Training Step

0 2000

8

6

4

0.42 2

0.36 0

500

1000 1500 Training Step

Average Validation Accuracy

2

0.36

0 2000

0

500

1000 1500 Training Step

Loss (×1000)

π105 0.60

0 2000

Loss

Figure 6: Validation Accuracy over Extended Training in 1-Shot OPD. We plot the validation trajectories over 2,000 steps for 1-shot OPD models trained on three representative examples of varying difficulty levels: {π105 } (Easy, left), {π794 } (Medium, middle), and {π973 } (Hard, right). Despite extended training, the validation performance across all models remains exceptionally stable without policy collapse.

12

0.55 Average Validation Accuracy

2.0

Entropy

1.6

1.2

0.8

0.4 0

100 Training Step

T = 0.6

200

279

T = 0.8

0.50

0.45

0.40

0.35 0

100 Training Step

T = 1.0

200

279

T = 1.1

Figure 7: Study on student rollout temperature T for the easy example π178 during 1-shot OPD. The left subfigure monitors the average per-token entropy of the student’s rollouts across training steps under different sampling temperatures T ∈ {0.6, 0.8, 1.0, 1.1}. The right subfigure tracks the corresponding Average Validation Accuracy (computed over AMC 2023, AIME 2024, and AIME 2025). B.2

1-shot OPD under Different Temperature

We conduct an auxiliary experiment by varying the student rollout temperature T in {0.6, 0.8, 1.0, 1.1} during 1-shot OPD. This experiment is specifically designed around one easy problem π178 . By adjusting the sampling temperature of the rollout, we artificially inflate the tokenlevel entropy of the student model’s rollouts without altering the inherent difficulty of the training example itself.

C

Limitations

A key limitation of our work is that our difficulty-driven data selection is a simple empirical choice rather than a mathematically optimal one. Finding the theoretically optimal training sets for OPD may be computationally intractable. However, despite not being theoretically optimal, our method is extremely simple to use. It requires zero complex optimization, zero extra training parameters, and can be implemented simply by sorting accuracy of examples. We believe this extreme simplicity and ease of implementation make our findings highly valuable for practical applications.

D

Discussions

CoT Length as a Practical Proxy in Verifier-Free Domains. In mathematical and coding domains, rule-based verifiers are readily available to compute rollout accuracy for task selection. However, in many general reasoning domains, such verifiers do not exist, making diffculty-based selection impossible. For these domains, our findings suggest a simple and effective alternative: using the length of the CoT as a proxy for data selection. As demonstrated in Section 4, CoT length is the key driver of policy alignment. Distilling on longer CoT sequences not only enables the student model to maintain alignment stability over a long reasoning horizon, but also successfully activates advanced reasoning patterns (e.g. reflection), that may miss in short trajectories. The Limits of Scaling Rollout Horizon and the Exposure Bias Trade-off. A natural question is whether continuously increasing the maximum response length will always yield better policy alignment. We argue that there exists a empirical trade-off: while longer trajectories contain richer reasoning patterns, excessively long on-policy rollouts suffer from severe exposure bias. Because OPD optimizes the student policy against its own dynamically sampled trajectories, any reasoning errors made by the student model will compound over a long horizon. This compounding error drifts the student into out-of-distribution states, which significantly degrades the quality of the teacher’s policy supervision. The maximum rollout length 7,618 is a empirical choice following Li et al. (2026a). To effectively train on even longer CoT sequences without policy decay, we leave the 13

design of advanced training algorithms (such as Prune-OPD (Yang et al., 2026)) as a promising direction for future work.

E

Example Details

E.1

Accuracy of Sampled Examples

To classify the difficulty of tasks in our example pool and select representative problems, we evaluate the baseline performance of both the student and teacher models. Specifically, we run 16 independent rollouts per problem for the student model (DeepSeek-R1-Distill-Qwen-1.5B) and the teacher model (JustRL-DeepSeek-1.5B), respectively. For each problem i, let Si and Ti denote the rollout accuracies of the student and teacher models. We then compute the average accuracy as Ai = (Si + Ti )/2 to represent the task difficulty. Based on this metric, we rank all training examples and group them into three categories: (1) Easy (Ai > 0.9), (2) Medium (0.1 ≤ Ai ≤ 0.9), and (3) Hard (Ai < 0.1). Table 5 reports the detailed rollout accuracies of both models on the representative training problems selected in our experiments. Table 5: OPD Dataset Preparation Accuracies. We report the rollout accuracies of both the student and teacher models across 16 rollouts per problem during our data preparation phase. Easy Problems

E.2

Medium Problems

Hard Problems

ID

Student

Teacher

ID

Student

Teacher

ID

Student

Teacher

π5 π20 π21 π63 π84 π105 π178 π192

100.0 100.0 100.0 93.8 93.8 87.5 87.5 87.5

100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0

π316 π487 π530 π543 π661 π763 π794 π820

68.8 37.5 25.0 50.0 25.0 6.3 0.0 6.3

100.0 100.0 100.0 75.0 68.8 37.5 37.5 18.8

π874 π890 π948 π954 π961 π973 π987 π997

0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0

6.3 0.0 0.0 0.0 0.0 0.0 0.0 0.0

Content of Sampled Examples Table 6: Details of example π5 .

Prompt: Let $a$ and $b$ be positive real numbers . Find the minimum value of \[ a ^2 + b ^2 + \ frac {1}{( a + b ) ^2}.\] The answer is in the form k \ sqrt { m }+ n ,. Please provide the value of k + m + n . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 3. Table 7: Details of example π20 . Prompt: In trapezoid $ABCD$ the lengths of the bases $AB$ and $CD$ are 8 and 17 respectively . The legs of the trapezoid are extended beyond $A$ and $B$ to meet at point $E$ . What is the ratio of the area of triangle $EAB$ to the area of trapezoid $ABCD$ ? Express your answer as a common fraction . The answer is in the form rac { m }{ n } , where gcd (m , n ) = 1. Please provide the value of m + n . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 289.

14

Table 8: Details of example π21 . Prompt: Lines $l_1 ^{} $ and $l_2 ^{} $ both pass through the origin and make first quadrant angles of $ \ frac {\ pi }{70} $ and $ \ frac {\ pi }{54} $ radians , respectively , with the positive $x$ - axis . For any line $l$ , the tran sformati on $R ( l ) $ produces another line as follows : $l$ is reflected in $l_1$ , and the resulting line is reflected in $l_2$ . Let $R ^{(1) }( l ) = R ( l ) $ and $R ^{( n ) }( l ) = R \ left ( R ^{( n -1) }( l ) \ right ) $ . Given that $l$ is the line $y =\ frac {19}{92} x$ , find the smallest positive integer $m$ for which $R ^{( m ) }( l ) = l$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 945. Table 9: Details of example π63 . Prompt: On the party , every boy gave $1$ candy to every girl , and every girl gave $1$ candy to every boy . Then every boy ate $2$ candies , and every girl ate $3$ candies . It is known that $ \ frac {1}{4} $ of all candies were eaten . Find the greatest possible number of children at the party . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 35. Table 10: Details of example π84 . Prompt: The prime numbers $a$ , $b$ , and $c$ satisfy the equation $a + b ^2 = 4 c ^2 $ . Determine the sum of all possible values of $a + b + c$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 31. Table 11: Details of example π105 . Prompt: Compute the smallest positive integer $N$ for which $N \ cdot 2^{2024} $ is a multiple of $2024$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 253. Table 12: Details of example π178 . Prompt: Determine the number of pairs $ (a , b ) $ of real numbers such that $10 , a , b , ab$ is an arithmetic progression . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 2. Table 13: Details of example π192 . Prompt: Compute the remainder when $2 ^{3^5}+ 3^{5^2}+ 5^{2^3} $ is divided by $30$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 6.

15

Table 14: Details of example π316 . Prompt: A dartboard consists of three concentric circles with radii 4 , 6 , and 8. Three darts are thrown at the board , sticking at random locations . Determine the probability that each dart lands in a different region of the dartboard . The probability can be expressed as \( \ frac { m }{ n } \) , where \( m \) and \( n \) are relatively prime positive integers . Calculate \( m + n \) . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 617.

Table 15: Details of example π487 . Prompt: In how many ways can the integers from 1 to n be ordered subject to the condition that , except for the first integer on the left , every integer differs by 1 from some integer to the left of it ? Please provide the number of ways for $n = 6 $ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 32.

Table 16: Details of example π530 . Prompt: Determine the largest integer $N$ for which there exists a $6 \ times N$ table $T$ that has the following properties :\ n \n - Every column contains the numbers $1 , 2 , \ ldots , 6 $ in some ordering .\ n - For any two columns $i \ ne j$ , there exists a row $r$ such that $T (r , i ) = T (r , j ) $ .\ n - For any two columns $i \ ne j$ , there exists a row $s$ such that $T (s , i ) \ ne T (s , j ) $ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 120.

Table 17: Details of example π543 . Prompt: Let f be a polynomial of degree $3$ with integer coefficients such that $f (0) = 3 $ and $f (1) = 11 $ . If f has exactly $2$ integer roots , how many such polynomials $f$ exist ? Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 0.

Table 18: Details of example π661 . Prompt: Find the straight sqrt { x ^3 and put

area of the figure on the coordinate plane bounded by the lines $x = 0$ , $x = 2$ , and the graphs of the functions $y = \\ + 1} $ and $y = -\\ sqrt [3]{ x ^2 + 2 x } $ . Please reason step by step , your final answer within \\ boxed {}.

Ground truth: 10.

16

Table 19: Details of example π763 . Prompt: An organization has $30$ employees , $20$ of whom have a brand A computer while the other $10$ have a brand B computer . For security , the computers can only be connected to each other and only by cables . The cables can only connect a brand A computer to a brand B computer . Employees can communicate with each other if their computers are directly connected by a cable or by relaying messages through a series of connected computers . Initially , no computer is connected to any other . A technician arbitrarily selects one computer of each brand and installs a cable between them , provided there is not already a cable between that pair . The technician stops once every employee can communicate with each other . What is the maximum possible number of cables used ? Please reason step by step , and put your final answer within \\ boxed {}.

Ground truth: 6.

Table 20: Details of example π794 . Prompt: Tanya wrote numbers in the form $n ^7 - 1 $ for $n = 2 , 3 , \ ldots$ and noticed that for $n = 8$ , she obtained a number divisible by $337$ . For what minimal $n$ did she get a number divisible by $2022$ ? Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 79.

Table 21: Details of example π820 . Prompt: Vasya has $n$ candies of several types , where $n > 145 $ . It is known that for any group of at least 145 candies , there is a type of candy which appears exactly 10 times . Find the largest possible value of $n$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 160.

Table 22: Details of example π874 . Prompt: Find the number of subsets of $ \{1 ,3 ,5 ,7 ,9 ,11 ,13 ,15 ,17 ,19\} $ where the elements in the subset add to $49$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 22.

Table 23: Details of example π890 . Prompt: Let $x_1 , x_2 , \ dots , x_ {100} $ be real numbers such that $ | x_1 | = 63 $ and $ | x_ { n +1}| = | x_n + 1| $ for $n = 1 , 2 , \ dots , 99 $ . Find the largest possible value of $ ( - x_1 - x_2 - \ cdots - x_ {100}) $ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 2034.

17

Table 24: Details of example π948 . Prompt: Two real numbers $x$ and $y$ are chosen at random in the interval (0 ,1) with respect to the uniform distribution . What is the probability that the closest integer to $x / y$ is even ? The original answer is in the form $r + s \ pi$ , please give the value of $r + s$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 4.

Table 25: Details of example π954 . Prompt: Let $f :\{1 ,2 ,\ dots ,2019\}\ to \{ -1 ,1\} $ be a function , such that for every $k \ in \{1 ,2 ,\ dots ,2019\} $ , there exists an $ \ ell \ in \{1 ,2 ,\ dots ,2019\} $ such that $$ \ sum_ { i \ in \ mathbb { Z }:(\ ell - i ) (i - k ) \ geqslant 0} f ( i ) \ leqslant 0. $$ Determine the maximum possible value of $$ \ sum_ { i \ in \ mathbb { Z }:1\ leqslant i \ leqslant 2019} f ( i ) . $$ Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 673.

Table 26: Details of example π961 . Prompt: A herder has forgotten the number of cows she has and does not want to count all of them . She remembers these four facts about the number of cows :\ n \n - It has $3$ digits .\ n - It is a palindrome .\ n - The middle digit is a multiple of $4$ .\ n - It is divisible by $11$ .\ n \ nWhat is the sum of all possible numbers of cows that the herder has ?\ n \ n$ \\ textbf {( A ) }343 \\ \\ textbf {( B ) }494 \\ \\ textbf {( C ) }615 \\ \\ textbf {( D ) }635 \\ \\ textbf {( E ) }726 $ Please reason step by step , and put your final answer within \\ boxed {}.

Ground truth: 6.

Table 27: Details of example π973 . Prompt: Find all ordered pairs $ (a , b ) $ of positive integers for which \ n$$ \ n \ frac {1}{ a }+\ frac {1}{ b }=\ frac {3}{2018} .\ n$$ \ nPlease provide the sum of all integers in the ordered pairs . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 1438383.

Table 28: Details of example π987 . Prompt: Find the smallest positive integer $n$ with the following property : for every sequence of positive integers $a_1 , a_2 ,\ ldots , a_n$ with $a_1 + a_2 +\ ldots + a_n =2013 $ , there exist some ( possibly one ) consecutive term ( s ) in the sequence that add up to $70$ . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 1033.

18

Table 29: Details of example π997 . Prompt: A quadratic polynomial $f ( x ) $ is called sparse if its degree is exactly 2 , if it has integer coefficients , and if there exists a nonzero polynomial $g ( x ) $ with integer coefficients such that $f ( x ) g ( x ) $ has degree at most 3 and $f ( x ) g ( x ) $ has at most two nonzero coefficients . Find the number of sparse quadratics whose coefficients lie between 0 and 10 , inclusive . Please reason step by step , and put your final answer within \ boxed {}.

Ground truth: 228.

19

Record · ID 660869 · SHA-256 14dc85ffeb179850
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.