OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Xinyu Ma1,*, † , Mingzhou Xu2,* , Xuebo Liu1,† , Chang Jin2 , Qiang Wang2 , Derek F. Wong3 , Min Zhang1 1
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China 2 Hithink RoyalFlush Information Network, Hangzhou, China 3 Computer and Information Science, University of Macau, Macau, China Q : [email protected], [email protected]
Abstract
et al., 2025; Wang et al., 2025b; Deng et al., 2026) have focused on augmenting model performance by refining the Group Relative Policy Optimization (GRPO) (Shao et al., 2024) algorithm. Despite these strides, recent findings (Yue et al., 2025; Nguyen et al., 2025) consistently highlight a critical bottleneck: while RLVR significantly elevates the performance of the foundational model, it often fails to foster truly novel problem-solving capabilities. Zhao et al. (2025) identify an “echo chamber” effect, wherein reinforcement learning (RL) converges toward dominant pre-existing distributions. Similarly, Nguyen et al. (2025); Yue et al. (2025) argue that self-sampling primarily reinforces known correct patterns rather than catalyzing the discovery of original reasoning trajectories. To address these limitations, research has branched into two primary directions: incorporating offline guidance and applying entropy-based regularization. On one hand, frameworks such as Luffy (Yan et al., 2025) and Chord (Zhang et al., 2025a) leverage high-quality teacher trajectories to provide gold-standard demonstrations, facilitating a rapid transition from supervised imitation to autonomous reasoning. On the other hand, entropydriven approaches (Cui et al., 2025b; Wang et al., 2025c) mitigate “entropy collapse” by maintaining the model within a high-entropy regime. By sustaining diverse exploration and preventing premature convergence to suboptimal, these methods ensure the continued discovery of effective paths. However, while external teacher data provides an intuitive performance boost, existing methods often lack a seamless integration between offline and online components. Furthermore, entropy-driven strategies remain fundamentally constrained by the model’s inherent capacity. To bridge this gap, we propose OGER, a framework that enhances reasoning capabilities by unifying offline teacher guidance and online RL through a specialized reward modeling lens. Specifically, our framework lever-
arXiv:2604.18530v1 [cs.AI] 20 Apr 2026
Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore novel trajectories beyond their initial latent space. While offline teacher guidance and entropy-driven strategies have been proposed to address this, they often lack deep integration or are constrained by the model’s inherent capacity. In this paper, we propose OGER, a novel framework that unifies offline teacher guidance and online reinforcement learning through a specialized reward modeling lens. OGER employs multi-teacher collaborative training and constructs an auxiliary exploration reward that leverages both offline trajectories and the model’s own entropy to incentivize autonomous exploration. Extensive experiments across mathematical and general reasoning benchmarks demonstrate that OGER significantly outperforms competitive baselines, achieving substantial gains in mathematical reasoning while maintaining robust generalization to out-of-domain tasks. We provide a comprehensive analysis of training dynamics and conduct detailed ablation studies to validate the effectiveness of our entropy-aware reward modulation. Our code is available at https: //github.com/ecoli-hit/OGER.git.
1
Introduction
Current Large Language Models (LLMs) have demonstrated increasingly sophisticated reasoning capabilities (Liu et al., 2024; Yang et al., 2025a). This progress is primarily catalyzed by Reinforcement Learning with Verifiable Rewards (RLVR), which fosters the emergence of extended chains of thought and strengthens systemic reasoning. Following the paradigm established by DeepSeek-R1 (Guo et al., 2025), subsequent advancements (Du et al., 2025; Feng et al., 2025; Liu *: Equal contribution. †: Correspondence.
1
ages multi-source teacher trajectories for collaborative training to solidify foundational reasoning capabilities. Concurrently, we construct a divergencebased exploration reward to quantify the semantic disparity between online and offline trajectories, facilitating a profound synergy between expert imitation and autonomous discovery. To ensure training stability, we implement a hybrid sampling mechanism that integrates offline expert data directly into the online training batches. Furthermore, we refine this exploration signal by incorporating the policy model’s token-level entropy distribution, enabling fine-grained control to incentivize novel reasoning behaviors and mitigate the risk of premature convergence. Our main contributions are as follows:
More recently, advancements have focused on the adaptive mixing of offline teacher data with online samples, utilizing importance sampling and sophisticated weighting schemes to shift the model’s focus during training (Yan et al., 2025; Zhang et al., 2025a; Bartoldson et al., 2025). While our framework, OGER, aligns with the paradigm of hybrid data utilization, it departs from conventional data-level blending or objectivefunction fusion. Instead, we unlock the latent potential of offline data by employing it as a foundational reference for an auxiliary exploration reward. This mechanism provides explicit signal guidance, incentivizing the model to explore a broader search space beyond the reach of standard on-policy RL.
• We propose OGER, a sophisticated RLVR framework that integrates multi-teacher offline trajectories into RL. By introducing a novel offline-guided exploration reward and leveraging entropy for reward shaping, OGER effectively harmonizes offline knowledge with online discovery.
2.2
Despite the success of RLVR, recent studies (Yue et al., 2025; Karan and Du, 2025; Xie et al., 2025; Ke et al., 2025; Deng et al., 2025) suggest that RL primarily amplifies pre-existing capabilities within the base model rather than inducing entirely novel reasoning paradigms. Cui et al. (2025b) analyzes this phenomenon through the lens of information theory, identifying the challenge of entropy collapse: without robust regularization, policy entropy diminishes precipitously during early training, leading the model to converge prematurely to a narrow subset of high-reward trajectories. Consequently, sustaining a relatively high-entropy regime is recognized as a pivotal factor in preventing entropy collapse and facilitating more profound reasoning discovery. Current research has explored various entropycentric methodologies to revitalize training. For instance, Wang et al. (2025c) demonstrates that selectively training on the top 20% of tokens with the highest entropy outperforms standard full-token fine-tuning. Cheng et al. (2025) introduces entropybased advantage shaping to bolster exploratory reasoning, while Zhang et al. (2025b) proposes an adaptive regularization framework to mitigate entropy decay throughout the training lifecycle. Furthermore, recent works (Su et al., 2025; Wang et al., 2025a) leverage entropy for granular analysis of model generations to achieve a more nuanced balance between exploration and exploitation. Building on these insights, OGER adopts entropy as a fundamental metric for reward shaping. By incorporating self-entropy into the reward modeling process, we exert fine-grained control over exploratory behavior, ensuring the model remains
• We demonstrate that OGER functions as a high-quality reward mechanism, achieving significant performance gains over competitive baselines, including 4% and 7.9% improvement on average of mathematics and general evaluation with 1.5B and 7B models. • We provide in-depth analysis of our OGER framework’s efficacy, analyzing the evolution of scores and entropy during training, and conducting comprehensive ablation studies of its modules to validate individual contributions.
2
Related Work
2.1
Off-policy Guided Reinforcement Learning
The Entropy Mechanism in RLVR
Integrating off-policy data into RLVR has emerged as a promising frontier for augmenting LLMs’ reasoning capabilities, leading to several distinct research trajectories. One approach emphasizes the collection, filtration, and replay of historical trajectories generated by the online model (Zhan et al., 2025; Dou et al., 2025), aiming to maximize sample efficiency and exploit the model’s past experiences. Another line of inquiry investigates the fusion of Supervised Fine-Tuning (SFT) and RL, where the relative weighting of losses is dynamically adjusted to balance imitation and exploration (Lanchantin et al., 2025; Lv et al., 2025). 2
robust against premature convergence.
training.
3
Embedding Trajectories for Comparison During the training procedure, for a given query q, let on } denote a set of N trajecTon = {τ1on , . . . , τN tories generated by the online policy model, and of f Tof f = {τiof f , . . . , τM } represent M offline trajectories curated from teacher models. Rather than treating offline trajectories as static replacements for online batches, we utilize them as a semantic reference to steer the online exploration process. We first embed both Tof f and Ton into a shared d-dimensional latent space. For the i-th trajectory in Ton and the j-th trajectory in Tof f , the latent representations are defined as:
OGER: Offline-Guided Exploration Reward for Hybrid RL
Current research underscores the critical need for autonomous exploration to enhance LLM reasoning, particularly through hybrid training paradigms and entropy-based regularization. In this work, we introduce OGER, a framework that fosters synergistic integration of offline expert guidance and online exploratory discovery. Our approach departs from traditional data-mixing strategies by shifting the fusion of loss to the reward modeling level. Specifically, OGER constructs an auxiliary exploration reward that benchmarks online trajectories against an ensemble of high-quality multi-teacher offline trajectories by divergence. This provides the model with an explicit directional signal to navigate the complex reasoning search space. To ensure exploration is both diverse and well-calibrated, this guidance is further refined by a fine-grained mechanism that uses the model’s internal entropy of the distributions at the last token of the trajectory to shape the exploration reward. By unifying these dimensions, OGER effectively mitigates the risk of entropy collapse while transcending the limitations of passive imitation, thereby facilitating the discovery of novel reasoning paths. The comprehensive architecture of our method is illustrated in fig. 1. 3.1
Eion = Enc(τion ),
Ejof f = Enc(τjof f ) ∈ Rd (1)
Exploration Reward for On-policy Trajectories In standard RLVR frameworks, the verifiable reward function is typically defined as follows: for a reward function R(x, y) ∈ [0, 1], x denotes the input query and y represents the response generated by policy model πθ . Current research predominantly assigns rewards via binary classification based on final answer correctness (Zhan et al., 2025; Zheng et al., 2025). The primary objective of RLVR training is to maximize the expected reward, formally expressed as: LRLVR (θ) = Ex∼D, y∼πθ (·|x) [ R(x, y) ]
Offline-Guided On-policy Exploration Reward
(2)
where θ denotes the model parameters and D represents the training distribution. However, existing methods typically assign uniform rewards to all correct trajectories, which inadvertently restricts the model’s ability to selectively learn from a diverse manifold of reasoning paths. To address this limitation, we construct an exploration reward that distinguishes reward signals across different correct trajectories. Building upon the latent representations Eion and Ejof f defined in eq. (1), we then compute the cosine similarity between the online and offline trajectories: si,j = Cosine(Eion , Ejof f )
Ensembling Offline Teacher Demonstrations A core component of the OGER framework is the construction of a high-quality dataset synthesized from multiple state-of-the-art teacher models. By generating supplementary trajectories across multiple teacher models, we capture a rich manifold of reasoning structures, particularly the long-form Chain-of-Thought (CoT) patterns that characterize advanced systemic reasoning. This multi-source synthesis yields a robust ensemble of standard trajectories, providing a high-fidelity offline reference for our reward-level fusion. To ensure both data integrity and computational tractability, we implement two rigorous filtering constraints: (1) we apply a length-consistency filter to ensure all teacher trajectories align with a uniform maximum reasoning length, and (2) we employ a correctness filter for each trajectory, retaining only those with a correct final answer for
=
Eion · Ejof f
(3)
∥Eion ∥2 ∥Ejof f ∥2
To quantify the proximity of each online sample to the offline distribution, we calculate the mean 3
§2.1: Offline Guided On-policy Exploration Reward
§2.2 & 2.3: OGER Reward with Hybrid Set Upon GRPO
Divergence-Aware Replacement
Query
Offline Teacher Sampling
Online Trajectory
Offline Trajectory
Random Trajectory
Online Policy Sampling
Hybrid Trajectory Set
···
Teacher Models
Lowest Divergence
Policy Model
Filtering (Length & Correctness) Offline Trajectories
Gated Mechanism Based on Source
Online Trajectories Offline Subset(
Verifiable Reward
Encoder(Enc): Latent Embeddings
Total Reward
Divergence As Foundational Exploration Reward
Similarity
Divergence
)
Online Subset(
Verifiable Reward
Last Token Entropy
Divergence
OGER Exploration Reward
Total Reward
Group Advantage Estimation And Optimization
Mean
)
Figure 1: The overall architecture of the OGER framework. We first construct a comprehensive, high-quality offline dataset from trajectories generated by multiple teacher models and perform hybrid offline-online integration by mapping online sampled trajectories and offline references into a shared representation space to compute divergence. This divergence serves as the basis for our exploration reward and offline replacement. Furthermore, we leverage the Shannon entropy of the last token to refine the exploration reward, providing a fine-grained, uncertainty-aware modulation of the training signal.
meaningful discovery while suppressing erratic or unconstrained divergence. Specifically, we incorporate the Shannon entropy of the last token in τion as a proxy for policy confidence to refine the foundational exploration reward Di from eq. (5). Let V denote the vocabulary of the policy model and p(v) represent the probability distribution over tokens v ∈ V. The entropy of the distribution at the final token position of trajectory τion is defined as: X Hilast = − p(v) log p(v) (6)
similarity for each online trajectory: M
simi =
1 X si,j M
(4)
j=1
This similarity score reflects the degree to which the online model replicates the teacher’s reasoning patterns; a higher similarity score indicates that the model has effectively converged toward the offline distribution through imitation. To incentivize the model to venture beyond these established patterns and explore a broader reasoning manifold, we utilize semantic divergence as the foundational exploration reward Di for trajectory τion : Di = 1 − simi 3.2
v∈V
By integrating the fundational exploration reward Di with the last-token entropy Hilast , the refined exploration reward for trajectory τion is formulated as:
(5)
RiOGER = Di · exp(−Hilast ) · Rim M X 1 = 1 − si,j · exp(−Hilast ) · Rim M
OGER Reward with Hybrid Set
Entropy-aware Reward Refinement In RLVR, rewards are typically sparse and anchored to the objective correctness of the final answer. However, evaluating the validity of the intermediate reasoning process remains a significant challenge, often leading to reward hacking (Skalse et al., 2022), where the model arrives at a correct solution via logically flawed trajectories. Drawing on the intuition that a model’s entropy reflects its inherent aleatoric uncertainty (Yang et al., 2025b), we leverage this signal to modulate the exploration process. This ensures the model is incentivized toward
j=1
(7) where Rim ∈ {0, 1} denotes the standard verifiable reward. This formulation ensures that RiOGER modulates the reward exclusively for correct online trajectories Divergence-Aware Trajectory Replacement In the joint training procedure, for each problem query q, we integrate the online and offline distributions by replacing the trajectory in Ton that exhibits the 4
Model
Valid Samples
Avg. Length
Accuracy (%)
R1 Qwen GLM
45,462 36,958 17,887
4,021.14 5,252.09 10,318.07
99.28 94.90 82.14
Table 1: Statistics of the curated offline dataset generated by multiple teacher models. The Valid Samples refer to trajectories that passed the correctness verification and length constraints. Figure 2: Distribution of sequence lengths for trajectories generated by multiple teacher models. We filter out samples with more than 8k tokens. R1 denotes DeepSeek-R1, Qwen means Qwem3-32B, and GLM represents GLM-4.5 Air.
lowest divergence with a randomly sampled trajectory from Tof f . This results in a hybrid trajectory set Thyb comprising N samples. Notably, we apply the proposed exploration reward exclusively to on-policy trajectories. For offline teacher trajectories, we utilize only the standard verifiable reward Rm . This categorical separation ensures that exploration signals—which are intrinsically coupled to the model’s internal uncertainty—do not introduce noise into the standard teacher demonstrations. The final reward used to compute group advantages is defined as a gated composition of verifiable and exploratory signals. For any trajectory τi ∈ Thyb , the total reward Ritotal is formulated as: Ritotal =
( Rim + RiOGER , Rim ,
if τi ∈ Ton if τi ∈ Tof f
parameters θ is formulated as: 1
JGRPO (θ) = PN
i=1 |τi | i=1 t=1
(10) where ri,t (θ) = πθ (τi,t | q, τi,<t )/πθold (τi,t | q, τi,<t ) represents the importance sampling ratio. We follow the advantage-shaping methodology proposed by Yan et al. (2025) to optimize the importance sampling term. The Kullback-Leibler (KL) divergence term DKL [πθ ||πref ], scaled by the hyperparameter β, traditionally regularizes the policy against a reference model πref to mitigate catastrophic forgetting or distribution collapse. However, consistent with recent advancements, which suggest that the KL penalty may become redundant when employing robust advantage normalization and intrinsic exploration rewards (Hu et al., 2025; Yu et al., 2025; Cui et al., 2025a), we omit the KL term in our primary training objective.
(8)
Implementation with GRPO
In the GRPO framework, πθold denotes the policy model used for trajectory sampling, while πθ represents the policy undergoing optimization. Each trajectory in Thyb is evaluated via eq. (8) to further determine its relative advantage within the group. For a specific trajectory τi ∈ Thyb , the group-relative advantage Ai is computed as:
Ai =
Ritotal − mean({Rjtotal |j = 1, · · · , N }) std({Rjtotal |j = 1, · · · , N })
CLIP(ri,t (θ), Ai , ϵ)
− β · DKL [πθ ||πref ]
In summary, within the OGER framework, the auxiliary exploration reward is exclusively assigned to correct trajectories generated by the online policy. For offline teacher demonstrations and incorrect online samples, the reward remains aligned with the standard verifiable baseline. 3.3
|τi | N X X
4
Experiments
4.1
Datasets
Training Dataset We utilize a subset of the OpenR1-Math-220k corpus, curated and filtered by Yan et al. (2025), as our foundational dataset. This base comprises 45k instances, each paired with a reasoning trajectory generated by DeepSeekR1 for tasks sourced from NuminaMath 1.5 (LI et al., 2024). Building upon this foundation, we construct a heterogeneous multi-teacher ensemble as detailed
(9)
https://huggingface.co/datasets/open-r1/OpenR1-Math220k
The GRPO objective function for optimizing the 5
in section 3.1, incorporating additional trajectories distilled from Qwen3-32B (Yang et al., 2025a) and GLM-4.5 Air (Team et al., 2025). We rigorously validate these samples for correctness using MathVerify and impose a maximum sequence length constraint of 8k tokens. The statistical distribution and structural characteristics of the resulting dataset are illustrated in fig. 2 and table 1. While OGER leverages the full multi-teacher ensemble, we restrict the Luffy baseline (Yan et al., 2025) to DeepSeek-R1 samples only; this ensures strict adherence to its original configuration and facilitates a fair comparison against its reported benchmarks.
using Math-Verify. Following the advantage shaping mechanism proposed by Yan et al. (2025), we apply token-level reward signals to enhance training stability and credit assignment. All primary experiments are conducted using Qwen2.5-Math7B and Qwen2.5-Math-1.5B model (Yang et al., 2024). For embedding, we employ the bge-largeen-v1.5 (Xiao et al., 2023) model via the FlagEmbedding framework. Baselines To rigorously evaluate the efficacy of OGER, we compare it against the following baselines: (1) Base: evaluates the zero-shot performance of the pre-trained backbone model without further fine-tuning, serving as the performance floor. (2) GRPO: the standard on-policy training paradigm using the GRPO algorithm (Shao et al., 2024) applied to our primary dataset, representing a purely online reinforcement learning approach. (3) Luffy: a competitive hybrid baseline that incorporates DeepSeek-R1 trajectories during training with token-level advantage shaping. Detailed experimental configurations and hyperparameter settings can be found in Appendix A.
Evaluation Benchmarks To assess the robustness and generalization of our framework, we evaluate all models across six prominent benchmarks for mathematical reasoning and three diverse outof-domain (OOD) tasks. The tasks we use are as follows: (1) for Mathematical Reasoning: we conduct evaluations on AIME 2024, AIME 2025, AMC (Li et al., 2024), Minerva (Lewkowycz et al., 2022), OlympiadBench (Olympaid) (He et al., 2024), and MATH-500 (Hendrycks et al., 2021). For benchmarks with limited test samples (AIME 2024, AIME 2025, and AMC), we report the pass@1 score at 32 rollouts to mitigate variance and better capture the model’s reasoning consistency. For the remaining benchmarks, we report pass@1 score at rollout 1. (2) for Out-of-Domain Generalization: to evaluate systemic reasoning capacity beyond pure mathematics, we include ARC-Challenge (Clark et al., 2018), GPQA-Diamond (Rein et al., 2024), and MMLUPro (Wang et al., 2024), and report the average score (OOD) of the pass@1 score at rollout 1. We evaluate all tasks at sampling temperature of 0.8 to balance generation diversity and precision. 4.2
4.3
Main Results
The primary experimental results are summarized in Table table 2. Overall, our proposed framework, OGER, achieves substantial performance gains across both the 1.5B and 7B model scales, reaching 36.77 on the 1.5B model and 52.03 on the 7B model. Compared to the vanilla GRPO baseline, OGER yields remarkable relative improvements of 28.2% and 33.0%, respectively. When compared with the stronger Luffy baseline, OGER achieves absolute score increases of 1.53 and 3.36 points, respectively. These results underscore the scalability of OGER, confirming its effectiveness in enhancing the mathematical reasoning capabilities of both small-scale and large-scale language models. Notably, the performance gains are more pronounced as model size increases, with the 7B variant outperforming all baselines across all evaluation tasks. Specifically, on the challenging AIME 2024 and AIME 2025 benchmarks, OGER 7B achieves 31.77 and 25.10, representing a significant leap over Luffy’s 26.67 and 21.04. For the 1.5B model, OGER also demonstrates significant progress on highly complex reasoning benchmarks; for instance, it achieves 13.54 on AIME 2025, outperforming Luffy by nearly 40%.
Implementation Details
For the training phase, we omit the KL divergence penalty and fix the entropy loss coefficient at 0.01. Training is conducted using a global batch size of 128 and a micro-batch size of 64. For each query, we perform online sampling of N = 8 trajectories with a temperature of 1.0. Within the OGER framework, we implement the hybrid update strategy described in section 3.2, in which each online trajectory per query is replaced by a high-quality offline trajectory from our teacher ensemble. The logical correctness of all trajectories is rigorously validated https://github.com/huggingface/Math-Verify
https://github.com/FlagOpen/FlagEmbedding/tree/master
6
Method
AIME 2024
AIME 2025
AMC
MATH-500
Base
4.17 (-5.21)
0.94 (-6.25)
22.29 (-19.39)
29.80 (-40.00)
GRPO
9.38 (0.00)
7.19 (0.00)
41.68 (0.00)
Luffy
14.37 (+4.99)
9.69 (+2.50)
OGER
14.27 (+4.89)
Base
Minerva
Olympiad
OOD
AVG.
5.51 (-20.96)
13.63 (-17.33)
21.32 (+5.96)
13.95 (-14.74)
69.80 (0.00)
26.47 (0.00)
30.96 (0.00)
15.37 (0.00)
28.69 (0.00)
47.33 (+5.65)
75.20 (+5.40)
27.21 (+0.74)
39.85 (+8.89)
33.07 (+17.70)
35.25 (+6.55)
13.54 (+6.35)
46.76 (+5.08)
76.60 (+6.80)
29.41 (+2.94)
39.11 (+8.15)
37.70 (+22.33)
36.77 (+8.08)
13.65 (-3.75)
5.63 (-5.52)
44.92 (-9.82)
65.00 (-15.20)
14.71 (-25.00)
30.07 (-9.34)
26.28 (-4.93)
28.61 (-10.51)
GRPO
17.40 (0.00)
11.15 (0.00)
54.74 (0.00)
80.20 (0.00)
39.71 (0.00)
39.41 (0.00)
31.21 (0.00)
39.12 (0.00)
Luffy
26.67 (+9.27)
21.04 (+9.89)
63.59 (+8.85)
87.00 (+6.80)
42.28 (+2.57)
48.74 (+9.33)
51.33 (+20.12)
48.66 (+9.55)
OGER
31.77 (+14.37)
25.10 (+13.95)
68.64 (+13.90)
88.40 (+8.20)
45.22 (+5.51)
53.48 (+14.07)
51.61 (+20.40)
52.03 (+12.91)
Qwen2.5-Math-1.5B
Qwen2.5-Math-7B
Table 2: Comprehensive performance comparison of OGER against other methods using the Qwen2.5-Math-1.5B and Qwen2.5-Math-7B backbones. Base denotes the performance of the original foundational models, while GRPO and Luffy serve as the primary reinforcement learning baselines. Values in parentheses denote the absolute performance gain or loss relative to the GRPO baseline. We highlight the best and second-best results for each model scale using bold and underlined text, respectively.
Furthermore, for OOD tasks, the 1.5B configuration achieves 37.70, a substantial 4.63 points lead over Luffy (33.07). This trend is maintained in the 7B model, where OGER hits a peak of 51.61. These results indicate that our entropy-refined exploration reward does not merely facilitate the memorization of teacher distributions but actively incentivizes the model to discover more robust and transferable reasoning paths. Even in specialized domains such as MATH-500 and Minerva, OGER consistently outperforms across both model sizes, validating the universality of our proposed exploration mechanism.
5
Analysis
5.1
Ablation Studies
w/o Refinement achieves a 2.24 point improvement over the Luffy baseline, it still falls 1.13 points short of the full OGER model on average. Furthermore, when the exploration reward is entirely removed (OGER w/o Reward), and only offline trajectory replacement is used, the model yields only marginal gains over Luffy and remains significantly inferior to the full OGER configuration. These results underscore that while offline data provides a valuable semantic reference, the entropymodulated exploration reward is the primary catalyst for superior reasoning performance and training stability. Crucially, our findings suggest that merely increasing the diversity of offline trajectories fails to yield substantial gains. This indicates that static imitation alone is insufficient and the dynamic interaction between offline alignment and entropy-driven discovery is essential.
Effectiveness of Reward Components To rigorously evaluate the contribution of each proposed component, we analyze two variants of our framework: (1) OGER w/o Refinement: this variant excludes the entropy-aware refinement mechanism, relying solely on the foundational exploration reward. (2) OGER w/o Reward: this configuration utilizes the standard verifiable reward and offline data integration without exploration reward, serving as a baseline to validate the necessity of the exploration reward. As summarized in table 3, removing the entropyaware refinement results in consistent performance degradation across all benchmarks relative to the full OGER framework. Specifically, while OGER
Impact of Offline Replacement Density To investigate the influence of offline teacher participation during training, we conducted another ablation study on the replacement strategy. In our primary OGER configuration, we implement a singlesample replacement per query group (N = 8). To evaluate the model’s sensitivity to offline guidance intensity, we compare this against OGER w/ 2Replace and OGER w/ 3-Replace, which substitute two and three online trajectories per batch, respectively. As illustrated in table 3, while more replacements achieve results slightly above the Luffy base7
Method
AIME 2024
AIME 2025
AMC
MATH-500
Minerva
Olympiad
OOD
AVG.
OGER Variants OGER w/o Refinement
30.31 (-1.46)
23.75 (-1.35)
66.42 (-2.22)
88.40 (0.00)
41.18 (-4.04)
53.04 (-0.44)
53.22 (+1.61)
50.90 (-1.13)
OGER w/o Reward
27.08 (-4.69)
23.13 (-1.97)
66.64 (-2.00)
88.40 (0.00)
36.03 (-9.19)
50.81 (-2.67)
51.19 (-0.41)
49.04 (-2.99)
OGER w/ 2-Replace
27.92 (-3.85)
25.21 (+0.11)
67.09 (-1.55)
87.40 (-1.00)
38.60 (-6.62)
52.89 (-0.59)
51.66 (+0.05)
50.11 (-1.92)
OGER w/ 3-Replace
28.75 (-3.02)
24.58 (-0.52)
64.65 (-3.99)
88.20 (-0.20)
38.60 (-6.62)
51.11 (-2.37)
49.72 (-1.88)
49.37 (-2.66)
OGER with Different Replace Density
Table 3: Ablation studies of the OGER framework on mathematical reasoning benchmarks. OGER w/o Refinement and OGER w/o Reward denote variants where the entropy-aware refinement and the exploration reward are, respectively, removed. We further evaluate OGER under varying replacement densities (OGER w/ 2-Replace and OGER w/ 3-Replace). Values in parentheses denote the absolute performance gain or loss relative to the OGER.
line, they exhibit a significant performance gap compared to the standard OGER. This disparity underscores that, in the hybrid paradigm of offline and online learning, the density of offline guidance must be carefully calibrated. Excessive reliance on static expert trajectories can inadvertently stifle the model’s autonomous exploratory potential, whereas OGER effectively bridges this gap by incentivizing novel reasoning paths that transcend mere imitation of the teacher distribution. 5.2
Method
GPU Hours
GRPO Luffy OGER
75*8 120*8 168*8
Data Usage Online Offline 45K*8 45K*7 45K*7
0K 45K 128K
Table 4: Comparison of resource requirements between OGER and other methods.
maintains continuous exploration throughout the learning process. Furthermore, fig. 3c shows that our method exhibits a markedly lower failure rate as training progresses. This decline suggests that OGER effectively navigates the reasoning manifold, translating exploratory signals into tangible improvements in problem-solving coverage. Regarding computational overhead, we summarize resource utilization in table 4. While OGER requires additional training time per step due to exploration reward computations and the overhead of maintaining higher sample diversity, this investment directly translates into substantial gains in final model accuracy and robustness. The higher compute cost is thus justified by the model’s ability to transcend the imitation bottleneck and achieve better performance.
Training Dynamics
To gain deeper insights into the underlying optimization process, we provide an extensive analysis of the OGER framework utilizing the Qwen2.5Math-7B model. Specifically, we examine the evolution of policy entropy, mean test accuracy, and the proportion of failed responses in the training batch. These metrics are visualized in fig. 3a, fig. 3b, and fig. 3c, respectively. Additionally, training resource consumption is detailed in table 4. As illustrated in fig. 3a, both OGER and its variant OGER w/o Refinement maintain significantly higher entropy levels throughout the training lifecycle compared to the Luffy and GRPO. While standard on-policy RL often suffers from rapid entropy collapse—where the policy prematurely concentrates on a narrow set of high-reward trajectories—our framework effectively preserves a broader exploratory landscape. This sustained entropy correlates with the performance trends in fig. 3b, where OGER demonstrates superior training stability. Unlike the baselines, which exhibit performance plateaus or degradation as the policy becomes overly confident, OGER facilitates longer optimization cycles by consistently discovering novel reasoning paths. The detailed analysis of our OGER reward during training, as provided in Appendix B, further demonstrates that our method
5.3
Coverage of the Reasoning Manifold
To evaluate the breadth of the reasoning space the model has mastered, we assess pass@k performance for OGER and the selected baselines on the AIME 2024 and AIME 2025 benchmarks. As illustrated in fig. 4, we generated 256 rollout samples per problem to compare solvability coverage across varying computational budgets. Our proposed OGER framework, along with its variants, consistently outperforms both Luffy and GRPO across the entire range of k. 8
(a) Entropy
(b) Average Score
(c) Failed Response Ratio
Figure 3: Comparative analysis of training dynamics across OGER, its variant OGER w/o Refinement, and baselines (GRPO and Luffy). Left: evolution of policy entropy throughout the training process. Middle: progression of average accuracy across six major mathematical benchmarks. Right: the failure ratio within training responses (GRPO’s data is not available).
hancing the model’s intrinsic reasoning capabilities. The details could be found in Appendix C.
6
Conclusion
We propose OGER, a simple yet powerful hybrid training framework that integrates offline guidance with online entropy to build auxiliary exploration rewards. Experimental results across multiple benchmarks demonstrate that OGER significantly enhances mathematical reasoning performance and exhibits superior generalization to outof-domain tasks, achieving 4% and 7.9% improvements on average for both small- and large-scale models compared to a competitive baseline. These findings underscore the feasibility and immense potential of entropy, alongside offline demonstrations, to guide model capability growth. Future work will focus on extending the OGER framework to more complex scenarios.
Figure 4: pass@k performance on AIME 2024 and AIME 2025 using 256 rollouts. Our proposed OGER method consistently outperforms all baselines across various k values, demonstrating a significantly higher convergence rate in solvability coverage.
These results demonstrate that our entropyaware reward refinement not only enhances the precision of individual reasoning trajectories but also significantly expands the aggregate coverage of solvable problems. By maintaining a high-entropy exploratory regime, OGER ensures that the model navigates a more diverse and robust solution manifold. This broader coverage indicates that our method successfully facilitates the discovery of alternative reasoning paths without compromising the model’s fundamental generalization capacity.
References Brian Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan Obando-Ceron, Yoshua Bengio, and Bhavya Kailkhura. 2025. Trajectory balance with asynchrony: Decoupling exploration and learning for fast, scalable llm post-training. arXiv preprint arXiv:2503.18929. Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai, Wayne Xin Zhao, Zhenliang Zhang, and Furu Wei. 2025. Reasoning with exploration: An entropy perspective. arXiv preprint arXiv:2506.14758.
Furthermore, we conducted extensive experiments across a range of decoding temperatures to evaluate the model’s inference-time robustness. As illustrated by the results, our proposed OGER consistently outperforms both Luffy and GRPO across the entire temperature spectrum, which further substantiates the effectiveness of our approach in en-
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
9
Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, Jiacheng Chen, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, and 1 others. 2025a. Process reinforcement through implicit rewards. arXiv preprint arXiv:2502.01456.
on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual. Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, Xiangyu Zhang, and Heung-Yeung Shum. 2025. Openreasoner-zero: An open source approach to scaling up reinforcement learning on the base model. arXiv preprint arXiv:2503.24290.
Ganqu Cui, Yuchen Zhang, Jiacheng Chen, Lifan Yuan, Zhi Wang, Yuxin Zuo, Haozhan Li, Yuchen Fan, Huayu Chen, Weize Chen, and 1 others. 2025b. The entropy mechanism of reinforcement learning for reasoning language models. arXiv preprint arXiv:2505.22617.
Aayush Karan and Yilun Du. 2025. Reasoning with sampling: Your base model is smarter than you think. arXiv preprint arXiv:2510.14901.
Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Jing Li, Min Zhang, and Zhaopeng Tu. 2025. DRPruning: Efficient large language model pruning through distributionally robust optimization. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 29152–29173, Vienna, Austria. Association for Computational Linguistics.
Xiaopeng Ke, Hexuan Deng, Xuebo Liu, Jun Rao, Zhenxi Song, Jun Yu, and Min Zhang. 2025. AQuilt: Weaving logic and self-inspection into low-cost, highrelevance data synthesis for specialist LLMs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 5752– 5785, Suzhou, China. Association for Computational Linguistics.
Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Jun Rao, and Min Zhang. 2026. REA-RL: Reflection-aware online reinforcement learning for efficient reasoning. In The Fourteenth International Conference on Learning Representations.
Jack Lanchantin, Angelica Chen, Janice Lan, Xian Li, Swarnadeep Saha, Tianlu Wang, Jing Xu, Ping Yu, Weizhe Yuan, Jason E Weston, and 1 others. 2025. Bridging offline and online reinforcement learning for llms. arXiv preprint arXiv:2506.21495.
Shihan Dou, Muling Wu, Jingwen Xu, Rui Zheng, Tao Gui, Qi Zhang, and Xuanjing Huang. 2025. Improving rl exploration for llm reasoning through retrospective replay. arXiv preprint arXiv:2504.14363.
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving quantitative reasoning problems with language models. In Advances in Neural Information Processing Systems, volume 35, pages 3843–3857. Curran Associates, Inc.
He Du, Bowen Li, Chengxing Xie, Chang Gao, Kai Chen, and Dacheng Tao. 2025. Confidence as a reward: Transforming llms into reward models. arXiv preprint arXiv:2510.13501. Yunzhen Feng, Parag Jain, Anthony Hartshorn, Yaqi Duan, and Julia Kempe. 2025. Don’t waste mistakes: Leveraging negative rl-groups via confidence reweighting. arXiv preprint arXiv:2510.08696.
Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, and 1 others. 2024. Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions. Hugging Face repository, 13(9):9.
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, and 1 others. 2024. Numinamath. [https://huggingface. co/AI-MO/NuminaMath-1.5](https://github.c om/project-numina/aimo-progress-prize/blo b/main/report/numina_dataset.pdf).
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, and 1 others. 2024. OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3828–3850, Bangkok, Thailand. Association for Computational Linguistics.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, and 1 others. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437. Runze Liu, Jiakang Wang, Yuling Shi, Zhihui Xie, Chenxin An, Kaiyan Zhang, Jian Zhao, Xiaodong Gu, Lei Lin, Wenping Hu, and 1 others. 2025. Attention as a compass: Efficient exploration for processsupervised rl in reasoning models. arXiv preprint arXiv:2509.26628.
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track
10
Xingtai Lv, Yuxin Zuo, Youbang Sun, Hongyi Liu, Yuntian Wei, Zhekai Chen, Lixuan He, Xuekai Zhu, Kaiyan Zhang, Bingning Wang, and 1 others. 2025. Towards a unified view of large language model posttraining. arXiv preprint arXiv:2509.04419.
2024. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024.
Phuc Minh Nguyen, Chinh D. La, Duy M. H. Nguyen, Nitesh V. Chawla, Binh T. Nguyen, and Khoa D. Doan. 2025. The reasoning boundary paradox: How reinforcement learning constrains language models. arXiv preprint arXiv:2510.02230.
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-pack: Packaged resources to advance general chinese embedding. arXiv preprint arXiv:2309.07597.
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2024. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling.
Can Xie, Ruotong Pan, Xiangyu Wu, Yunfei Zhang, Jiayi Fu, Tingting Gao, and Guorui Zhou. 2025. Unlocking exploration in rlvr: Uncertainty-aware advantage shaping for deeper reasoning. arXiv preprint arXiv:2510.10649.
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300.
Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang, Ganqu Cui, Xiaoye Qu, Yu Cheng, and Yue Zhang. 2025. Learning to reason under off-policy guidance. arXiv preprint arXiv:2504.14945. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025a. Qwen3 technical report. arXiv preprint arXiv:2505.09388.
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. Defining and characterizing reward gaming. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, and 1 others. 2024. Qwen2.5-math technical report: Toward mathematical expert model via self-improvement. arXiv preprint arXiv:2409.12122.
Zhenpeng Su, Leiyu Pan, Minxuan Lv, Yuntao Li, Wenping Hu, Fuzheng Zhang, Kun Gai, and Guorui Zhou. 2025. Ce-gppo: Controlling entropy via gradientpreserving clipping policy optimization in reinforcement learning. arXiv preprint arXiv:2509.20712.
Wenkai Yang, Weijie Liu, Ruobing Xie, Yiju Guo, Lulu Wu, Saiyong Yang, and Yankai Lin. 2025b. Laser: Reinforcement learning with last-token selfrewarding. arXiv preprint arXiv:2510.14943.
GLM Team, Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, and 1 others. 2025. Glm4.5: Agentic, reasoning, and coding (arc) foundation models. arXiv preprint arXiv:2508.06471.
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, YuYue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, and 1 others. 2025. DAPO: An open-source LLM reinforcement learning system at scale. In The Thirty-ninth Annual Conference on Neural Information Processing Systems.
Chen Wang, Zhaochun Li, Jionghao Bai, Yuzhi Zhang, Shisheng Cui, Zhou Zhao, and Yue Wang. 2025a. Arbitrary entropy policy optimization breaks the exploration bottleneck of reinforcement learning. arXiv preprint arXiv:2510.08141.
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, and Gao Huang. 2025. Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? arXiv preprint arXiv:2504.13837.
Jiakang Wang, Runze Liu, Lei Lin, Wenping Hu, Xiu Li, Fuzheng Zhang, Guorui Zhou, and Kun Gai. 2025b. Aspo: Asymmetric importance sampling policy optimization. arXiv preprint arXiv:2510.06062.
Runzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu, Dongrui Liu, Jing Shao, Derek F. Wong, and Yu Cheng. 2025. Exgrpo: Learning to reason from experience. arXiv preprint arXiv:2510.02245.
Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, Kai Dang, Xionghui Chen, Jianxin Yang, Zhenru Zhang, and 1 others. 2025c. Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning. arXiv preprint arXiv:2506.01939.
Wenhao Zhang, Yuexiang Xie, Yuchang Sun, Yanxi Chen, Guoyin Wang, Yaliang Li, Bolin Ding, and Jingren Zhou. 2025a. On-policy rl meets off-policy experts: Harmonizing supervised fine-tuning and reinforcement learning via dynamic weighting. arXiv preprint arXiv:2508.11408.
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, and 1 others.
11
Xiaoyun Zhang, Xiaojian Yuan, Di Huang, Wang You, Chen Hu, Jingqing Ruan, Kejiang Chen, and Xing Hu. 2025b. Rediscovering entropy regularization: Adaptive coefficient unlocks its potential for llm reinforcement learning. arXiv preprint arXiv:2510.10959. Rosie Zhao, Alexandru Meterez, Sham Kakade, Cengiz Pehlevan, Samy Jelassi, and Eran Malach. 2025. Echo chamber: Rl post-training amplifies behaviors learned in pretraining. arXiv preprint arXiv:2504.07912. Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, and 1 others. 2025. Group sequence policy optimization. arXiv preprint arXiv:2507.18071.
A
Figure 5: The evolution of the OGER exploration reward during training. The transition from low to stabilized reward levels reflects the shift from imitation to continues exploration.
Experiments Detail
Training Configurations For the Qwen2.5-Math backbone, the original context length is 4,096 tokens. To accommodate the requirements of complex reasoning tasks and maintain alignment with the offline teacher trajectory lengths, we extended the context window to 16,384 tokens. Accordingly, we adjusted the RoPE theta from 10,000 to 40,000. For all reinforcement learning experiments, the learning rate was set to 1 × 10−6 with a tensor parallelism (TP) degree of 2. For the Qwen2.5Math-7B model, we trained for up to 1,950 steps, while the smaller Qwen2.5-Math-1.5B variant was trained for 1,500 steps. For all RL training sessions, the training was performed on a single cluster of 8× NVIDIA H200 GPUs.
Figure 6: The pass@8 performance across different inference temperatures on AIME 2024 and AIME 2025, we illustrate the average score.
high-quality reasoning patterns from the offline teacher trajectories. As training progresses into the mid-to-late stages, the model’s intrinsic reasoning proficiency matures, leading to an increase in exploratory signals. The OGER reward subsequently stabilizes at a specific plateau, facilitating a dual-objective learning process: it continuously absorbs the sophisticated reasoning logic of the teacher models while simultaneously leveraging the exploration reward to autonomously discover novel reasoning paths that deviate from the static offline demonstrations. This dynamic transition effectively balances expert-guided imitation with autonomous latent space exploration.
System Prompt For all experiments, we share the same system prompt in section for training. For evaluation, we use the same system prompt for inference for the mathematical tasks, and remove the system prompt at OOD tasks.
B
The Balance of Imitation and Exploration
To further dissect the behavioral shift of the model during the training process, we analyze the evolution of the OGER reward (both Mean and Max values), as illustrated in fig. 5. Our exploration reward is formulated based on the semantic divergence between online and offline trajectories. In the early stages of training, due to the model’s limited initial reasoning capacity, the generation of correct online samples is relatively scarce. Consequently, the optimization focus primarily gravitates toward internalizing the
C
OGER Maintains Exploration During Inference
To evaluate the robustness of our framework across varying degrees of exploration during inference, we compare the average pass@8 scores of OGER, 12
System Prompt Your task is to follow a systematic, thorough reasoning process before providing the final solution. This involves analyzing, summarizing, exploring, reassessing, and refining your thought process through multiple iterations. Structure your response into two sections: Thought and Solution. In the Thought section, present your reasoning using the format: "<think>\n {thoughts} </think>\n". Each thought should include detailed analysis, brainstorming, verification, and refinement of ideas. After "</think>\n", in the Solution section, provide the final, logical, and accurate answer, clearly derived from the exploration in the Thought section. If applicable, include the answer in \boxed{} for closed-form results like multiple choices or mathematical solutions. User: This is the problem: {Question} Assistant: <think> Table 5: The system prompt used for all experiments.
Luffy, and GRPO on the AIME 2024 and AIME 2025 benchmarks. We systematically vary the decoding temperature from 0.1 to 1.0 in increments of 0.1. As illustrated in fig. 6, our proposed OGER framework consistently and significantly outperforms both Luffy and GRPO across the entire temperature spectrum. This consistency highlights OGER’s stability under diverse exploratory constraints at test time. Furthermore, a horizontal comparison reveals that OGER maintains a relatively stable performance profile as the temperature increases. In contrast, GRPO exhibits a marked performance degradation at higher temperatures. This phenomenon underscores OGER’s enhanced generalizable reasoning capability; whereas baseline models may become erratic or lose reasoning coherence when stochasticity increases, OGER preserves a structured and effective reasoning manifold, effectively bridging the gap between training-time exploration and test-time robustness.
13