ConceptioArchivearXiv CS
arXiv CSopen access

Visually-Guided Policy Optimization for Multimodal Reasoning

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Visually-Guided Policy Optimization for Multimodal Reasoning Zengbin Wang1 * , Feng Xiong1 * , Liang Lin1 , Xuecai Hu1† Yong Wang1† , Yanlin Wang2 , Man Zhang3 , Xiangxiang Chu1 1 AMAP, Alibaba Group 2 SYSU 3 BUPT

Abstract

Multimodal Question

arXiv:2604.09349v1 [cs.CV] 10 Apr 2026

Instruction:

Reinforcement learning with verifiable rewards (RLVR) has significantly advanced the reasoning ability of vision-language models (VLMs). However, the inherent text-dominated nature of VLMs often leads to insufficient visual faithfulness, characterized by sparse attention activation to visual tokens. More importantly, our empirical analysis reveals that temporal visual forgetting along reasoning steps exacerbates this deficiency. To bridge this gap, we propose Visually-Guided Policy Optimization (VGPO), a novel framework to reinforce visual focus during policy optimization. Specifically, VGPO initially introduces a Visual Attention Compensation mechanism that leverages visual similarity to localize and amplify visual cues, while progressively elevating visual expectations in later steps to counteract visual forgetting. Building on this mechanism, we implement a dual-grained advantage re-weighting strategy: the intra-trajectory level highlights tokens exhibiting relatively high visual activation, while the inter-trajectory level prioritizes trajectories demonstrating superior visual accumulation. Extensive experiments demonstrate that VGPO achieves better visual activation and superior performance in mathematical multimodal reasoning and visual-dependent tasks. The code has been released at https: //github.com/wzb-bupt/VGPO.

1

Introduction

Vision-Language Models (VLMs) (Hurst et al., 2024; Bai et al., 2025b; Comanici et al., 2025), renowned for their multimodal reasoning over both image and text, have achieved remarkable progress across multimodal tasks (Lu et al., 2024; Zhang et al., 2024b; Yue et al., 2025; Zhou et al.; Qiao et al., 2025). Recently, the emerging Reinforcement Learning with Verifiable Rewards * Equal contribution. †

Project leads and corresponding authors.

Vision Input

You first think through the reasoning process as an internal monologue, enclosed within <think> </think> tags. Then, provide your final answer enclosed within \boxed{}.

Qwen2.5-VL-7B-Instruct <think>To determine the total cost for Arianna to buy one of each type of quartz, we need need to sum add the prices of all the items listed. The prices are : - yellow tiger's eye: $0.36; - pumice stone: $0.27; Visual-Dependent - quartz crystal: $0.28; <···>; - piece of green slate: $0.59. Now, let's add these prices together: Text-Dependent $0.36 + $0.27 + $0.28 + $0.26 + $0.99 + $0.59. <···> So, the total cost for Arianna to buy one of each type of quartz is $2.75. <\think> \boxed{2.75}.

VGPO (Visually-Guided Policy Optimization, Ours) <think>To determine how much it will cost Arianna to buy one of each type of quartz, we need to add up the prices of each type of quartz listed in the table: - yellow tiger's eye: $0.36; <···>; - quartz crystal: $0.28; Visual-Dependent - piece of rose quartz: $0.26; <···>; - piece of green slate: $0.59. However, the question specifies that Arianna wants to buy one of each type of Text-Dependent quartz. The types of quartz in the table are: - quartz crystal $0.28; - piece of rose quartz $0.26; - piece of smoky quartz $0.99. Visual-Dependent We will add the prices of these three types of quartz: $0.28 + $0.26 + $0.99 = $1.53 Text-Dependent Therefore, the cost for Arianna to buy one of each type of quartz is $1.53.<\think> \boxed{1.53}.

··· Query Extraction

···

··· Vision Extraction (early)

···

Continuous Vision Focus (late)

Figure 1: A multimodal reasoning example with visual input. Our VGPO not only attends to the query and visual elements in the early reasoning steps, like the base model (Bai et al., 2025b), but also exhibits better visual activation and sustains visual focus in the later steps to ensure consistent visual faithfulness.

(RLVR) (Shao et al., 2024; Yu et al., 2025; Zheng et al., 2025; Chu et al., 2025; Dai et al., 2026) has further advanced into complex step-by-step reasoning, yielding enhanced logical coherence and superior performance (Huang et al., 2025b; Meng et al., 2025; Wang et al., 2025a; Huang et al., 2025a; Yang et al., 2025c; Wang et al., 2025e). Despite exhibiting better logical reasoning, the inherent text-dominated inference nature of VLMs often leads to insufficient visual faithfulness (Wang et al., 2024; Fu et al., 2024; Liu et al., 2025c,b). As discussed in recent literature (Jian et al., 2025; Yang et al., 2025b; Favero et al., 2024) and evidenced by Figures 1 and 2, models often assign

lower attention to input image tokens compared to query tokens and generated text tokens, resulting in sparse activation of input image. More critically, this deficit will be exacerbated by temporal visual forgetting: as reasoning chains extend, attention to visual inputs progressively decays. This over-reliance on textual priors rather than visual facts often leads to hallucinations or reasoning errors (Sun et al., 2025b; Wang et al., 2025e; Tian et al., 2025). Therefore, a natural question arises: Q1: How can we effectively amplify visual activation and mitigate temporal visual forgetting to ensure visual faithfulness? Recent works primarily focused on the former aspect “amplify visual response” through external interventions. For example, (1) Look-Back (Yang et al., 2025b) and latent reasoning methods (Li et al., 2025a; Sun et al., 2025a; Yang et al., 2025d) aimed to introduce external learnable special tokens (e.g., <back>, <latent_strat>, <latent_end>) that act as explicit triggers to revisit or reconstruct visual input through initial supervised fine-tuning and subsequent autonomous generation. (2) VAPO (Tian et al., 2025) explored an additional GPT-5 model to verify the visual faithfulness during intermediate steps and re-activate visual responses through probing visual-related questions. (3) PAPO (Wang et al., 2025e) and VPPO (Huang et al., 2025a) identified and highlighted visual tokens by contrasting the KL divergence between original versus noisy images through two forward processes. While effective, these methods inevitably introduce extra training tokens, auxiliary models, or additional forward passes. In this case, another question arises: Q2: How can we leverage the model’s own internal states to localize visual activation, without any external dependencies? To comprehensively answer these two questions, the key solution lies in (i) precisely localizing visual activations; (ii) amplifying visual activations; and (iii) sustaining visual expectation in the later reasoning steps. For the first part, our empirical experiment in Figure 2 reveals that the inherent hidden states similarity between generated tokens and input image tokens serves as a precise Visual Focus Score, enabling spontaneous localization of visually grounded tokens without external supervision. For the second and third parts, visual focus score serves as the basis for further designing a Visual Attention Compensation mechanism that

re-weights visually grounded tokens and progressively improves visual expectations in later reasoning steps to combat temporal visual forgetting. Building on the above two insights, we propose Visually-Guided Policy Optimization (VGPO). Relying solely on the model’s self-generated hidden states to localize visual tokens, VGPO integrates the Visual Attention Compensation mechanism into policy optimization through a dualgrained advantage re-weighting strategy. At intratrajectory level, it dynamically re-weights token advantages to amplify specific visual cues, while applying progressive incentives to compensate for temporal visual decay in later steps. At intertrajectory level, it globally prioritizes superior visual accumulation, effectively compensating for holistic visual negligence by favoring trajectories that sustain consistent grounding. Extensive experiments across mathematical and visiondependent multimodal reasoning tasks demonstrate that VGPO achieves better visual activation and less visual forgetting. In summary, our contributions are four-fold. • We reveal that the inherent hidden states similarity between generated tokens and image tokens serves as a reliable Visual Focus Score to localize visual tokens, thereby facilitating targeted attention modulation without external supervision. • We design a Visual Attention Compensation mechanism that leverages visual focus score to amplify visual cues, while progressively elevating visual expectations in the latter reasoning step to counteract temporal visual forgetting. • We propose VGPO, which integrates the visual attention compensation mechanism into policy optimization using a Dual-Grained Advantage Re-Weighting strategy to amplify visual activation and mitigate visual forgetting across intratrajectory and inter-trajectory levels. • Extensive experiments demonstrate that VGPO achieves better visual activation and state-ofthe-art performance on mathematical and visualdependent multimodal reasoning tasks.

2

Preliminary

2.1

Group Relative Policy Optimization

Group Relative Policy Optimization (GRPO) (Shao et al., 2024) is a resource-efficient variant of reinforcement learning, which eliminates the value model (Schulman et al., 2017) by estimating advantages through group-based computation. Formally,

(a)

(b)

(c)

Query: If Arianna wants to buy one of each type of quartz, how much will it cost her?

LogicVista

CLEVR Counting

MMMU-Pro

MathVerse-V

Figure 2: Analysis of the inference nature of multimodal reasoning trajectory (based on Qwen2.5-VL-7B (Bai et al., 2025b)). (a) An example of attention allocation across image, query, and generated text tokens (normalized to 1 at each step). (b) Average attention statistics on four visual-dependent benchmarks (Xiao et al., 2024; Li et al., 2023; Yue et al., 2025; Zhang et al., 2024b). (c) Distribution of late/early visual accumulation ratios for incorrect (left) vs. correct (right) samples of these four benchmarks. Incorrect samples often exhibit higher visual forgetting.

we consider visual input (I, q, a), comprising a visual input I, a textual query q, and the corresponding ground truth answer a, the algorithm begins by sampling a group of G candidate responses {oi }Gi=1 from the current policy πθold . For each response oi , a binary reward ri ∈ {0, 1} is assigned based on the exact match with the ground truth a. Subsequently, the advantage Âi for each response is computed by normalizing the assigned rewards within the sampled group: Âi =

ri −mean({rk }G k=1 ) std({rk }G k=1 )

.

The policy πθ is then updated to maximize the following surrogate objective: " JGRPO (θ) =E{oi }G

i=1 ∼πθold (·|I,q)

|oi | G 1X 1 X G i=1 |oi | t=1

  min ri,t (θ)Âi , clip ri,t (θ), 1 − ε, 1 + ε Âi π (o

|q,I,o

!# (1) ,

)

where ri,t (θ) = πθ θ (oi,ti,t |q,I,oi,<t represents probi,<t ) old ability ratio. Following DAPO Yu et al. (2025), we adopt clip-higher setting, asymmetric clipping hyper-parameters, and exclude the KL penalty. 2.2

Key Findings in Multimodal Reasoning

To investigate the underlying flaw in multimodal reasoning, we conduct an empirical analysis of how internal attentions are allocated and evolved throughout VLMs’ generation in Figure 2. This reveals three key findings that inspire our method. (1) Text-dominated Inference and Sparse Visual Activation. First, we analyze the attention allocation among three types of tokens: input image I, input query q, and current generated text o<t . As illustrated in Figure 2(a)(b), we observe a significant dominance of textual priors: the model

heavily attends to both generated text history (blue line) and input text query (gray line), leaving visual attention (red line) lower than the sum of these text attentions. Furthermore, visual activation is often sparsely activated in the early stage, manifesting as a brief “glance” the image and followed by long periods of neglect in the later stage. (2) Temporal Visual Forgetting. Second, we identify a critical correlation between reasoning length and visual activation, which we term temporal visual forgetting. The intensity of visual attention (red line) in Figure 2(b) across four benchmarks shows a similar trend: visual activation initially improves and progressively decays as the generation step increases. To further validate the impact of this decay, we statistically compared the attention allocation of correct versus incorrect reasoning samples. Figure 2(c) reveals that correct samples often exhibit higher late/early visual accumulation ratios (Average, 0.680 vs. 0.532), indicating better temporal visual forgetting mitigation. (3) Inherent Visual Similarity as a Precise Visual Focus Score. Finally, despite the sparsity of visual activation, we observe a promising phenomenon in the bottom of Figure 2(a): when visual tokens are activated (i.e., the peaks in the red line), their attention maps on the original image are highly accurate and semantically grounded based on current base models. This indicates the model can locate relevant information but lacks sustained focus in the following reasoning steps. Therefore, this visual similarity can serve as a precise, intrinsic visual focus score to detect and re-weight critical visual tokens, eliminating the need for auxiliary models.

update Outcome Advantage 𝐴"!,#

⋯ ⋯ ⋯ ⋯

Policy Model

Visual Similarity

Input Query Tokens

⋯ ⋯ ⋯ ⋯

Visual 𝛍$ Prototype

Input Image Tokens

(a) Visual Focus Score ρ!,#

Generated Text Tokens

ρ!,# 1 + 𝐺! ρ!,# β

⋯ ⋯ ⋯ ⋯

⋯ ⋯ ⋯ ⋯

Verifier & Group Norm

𝑡 𝑇!

Decoding Step (t)

(b) Visual Compensation

Reshaped Advantage 𝐴"𝒱!,#

Intra-traj. 𝜓!,#

⋯ ⋯ ⋯ ⋯

Inter-traj. 𝜙!,#

amplify

🔥

Query: Is the cat's paw on the carrot?

Trajectories 𝜏!

⋯ ⋯ ⋯ ⋯

(c) Advantage Re-weighting

Figure 3: Overview of Visually-Guided Policy Optimization framework. Given query and image, (a) VGPO firstly utilizes the intrinsic hidden state similarity between generated tokens and visual prototype to derive a Visual Focus Score for visual token localization. (b) Then, Visual Attention Compensation (VAC) mechanism leverages this score to re-focus visual tokens, while progressively elevating visual expectations along decoding steps to counteract temporal visual forgetting. (c) Finally, Dual-grained Advantage Re-weighting strategy integrates VAC mechanism into intra- and inter-trajectory levels to explicitly incentivize sustained visual faithfulness during policy updates.

3

Methodology

In this section, we present Visually-Guided Policy Optimization (VGPO) framework. As in Figure 3, VGPO is structured into three key components, including Visual Focus Score (Section 3.1), Visual Attention Compensation mechanism (Section 3.2), and Dual-grained Advantage Re-weighting strategy (Section 3.3) to transform intrinsic visual attention into explicit signals to incentivize sustained visual focus throughout the reasoning process. 3.1

Visual Focus Score

To quantify the visual engagement of each generated token without external supervision, we introduce an intrinsic metric based on hidden state similarity, as we discussed in Section 2.2. Let d v Hv = {hvk }N k=1 ⊂ R denote the hidden states of the Nv input image tokens. To capture the global visual semantics, we derive a visual prototype µv : µv =

Nv X

αk hvk ,

s.t.

X

αk = 1,

(2)

k=1

where αk represents the weight of the k-th image token. For the i-th trajectory τi at decoding step t, we measure the cosine semantic similarity S(·, ·) between the current token hidden state hi,t and the established visual prototype µv : S(hi,t , µv ) =

(hi,t )⊤ µv , ∥hi,t ∥2 ∥µv ∥2 + ϵ

(3)

where ϵ is a smoothing term. Finally, the visual

focus score ρi,t for the t-th token in trajectory τi is obtained via a scaling function to normalize the values to the range [0, 1]: 1 ρi,t = (S(hi,t , µv ) + 1) . (4) 2 3.2 Visual Attention Compensation However, directly employing the raw visual focus score ρi,t is suboptimal due to the temporal visual forgetting phenomenon observed in Section 2.2. As the reasoning chain extends, intrinsic visual attention naturally decays, causing visually grounded tokens in later steps to often exhibit suppressed focus scores and lead to insufficient optimization signals. To mitigate this, we introduce a Visual Attention Compensation mechanism that counteracts this decay by linearly elevating the expectations in the later reasoning step. Specifically, for i-th trajectory τi with token length Ti :   t wi,t = ρi,t · 1 + Gi (ρi,t ) · β · , (5) Ti where the ratio Tti indicates a linear strategy to progressively amplify the visual optimization signal as the reasoning step extends. β is a hyper-parameter controlling visual compensation intensity. Gi (·) acts as a visual gate to selectively identify tokens with higher visual similarity (filtering out noisy text tokens) in the later reasoning step. Formally, ( 1, if t>(1−γ)Ti and ρi,t ≥Qκ (Ptail,i ) Gi (ρi,t )= , 0, otherwise (6)

where Ptail,i = {ρi,k | k > (1−γ)Ti } represents the set of score in the tail of trajectory τi , and Qκ (·) returns threshold for the top κ-percent of this set. 3.3

Dual-grained Advantage Re-weighting

To effectively integrate visual attention compensation into the policy optimization, we introduce a Dual-grained Advantage Re-weighting mechanism. This approach modulates the standard advantage function by assessing visual relevance at both the Intra-trajectory and Inter-trajectory levels. Intra-trajectory Re-weighting. This component aims to capture the distinctions of individual tokens based on their local visual saliency. To highlight visual significance within a specific reasoning path, we derive an intra-trajectory scaling factor ψi,t . Specifically, we first normalize the raw visual focus scores wi,t via Min-Max scaling to ensure numerical stability. Subsequently, we zero-center these scores relative to the trajectory’s mean, ensuring that only tokens with above-average visual activation are positively incentivized, while effectively suppressing non-visual tokens. Formally, wi,t − mink wi,k w bi,t = , maxk wi,k − mink wi,k + ϵ

(7)

T

ψi,t = w bi,t −

i 1 X w bi,k . Ti

(8)

k=1

Inter-trajectory Re-weighting. While the intratrajectory adjustment addresses local granularity, it is equally crucial to evaluate the global visual focus of generated sequences. Correspondingly, we design the inter-trajectory re-weighting strategy to incentivize the model to prioritize entire generations that exhibit superior aggregate visual score. For each trajectory τi , we compute a cumulative visual score si , which aggregates the compensated weights over all timesteps. As shown in Eq. 9, this score explicitly accounts for both the base relevance and the time-dependent compensation: si =

Ti X

wi,t =

t=1

Ti X

  Ti X t ρi,t + ρi,t ·Gi (ρi,t )· β· . (9) Ti t=1 t=1 | {z } | {z }

Visual Relevance Late-stage Compensation

Finally, to derive the final inter-trajectory scaling factor ϕi , we apply group-wise normalization and zero-centering across the rollout group G: sbi =

si − minτj ∈G sj , maxτj ∈G sj − minτj ∈G sj + ϵ

(10)

G

ϕi = sbi −

1X sbj . G

(11)

j=1

Overall Integration. By synergizing the intratrajectory re-weighting with the inter-trajectory reweighting, we formulate the final visual-focus advantage Âi,t . This mechanism essentially modulates the standard outcome-based advantage Âi (Eq. 1) to explicitly incorporate visual focus signals into the policy update. The final advantage is computed as a multiplicative integration of the base advantage and the dual-grained visual factors: ÂVi,t = Âi · (1 + ψi,t ) · (1 + ϕi ).

(12)

Here, the modulation terms (1 + ψi,t ) and (1 + ϕi ) serve as fine-grained and coarse-grained scaling factors, respectively. This composite advantage function reshapes the optimization landscape, ensuring that the policy is driven not merely by the correctness of the final answer, but also by the faithful utilization of visual focus.

4

Experiments

In this section, we present the experimental methodology and the analysis of the results. Further details and results are provided in the Appendix A. 4.1

Experimental Settings

Models and Datasets. Consistent with prior studies (Wang et al., 2025e), we adopt Qwen2.5-VLseries (3B, 7B, and 32B) (Bai et al., 2025b) as our backbone models. We train these models in ViRL39K (Wang et al., 2025a), Geo3K (Lu et al., 2021), and MMK12 (Meng et al., 2025), while validating in MMK12-val (Meng et al., 2025). Baselines and Evaluations. We conduct a comprehensive evaluation of our approach against a diverse set of leading VLMs at 7B scale. Our baselines include ThinkLite-VL-7B (Wang et al., 2025d), VL-Rethinker-7B (Wang et al., 2025a), MMEureka-7B (Meng et al., 2025), NoisyRollout7B (Liu et al., 2025a), PAPOD -7B (Wang et al., 2025e), and VPPO-RL-7B (Huang et al., 2025a). To assess our VGPO, we employ a suite of benchmarks categorized as follows: (1) General Mathematical & Geometric Reasoning. We utilize MathVista (Lu et al., 2024), MathVerse (Zhang et al., 2024b), We-Math (Qiao et al., 2025), MMK12 (Meng et al., 2025), GeoMath (Zhou et al.), and Geometry3K (Lu et al., 2021). (2) Vision-dependent Multimodal Reasoning. We adopt

General Mathematical & Geometric Reasoning

Models

Avg-Math

MathVista MathVerse WeMath MMK12 GeoMath Geo3k

Vision-dependent Multimodal Reasoning

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV

Open-Source General MLLMs (32B-72B) Qwen2.5-VL-72B Qwen3-VL-32B

74.8 83.8

65.2†

69.4†

71.5†

75.2

65.9† 63.1†

53.5† 45.7†

54.2† 65.4†

63.8 67.5

57.9† 62.2

80.5† 95.5†

51.1 65.3

57.6 76.8

61.8 75.0

44.3 47.0 47.9 47.0 45.9 48.8

88.5 76.5 79.0 88.5 89.0 88.5

32.3 37.6 40.6 38.2 40.2 40.2

30.9 64.8 63.6 51.7 66.6 67.6

49.0 56.5 57.8 56.3 60.4 61.3

45.2 49.4 47.4 49.4

76.5 85.5 85.5 95.5

36.4 38.6 39.0 40.5

36.6 61.8 66.6 67.6

48.7 58.8△20.7% 59.6△22.4% 63.3△30.0%

Open-Source Multimodal Reasoning Models (∼7B) ThinkLite-VL-7B† VL-Rethinker-7B† MM-Eureka-7B† NoisyRollout-7B† PAPOD -7B† VPPO-RL-7B†

73.6 65.2 66.3 72.9 72.3 70.5

33.7 68.8 66.7 58.8 69.5 71.1

43.7 67.5 66.6 64.4 69.4 70.6

50.5 66.0 61.7 51.2 80.8 81.8

44.6 50.1 49.6 51.6 52.3 53.2

45.3 44.4 41.1 55.2∗ 48.8 46.9

48.6 60.3 58.7 59.0 65.5 65.7

Qwen2.5-VL-7B + GRPO + DAPO + VGPO (Ours)

68.5 70.1 68.7 74.1

40.2 66.7 69.6 71.6

47.8 69.8 70.8 72.5

49.4 72.9 77.0 81.5

51.2 52.8 51.3 54.3

42.9 50.0 43.4 62.6△25.2% 45.6 63.8△27.6% 45.8 66.6△33.2%

VGPO (Ours)

Average Reward

0.8

GRPO

DAPO

0.7 0.6 0.5 0

25

50

75

100

Training Steps

125

(a) Training rewards.

150

Average Accuracy

Table 1: Performance comparisons across mathematical and vision-dependent multimodal reasoning benchmarks. “∗ ” indicates NoisyRollout is trained on Geo3K. “† ” indicates our reproduction with official checkpoints. Other values are sourced from the official technical report. Bold and underlined indicate best and second-best results. VGPO (Ours)

0.9

GRPO

DAPO

0.8 0.7 0.6 0.5 0.4 0

25

50

75

100

Training Steps

125

150

(b) Validation accuracy.

Figure 4: Training dynamics based on Qwen2.5-VL7B: (a) training rewards and (b) validation accuracy on MMK12 (Meng et al., 2025) across GRPO (Shao et al., 2024), DAPO (Yu et al., 2025), and our VGPO.

LogicVista (Xiao et al., 2024), SuperClevr Counting (Li et al., 2023), MMMU-Pro (Yue et al., 2025), and MathVerse-V (Zhang et al., 2024b). Specifically, we adopt these datasets from PAPOEval (Wang et al., 2025e), which provide filtered subsets enabling robust verification without reliance on LLM-as-a-judge (Chen et al., 2024a). Implementation Details. Following most settings in current methods (Wang et al., 2025e), we set the training epoch to 2, with learning rate of 1 × 10−6 , rollout batch size of 512, and maximum response length of 2,048. During the evaluation phase, we set temperature to 0.0 for all experiments. Regarding the specific hyper-parameters of our VGPO, we employ mean-pooling strategy for visual prototype and set compensation intensity β = 0.3, position γ = 0.5, and κ-percent κ = 0.2. 4.2

Main Results

Superior Performance on Multimodal Reasoning. As shown in Table 1, our proposed VGPO

achieves significant improvements over the base Qwen2.5-VL-7B model, delivering substantial relative gains of 33.2% on general mathematical reasoning and 30.0% on vision-dependent multimodal reasoning tasks. Furthermore, when compared to other advanced reasoning models initialized from the same backbone, VGPO consistently secures the top performance, surpassing previous state-ofthe-art results with superior average accuracies of 66.6% in general mathematical benchmarks and 63.3% in vision-dependent tasks. Notably, despite the significant disparity in model scale, our 7B model exhibits highly competitive performance comparable to the much larger Qwen2.5-VL-72B, highlighting the efficiency of our approach in maximizing the potential of smaller-scale models. Training Dynamics and Generalization. We further analyze the training stability and generalization capabilities in Figure 4. VGPO demonstrates a more stable and efficient learning trajectory, achieving consistently higher training rewards compared to baselines like GRPO and DAPO. This superiority can also extend to generalization, where VGPO maintains a clear lead in validation accuracy throughout the training steps. Unlike competitive methods that may suffer from instability or slower convergence, VGPO establishes a robust optimization path, ensuring steady performance gains. Scalability across Model Scales and Training Data. To validate the scalability, as summarized in Table 2, experiments ranging from 3B to 32B parameters (Qwen2.5-VL-based) demonstrate that VGPO consistently outperforms baselines, achieving significant relative gains (e.g., 13.8% on math-

General Mathematical & Geometric Reasoning

Models

Vision-dependent Multimodal Reasoning

Avg-Math

MathVista MathVerse WeMath MMK12 GeoMath Geo3k

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV

Exploring More Scalable Models (7B → − 3B, 32B) Qwen2.5-VL-3B + DAPO + VGPO

54.0 63.9 65.0

44.1 57.1 61.4

50.2 62.9 62.1

44.1 64.1 71.6

44.3 47.6 50.2

27.6 44.1 36.3 55.3△25.4% 36.1 57.7△30.8%

39.2 43.0 45.9

58.5 65.5 78.0

28.6 31.5 32.5

37.5 53.1 58.1

41.0 48.3△17.8% 53.6△30.7%

Qwen2.5-VL-32B + DAPO + VGPO

76.8 69.7 75.3

62.4 73.0 73.8

68.8 78.7 78.7

60.0 82.0 87.3

53.7 55.2 56.5

50.8 62.1 51.6 68.4△10.1% 52.3 70.7△13.8%

55.7 59.1 59.3

80.5 85.5 90.0

46.0 47.3 47.3

58.4 67.4 70.1

60.2 64.8△7.6% 66.7△10.8%

Qwen2.5-VL-7B + DAPO w/ Geo3K 2.1K + VGPO w/ Geo3K 2.1K

68.5 71.6 72.2

40.2 56.3 62.5

47.8 63.6 66.6

49.4 51.1 53.0

51.2 49.4 52.6

42.9 50.0 52.3 57.4△14.8% 55.6 60.4△20.8%

45.2 46.5 47.7

76.5 85.5 80.5

36.4 36.8 37.0

36.6 50.3 57.8

48.7 54.8△12.5% 55.8△14.6%

+ DAPO w/ MMK12 6.4K + VGPO w/ MMK12 6.4K

72.6 74.1

66.2 68.5

67.5 68.9

64.9 66.0

51.2 53.2

42.3 60.8△21.6% 43.4 62.4△24.8%

47.4 48.8

89.5 90.0

37.1 38.6

61.2 63.6

58.8△20.7% 60.3△23.8%

Exploring More Training DataSets (39K → − 2.1K, 6.4K)

Table 3: Ablation study on the impact of Intra- and Intertrajectory re-weighting strategies. Method Qwen2.5-VL-7B + DAPO + DAPO w/ Entropy + DAPO w/ KLperception + VGPO (Ours)

Avg-Math Avg-Vision Overall 50.0 63.8 65.6 65.7 66.6

48.7 59.6 61.9 61.3 63.3

49.5 62.2 64.1 63.9 65.3

Table 4: Comparisons with other advantage shaping strategies, including the Entropy-based (Cheng et al., 2025) and KL-based (Huang et al., 2025a) methods.

ematical tasks, 10.8% on visual-dependent tasks) based on 32B model. Additionally, evaluations on different training datasets with varying sizes reveal superior data efficiency, where our method surpasses the DAPO baseline even in data-constrained regimes. These results suggest that incorporating visual focus into the training process is a generalizable strategy to enhance multimodal reasoning. 4.3

Quantitative Analysis

Ablation Study on Re-weighting Strategy. As detailed in Table 3, while the baseline model provides a solid foundation, the addition of both intra- and inter-trajectory mechanisms provides extra gains in reliability and specific tasks. Their combination achieves peak average performance across all metrics, confirming that these two strategies are highly complementary and work synergistically to enhance multimodal reasoning capability.

1.2 1.0 0.8 0

50 100 Training Steps

150

Average Accuracy

62.2 64.6 64.0 65.3

= 0.5 = 0.8

(a) Training dynamics of β. = 0.0 = 0.2

1.2

= 0.4 = 0.6

= 0.8 = 1.0

1.0 0.8 0

50 100 Training Steps

150

1.2

= 0.6 = 0.7

0.8 50 100 Training Steps

65

66.6 65.3

150

(e) Training dynamics of γ.

= 0.5 = 0.8

65.8 64.4 63.3

62

62.0

61.6

60

Math

65 65.3

66.2

61.1

Vision = 0.4 = 0.6

= 0.8 = 1.0

65.6 65.6 63.3

62.4

61.6

61.8 61.8

60 55

60.7

55.3

54.5

Math

Vision

(d) Performance w.r.t. κ. = 1.0 = 0.5

= 0.8

1.0 0

68

= 0.0 = 0.2

(c) Training dynamics of κ. = 1.0 = 0.5

= 0.0 = 0.3

(b) Performance w.r.t. β. Average Accuracy

59.6 62.5 62.0 63.3

= 0.0 = 0.3

Average Accuracy

63.8 66.1 65.3 66.6

Late/Early Ratio

DAPO Baseline + Intra-trajectory + Inter-trajectory + Intra- & Inter-trajectory

Avg-Math Avg-Vision Overall

Late/Early Ratio

Method

Late/Early Ratio

Table 2: Performance comparisons across different model scales and training datasets.

= 0.6 = 0.7

= 0.8

66.6

65

65.3

65.0 63.3 62.1 60.8

60 Math

61.6

61.3 60.1

59.1

Vision

(f) Performance w.r.t. γ.

Figure 5: Ablation study of training dynamics of the late/early ratio and performance on hyperparameters β, κ, and γ. The Late/Early Ratio is calculated by dividing the visual attention score of the late reasoning stage by that of the early stage.

Comparison with Existing Advantage Shaping Methods. We evaluate the efficacy of our VGPO against established advantage shaping strategies, specifically regularized derivatives of DAPO. As shown in Table 4, while entropy and KL regularization effectively regulate advantage estimation relative to vanilla DAPO, VGPO achieves superior empirical performance with a peak average accuracy of 63.3%. This advantage is particularly pronounced in vision-centric tasks like LogicVista

Method Baseline (DAPO) + Step-Function + Exponential + Linear (Ours)

Avg-Math

Avg-Vision

Overall

63.8 64.7 65.1 66.6

59.6 60.7 61.0 63.3

62.2 63.1 63.5 65.3

Table 5: Ablation study on the impact of different compensation schedules. Method Baseline (DAPO) + Full-trajectory + Late-trajectory (Ours)

LogicVista

CLEVR Counting

MMMU-Pro

MathVerse-V

Avg-Math Avg-Vision Overall 63.8 53.0 66.6

59.6 54.2 63.3

62.2 53.5 65.3

Table 6: Ablation study on the impact of Full-trajectory versus Late-trajectory compensation.

and Counting, confirming that strengthening visual focus is a more effective strategy for enhancing multimodal reasoning than generic regularization. Sensitivity to Hyperparameters. To evaluate the impact of different parameters on our method and uncover the underlying relation between training dynamics and performance, we conduct a sensitivity analysis on three key hyperparameters: compensation intensity β, threshold κ, and tail ratio γ. As in Figure 5, while the model exhibits performance fluctuations across different settings, it achieves optimal accuracy with specific configurations (i.e., β = 0.3, κ = 0.2, and γ = 0.5). Crucially, by correlating these performance peaks with the evolution of the Late/Early Ratio, we identify a decisive factor for training stability: the model yields the best performance when this ratio converges to or stabilizes near 1. This convergence implies that the optimal hyperparameter configuration facilitates a necessary equilibrium between late and early visual focus. By harmonizing the contributions of these two stages, our method ensures robust performance across diverse reasoning tasks. Comparison of Different Compensation Schedules. To explore the impact of different visual compensation schedules, we compare our Linear strategy against Step-Function and Exponential schedules, as summarized in Table 5. The Linear strategy consistently achieves the best performance. This is because the Exponential schedule tends to overcorrect by placing too much emphasis on the final tokens (often formatting or calculation), while the Step-Function schedule introduces training instability by abruptly changing the compensation value. The Linear schedule, however, aligns well with the progressive and continuous decay of visual atten-

Figure 6: Comparison of the vision attention ratio distribution before and after our VGPO across four visualdependent multimodal reasoning benchmarks.

tion observed in our empirical analysis. Detailed results and analyses are provided in Appendix B.4. Comparison with Full-trajectory Compensation. We justify our choice of the late-trajectory compensation strategy by comparing it with a fulltrajectory approach, as shown in Table 6. The fulltrajectory compensation significantly drops performance across most benchmarks. This is because current VLMs naturally exhibit high visual attention in the early stages, and enforcing additional visual compensation early can distract the model from parsing the textual query or overemphasize early visual attention. Our late-trajectory design specifically targets the later stages where visual decay occurs, achieving superior results. Detailed comparisons are provided in Appendix B.6. Analysis of Visual Attention Allocation after our VGPO. To investigate whether our VGPO achieves better visual activation and temporal visual forgetting mitigation, we compare the visual attention ratio of input image before and after our VGPO in Figure 6. Our VGPO exhibits higher visual attention allocation throughout the entire generation process and sustains better temporal visual forgetting mitigation compared with the baseline. This prolonged visual grounding is critical for accurate long-chain reasoning, ensuring the model does not lose track of visual evidence in later stages.

5

Related Work

Multimodal Reasoning Challenges. Following the milestone of Large Language Models (LLMs) in complex step-by-step reasoning (Wei et al., 2022; Xia et al., 2025; Yang et al., 2025a), the research

focus has naturally shifted toward extending these abilities to Vision-Language Models (VLMs) for broader real-world applications through integrating vision encoders (Vaswani et al., 2017) and large language models (Zhang et al., 2024a; Li et al., 2025d; Zhu et al., 2025; Bai et al., 2025a). However, despite significant architectural advancements, current VLMs largely inherit the text-dominated inductive biases of their LLM backbones, frequently manifesting as insufficient visual faithfulness and severe visual hallucinations remain the primary bottleneck (Bai et al., 2024; Wu et al., 2024; Zhong et al., 2024; Liu et al., 2025c; He et al., 2025a). Mainstream Strategies for Multimodal Reasoning. To enhance multimodal reasoning, existing methods primarily focus on an RL paradigm that enables the autonomous refinement of reasoning trajectories via rollout sampling (Shao et al., 2024; Huang et al., 2025b; Shen et al., 2025; Li et al., 2025c; He et al., 2026). Following this paradigm, a series of works have explored two kinds of improvements: including (1) training strategies like improving rollout diversity via mixing normal image and noisy augmentation (Liu et al., 2025a) or entropy-based regulation (Cheng et al., 2025) to balance exploration and exploitation; (2) Visualcentric refinement like fine-grained visual enhancement via KL divergence comparison between normal and noisy images (Huang et al., 2025a; Wang et al., 2025e), re-activating specific reasoning paths via introducing specific tokens (Sun et al., 2025a; Li et al., 2025a; Yang et al., 2025d), or verifying intermediate processes as the reward via auxiliary model (Tian et al., 2025). While effective, how to mitigate the need for external models and solely utilize the inherent states of modal abilities to achieve visual enhancement remains an open question. Visual Perception Methods. Recent concurrent works have also explored enhancing visual perception in multimodal reasoning. PEARL (Zhang et al., 2025) and ViCrit (Wang et al., 2025c) enhance perception through external checklists or synthetic hallucination detection, but they require high computational or data construction costs. On the other hand, SSL4RL (Guo et al., 2025) and VisPlay (He et al., 2025b) utilize unlabeled data yet fail to maintain consistent visual attention during complex reasoning. VGPO overcomes these trade-offs by deriving a Visual Focus Score directly from the internal hidden state. By enforcing continuous visual guidance via an attention compensation mechanism, VGPO

mitigates visual decay and sparse activation. This ensures robust reasoning and cross-domain generalizability, eliminating the reliance on external models or outcome-based proxy tasks found in prior work.

6

Conclusion

In this paper, we explore the critical limitation of insufficient visual activation and its temporal decay during multimodal reasoning. To this end, we propose VGPO, a novel framework designed to alleviate this issue by leveraging intrinsic hidden states to autonomously ground visual focus. By strategically coupling a Visual Attention Compensation mechanism with a dual-grained advantage reweighting strategy, our method ensures consistent and sustained visual reasoning capability. Comprehensive experiments validate that VGPO significantly enhances visual activation and delivers state-of-the-art performance.

Limitations While VGPO effectively elevates visual expectations to amplify activation and mitigate forgetting during reasoning, we identify several avenues for future improvement: • The strategy of progressively elevating visual expectations acts as a robust heuristic that may not represent the globally optimal solution for all reasoning patterns. For instance, in scenarios where the final reasoning steps rely strictly on logical deduction or calculation independent of visual cues, a mandated high visual focus might be less critical. Nevertheless, we regard this work as a pivotal milestone in identifying and addressing the overlooked issue of visual decay. We hope this study serves as a foundation for future research to explore more adaptive mechanisms that can dynamically rectify visual reliance according to the specific context of each reasoning step. • As VGPO is designed to elicit and sustain vision intrinsic to the VLMs, its performance upper bound is inherently constrained by the representational quality of the visual encoder and projector. If the base model fails to encode critical visual features into the hidden states initially, our re-weighting strategy may not effectively recover this missing information. Consequently, our framework focuses on optimizing the utilization of visual cues rather than enhancing the raw perceptual capabilities of visual encoder itself.

References Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025a. Qwen3-vl technical report. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, and 1 others. 2025b. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923. Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. 2024. Hallucination of multimodal large language models: A survey. arXiv preprint arXiv:2404.18930. Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. 2024a. Mllm-asa-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark. In Forty-first International Conference on Machine Learning. Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. 2024b. M3 CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pages 8199–8221, Bangkok, Thailand. Association for Computational Linguistics. Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai, Wayne Xin Zhao, Zhenliang Zhang, and Furu Wei. 2025. Reasoning with exploration: An entropy perspective. arXiv preprint arXiv:2506.14758. Xiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei, and Yong Wang. 2025. Gpg: A simple and strong reinforcement learning baseline for model reasoning. arXiv preprint arXiv:2504.02546. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Yanqi Dai, Yuxiang Ji, Xiao Zhang, Yong Wang, Xiangxiang Chu, and Zhiwu Lu. 2026. Harder is better: Boosting mathematical reasoning via difficulty-aware grpo and multi-aspect question reformulation. arXiv preprint arXiv:2601.20614. Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto. 2024. Multi-modal hallucination control by visual information grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14303–14312.

Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, WeiChiu Ma, and Ranjay Krishna. 2024. Blink: Multimodal large language models can see but not perceive. In European Conference on Computer Vision, pages 148–166. Springer. Xiaojun Guo, Runyu Zhou, Yifei Wang, Qi Zhang, Chenheng Zhang, Stefanie Jegelka, Xiaohan Wang, Jiajun Chai, Guojun Yin, Wei Lin, and 1 others. 2025. Ssl4rl: Revisiting self-supervised learning as intrinsic reward for visual-language reasoning. arXiv preprint arXiv:2510.16416. Jinghan He, Junfeng Fang, Feng Xiong, Zijun Yao, Fei Shen, Haiyun Guo, Jinqiao Wang, and Tat-Seng Chua. 2026. Active zero: Self-evolving visionlanguage models through active environment exploration. arXiv preprint arXiv:2602.11241. Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, and Jinqiao Wang. 2025a. Cracking the code of hallucination in LVLMs with vision-aware head divergence. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3488–3501, Vienna, Austria. Association for Computational Linguistics. Yicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang, and Yonghui Yang. 2025b. Visplay: Selfevolving vision-language models from images. arXiv preprint arXiv:2511.15661. Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, and Yu Cheng. 2025a. Spotlight on token perception for multimodal reinforcement learning. arXiv preprint arXiv:2510.09285. Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Yao Hu, and Shaohui Lin. 2025b. Vision-r1: Incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276. Pu Jian, Junhong Wu, Wei Sun, Chen Wang, Shuo Ren, and Jiajun Zhang. 2025. Look again, think slowly: Enhancing visual reflection in vision-language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 9262–9281. Bangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang, Jialian Wu, Xiaodong Yu, Hao Chen, Emad Barsoum, Muhao Chen, and Zicheng Liu. 2025a. Latent visual reasoning. arXiv preprint arXiv:2509.24251.

Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2025b. LLaVAonevision: Easy visual task transfer. Transactions on Machine Learning Research. Renda Li, Hailang Huang, Fei Wei, Feng Xiong, Yong Wang, and Xiangxiang Chu. 2025c. Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting. arXiv preprint arXiv:2511.09478. Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, and 1 others. 2025d. Perception, reason, think, and plan: A survey on large multimodal reasoning models. arXiv preprint arXiv:2505.04921. Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, and Alan L Yuille. 2023. Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14963–14973. Xiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du, Longxu Dou, Haonan Wang, Tianyu Pang, and Michael Qizhe Shieh. 2025a. Noisyrollout: Reinforcing visual reasoning with data augmentation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, and 1 others. 2025b. Seeing but not believing: Probing the disconnect between visual attention and answer correctness in vlms. arXiv preprint arXiv:2510.17771. Zujing Liu, Junwen Pan, Qi She, Yuan Gao, and Guisong Xia. 2025c. On the faithfulness of visual thinking: Measurement and enhancement. arXiv preprint arXiv:2510.23482. Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, KaiWei Chang, Michel Galley, and Jianfeng Gao. 2024. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. In The Twelfth International Conference on Learning Representations. Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. 2021. Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, pages 6774–6786, Online. Association for Computational Linguistics. Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey

Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica. 2025. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. Fanqing Meng, Lingxiao Du, Zongkai Liu, Zhixiang Zhou, Quanfeng Lu, Daocheng Fu, Tiancheng Han, Botian Shi, Wenhai Wang, Junjun He, and 1 others. 2025. Mm-eureka: Exploring the frontiers of multimodal reasoning with rule-based reinforcement learning. arXiv preprint arXiv:2503.07365. Runqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, and Honggang Zhang. 2025. We-math: Does your large multimodal model achieve human-like mathematical reasoning? In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20023–20070, Vienna, Austria. Association for Computational Linguistics. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. 2024. Flashattention-3: Fast and accurate attention with asynchrony and low-precision. Advances in Neural Information Processing Systems, 37:68658–68685. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, and 1 others. 2025. Vlm-r1: A stable and generalizable r1style large vision-language model. arXiv preprint arXiv:2504.07615. Guohao Sun, Hang Hua, Jian Wang, Jiebo Luo, Sohail Dianat, Majid Rabbani, Raghuveer Rao, and Zhiqiang Tao. 2025a. Latent chain-of-thought for visual reasoning. arXiv preprint arXiv:2510.23925. Hai-Long Sun, Zhun Sun, Houwen Peng, and Han-Jia Ye. 2025b. Mitigating visual forgetting via takealong visual conditioning for multi-modal long cot reasoning. arXiv preprint arXiv:2503.13360. Kai Sun, Yushi Bai, Ji Qi, Lei Hou, and Juanzi Li. 2024. MM-MATH: Advancing multimodal math evaluation with process evaluation and fine-grained classification. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1358–1375, Miami, Florida, USA. Association for Computational Linguistics.

Xinyu Tian, Shu Zou, Zhaoyuan Yang, Mengqi He, Fabian Waschkowski, Lukas Wesemann, Peter Tu, and Jing Zhang. 2025. More thought, less accuracy? on the dual nature of reasoning in vision-language models. arXiv preprint arXiv:2509.25848.

Yu Xia, Rui Wang, Xu Liu, Mingyan Li, Tong Yu, Xiang Chen, Julian McAuley, and Shuai Li. 2025. Beyond chain-of-thought: A survey of chain-of-x paradigms for llms. In Proceedings of the 31st International Conference on Computational Linguistics, pages 10795–10809.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems, 30.

Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang. 2024. Logicvista: Multimodal llm logical reasoning benchmark in visual contexts. arXiv preprint arXiv:2407.04973.

Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu, Fangzhen Lin, and Wenhu Chen. 2025a. Vlrethinker: Incentivizing self-reflection of visionlanguage models with reinforcement learning. arXiv preprint arXiv:2504.08837.

An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025a. Qwen3 technical report. arXiv preprint arXiv:2505.09388.

Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Sharon Li, and Neel Joshi. 2024. Is a picture worth a thousand words? delving into spatial reasoning for vision language models. Advances in Neural Information Processing Systems, 37:75392– 75421. Peijie Wang, Zhong-Zhi Li, Fei Yin, Dekang Ran, and Cheng-Lin Liu. 2025b. Mv-math: Evaluating multimodal math reasoning in multi-visual contexts. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 19541–19551. Xiyao Wang, Zhengyuan Yang, Chao Feng, Yongyuan Liang, Yuhang Zhou, Xiaoyu Liu, Ziyi Zang, Ming Li, Chung-Ching Lin, Kevin Lin, and 1 others. 2025c. Vicrit: A verifiable reinforcement learning proxy task for visual perception in vlms. arXiv preprint arXiv:2506.10128. Xiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu, Linjie Li, Chung-Ching Lin, Kevin Lin, Furong Huang, and Lijuan Wang. 2025d. Sota with less: Mcts-guided sample selection for data-efficient visual reasoning self-improvement. arXiv preprint arXiv:2504.07934. Zhenhailong Wang, Xuehang Guo, Sofia Stoica, Haiyang Xu, Hongru Wang, Hyeonjeong Ha, Xiusi Chen, Yangyi Chen, Ming Yan, Fei Huang, and 1 others. 2025e. Perception-aware policy optimization for multimodal reasoning. arXiv preprint arXiv:2507.06448. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824– 24837. Junyi Wu, Weitai Kang, Hao Tang, Yuan Hong, and Yan Yan. 2024. On the faithfulness of vision transformer explanations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10936–10945.

Shuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye, Bin Lin, and Li Yuan. 2025b. Look-back: Implicit visual re-focusing in mllm reasoning. arXiv preprint arXiv:2507.03019. Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu, Bo Zhang, and Wei Chen. 2025c. R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 2376–2385. Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, and Chuang Gan. 2025d. Machine mental imagery: Empower multimodal reasoning with latent visual tokens. arXiv preprint arXiv:2506.17218. Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, YuYue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, LingJun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, and 17 others. 2025. DAPO: An open-source LLM reinforcement learning system at scale. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, and Graham Neubig. 2025. MMMU-pro: A more robust multi-discipline multimodal understanding benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15134–15186, Vienna, Austria. Association for Computational Linguistics. Chi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu, Zhixiong Zeng, Siqi Yang, Peng Shi, Lin Ma, and Jing Zhang. 2025. Perceptual-evidence anchored reinforced learning for multimodal reasoning. arXiv preprint arXiv:2511.18437. Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024a. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8):5625–5644.

Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Yu Qiao, Peng Gao, and Hongsheng Li. 2024b. Multi-modal llm truly see the diagrams in visual math problems? In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part VIII, page 169–186, Berlin, Heidelberg. SpringerVerlag. Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, and 1 others. 2025. Group sequence policy optimization. arXiv preprint arXiv:2507.18071. Weihong Zhong, Xiaocheng Feng, Liang Zhao, Qiming Li, Lei Huang, Yuxuan Gu, Weitao Ma, Yuan Xu, and Bing Qin. 2024. Investigating and mitigating the multimodal hallucination snowballing in large visionlanguage models. arXiv preprint arXiv:2407.00569. Yue Zhou, Litong Feng, Mengcheng Lan, Yiping Ke, Xue Jiang, and Wayne Zhang. Geomath: A benchmark for multimodal mathematical reasoning in remote sensing. Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, and 1 others. 2025. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479.

A

Experimental Settings

A.1

Training Datasets

We mainly conduct our training experiments on the ViRL39K (Wang et al., 2025a) dataset, a verifiable instruction-tuning benchmark designed for vision-language reasoning. This dataset comprises approximately 39,000 high-quality samples synthesized from proprietary collections and various upstream datasets, including Llava-OneVision (Li et al., 2025b), R1-OneVision (Yang et al., 2025c), MM-Eureka (Meng et al., 2025), MM-Math (Sun et al., 2024), M3CoT (Chen et al., 2024b), DeepScaleR (Luo et al., 2025), and MV-Math (Wang et al., 2025b). ViRL39K is distinguished by a rigorous filtering pipeline that removes unverifiable queries, thereby ensuring extensive coverage of topics ranging from general chart interpretation to complex STEM problem-solving. A.2

Baselines

• ThinkLite-VL (Wang et al., 2025d). ThinkLiteVL bypasses knowledge distillation in favor of an MCTS-guided selection strategy that curates a compact, high-quality dataset based on sample difficulty. By applying reinforcement fine-tuning to these challenging instances, the method attains superior visual reasoning results while reducing data requirements by an order of magnitude. • VL-Rethinker (Wang et al., 2025a). VLRethinker enhances VLMs’ slow-thinking capabilities through a distillation-free reinforcement learning framework that integrates Selective Sample Replay (SSR) to stabilize training and Forced Rethinking to incentivize self-reflection. By rehearsing high-value experiences and triggering explicit verification processes, the method effectively mitigates GRPO limitations and fosters the internalization of deliberate reasoning patterns. • MM-Eureka (Meng et al., 2025). MMEureka advances multimodal reasoning by combining the MMK12 dataset with a Qwen2.5-VL-based GRPO pipeline that employs online filtering to ensure gradient efficacy. To address training instability in large models, it utilizes a two-stage strategy involving initial training on MMK12 without KL divergence followed by fine-tuning on Geo3k with KL regularization. • NoisyRollout (Liu et al., 2025a). NoisyRollout enhances visual reasoning and robustness by

computing GRPO advantages from hybrid clean and distorted trajectories, restricting policy optimization to uncorrupted inputs. Furthermore, it incorporates a noise annealing schedule to gradually diminish distortion, thereby effectively balancing exploration and stability. • PAPO (Wang et al., 2025e). PAPO addresses perception errors by introducing an Implicit Perception Loss that maximizes the KL divergence between policy outputs on original and masked inputs to enforce visual grounding. To ensure training stability, it employs a Double Entropy Loss regularization, offering a model-agnostic framework that jointly optimizes perception and reasoning without requiring additional annotations or teacher models. • VPPO (Huang et al., 2025a). VPPO integrates visual perception into RL by quantifying tokenlevel visual dependency via KL divergence between policy distributions on original and perturbed image inputs. Based on this metric, it employs advantage modulation and sparse gradient masking to concentrate optimization exclusively on pivotal visual tokens. A.3

Evaluation Benchmarks

Mathematical & Geometric Reasoning:

Parameter

Configuration

General Settings Rollout Number Learning Rate Global Batch Size Rollout Batch Size Val Batch Size Max Prompt Length Max Response Length Reward GPU Usage

8 1e-6 128 512 1024 4096 2048 Binary Accuracy 8×H20, 96G Memory

Qwen2.5-VL-3B/7B on ViRL39K (39K samples) Training Episodes Total Optimization Steps

2 150

Qwen2.5-VL-7B on Geo3K (2.1K samples) Training Episodes Total Optimization Steps

15 60

Qwen2.5-VL-7B on MMK12 (6.4K samples) Training Episodes Total Optimization Steps

12 120

Qwen2.5-VL-32B on ViRL39K (39K samples) Training Episodes Total Optimization Steps GPU Usage

2 150 32×H20, 96G Memory

Table 7: Experimental hyperparameter configurations of our VGPO across 3B, 7B, and 32B-based backbones and different scalable training datasets.

• MathVista (Lu et al., 2024). MathVista is designed to integrate challenges across diverse mathematical and visual domains, necessitating both fine-grained visual perception and compositional reasoning. The benchmark consists of 6,141 examples amassed from a wide array of sources, including 28 existing multimodal mathematics datasets and 3 constructed datasets.

6,500 visual tasks spanning 67 distinct knowledge concepts. By employing a decompositional methodology that reduces composite problems into atomic sub-tasks, the benchmark facilitates a granular four-dimensional evaluation distinguishing among Insufficient Knowledge, Inadequate Generalization, Complete Mastery, and Rote Memorization.

• MathVerse (Zhang et al., 2024b). MathVerse is designed to mitigate the textual bias in existing benchmarks by curating 2,612 high-quality problems that necessitate genuine visual interpretation. Each problem is manually reformulated into six distinct versions with varying multimodal information density, resulting in a comprehensive corpus of approximately 15,000 test samples.

• MMK12 (Meng et al., 2025). MMK12 benchmark constitutes a comprehensive evaluation framework designed to assess multimodal reasoning capabilities at the K-12 educational level. Comprising 2,000 high-quality instances, the benchmark is stratified across four scientific disciplines, specifically Mathematics, Physics, Chemistry, and Biology, with 500 multimodal multiple-choice questions dedicated to each field.

• We-Math (Qiao et al., 2025). We-Math transcends conventional accuracy metrics to scrutinize the underlying problem-solving mechanisms of LMMs through a hierarchical framework of

• GeoMath (Zhou et al.). Designed to address scarcity of reasoning data in Earth observation, GeoMath comprises 3,773 high-quality queries

derived from aerial imagery, encompassing six distinct mathematical subjects across 20 subtopics. The visual data are acquired via proprietary drone flights capturing a wide range of altitudes and viewing angles. • Geometry3K (Lu et al., 2021). Geometry3K comprises 3,002 geometry problems enriched with dense annotations in formal language, serving as a challenging task for abstract problem understanding and symbolic reasoning, requiring axiomatic knowledge. Vision-Dependent Multimodal Reasoning: • LogicVista (Xiao et al., 2024). LogicVista is an evaluation benchmark designed to assess the integrated logical reasoning abilities of VLMs in visual contexts. Addressing the limitations of prior work in systematically evaluating logical proficiency. The dataset consists of 448 multiplechoice questions, each densely annotated with the correct answer and a human-written explanation. • Super-CLEVR (Li et al., 2023). Super-CLEVR constitutes a virtual benchmark designed to rigorously assess the out-of-distribution robustness and domain generalization of VQA models. Through controllable data synthesis, it decouples inherent multi-modal factors to facilitate the isolated analysis of four specific domain shifts: visual complexity, question redundancy, concept distribution, and concept compositionality. • MMMU-Pro (Yue et al., 2025). MMMU-Pro advances the MMMU benchmark by employing a stringent curation process that eliminates textsolvable queries and expands candidate options to rigorously evaluate intrinsic multimodal reasoning. Additionally, it introduces a vision-only paradigm where inquiries are embedded within images, thereby necessitating integrated visual and textual interpretation.

B

More Experimental Results

B.1 Detailed Results of Re-weighting Strategy Table 8 provides the whole performance breakdown across all datasets, complementing the aggregated results in the main text. Effectiveness of Intra-trajectory Strategy. This mechanism significantly enhances general mathematical reasoning. It yields substantial gains on

MathVista (+5.4%) and MMK12 (+4.7%), demonstrating that reinforcing self-consistency within individual reasoning paths is critical to re-focus visual activation for multimodal reasoning Impact of Inter-trajectory Strategy. The Intertrajectory module excels in geometry and visual counting tasks. It achieves a peak score of 54.9% on GeoMath and matches the top performance on the Counting dataset. These results suggest that cross-path comparison effectively mitigates visual hallucinations in object-centric scenarios. Synergy of the Two Strategies. Combining both modules yields the most robust performance across mathematical and vision-dependent categories. Notably on WeMath, the integrated approach reverses a slight decline caused by Intra-module alone, achieving a superior score of 72.5%. Furthermore, the combined strategy secures top results on challenging benchmarks like LogicVista and MMMUPro. This confirms a synergistic effect, where the Intra-module refines logical consistency while the Inter-module broadens the reasoning scope. B.2

Detailed Results of Advantage Shaping Methods

Table 9 presents a fine-grained comparison of VGPO against established advantage shaping strategies. Consistent with the main results, VGPO exhibits superior performance across both mathematical and vision-centric benchmarks. Robustness in General Reasoning. In contrast to generic regularization methods like Entropy and KL, which rely on implicit constraints on policy divergence, VGPO explicitly leverages visual feedback. This mechanism translates into substantial gains in General Mathematical & Geometric Reasoning. VGPO secures the highest scores on 4 out of 6 datasets, establishing significant leads on MathVista and MathVerse. This demonstrates that visually-guided optimization not only enhances reasoning stability but does so while preserving the general problem-solving proficiency. Efficacy in Visual Perception. Our detailed analysis highlights the pivotal role of VGPO in Visiondependent Multimodal Reasoning. The performance gap is particularly evident in the Counting dataset, a task requiring rigorous visual grounding. VGPO achieves 95.5% accuracy, surpassing the Entropy and KL baselines by a significant margin. Similar dominance is observed on LogicVista. These results validate our hypothesis that explicitly reinforcing visual attention through VGPO is far

General Mathematical & Geometric Reasoning

Models

Vision-dependent Multimodal Reasoning

Avg-Math

MathVista MathVerse WeMath MMK12 GeoMath Geo3k DAPO Baseline + Intra-trajectory + Inter-trajectory + Intra- & Inter-trajectory

68.7 74.1 72.1 74.1

69.6 71.1 68.8 71.6

70.8 69.8 69.9 72.5

77.0 81.7 80.7 81.5

51.3 54.3 54.9 54.3

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV

45.6 45.6 45.3 45.8

63.8 66.1 65.3 66.6

47.4 48.1 47.9 49.4

85.5 95.0 95.5 95.5

39.0 40.2 40.2 40.5

66.6 66.5 64.6 67.6

59.6 62.5 62.0 63.3

Table 8: The detailed ablation study on the impact of Intra- and Inter-trajectory re-weighting strategies. General Mathematical & Geometric Reasoning

Models

MathVista MathVerse WeMath MMK12 GeoMath Geo3k Qwen2.5-VL-7B + DAPO + DAPO w/ Entropy + DAPO w/ KLperception + VGPO (Ours)

68.5 68.7 70.2 70.5 74.1

40.2 69.6 72.1 71.1 71.6

47.8 70.8 70.8 70.6 72.5

49.4 77.0 80.9 81.8 81.5

51.2 51.3 53.9 53.2 54.3

Vision-dependent Multimodal Reasoning

Avg-Math

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV

42.9 45.6 45.6 46.9 45.8

50.0 63.8 65.6 65.7 66.6

45.2 47.4 48.3 48.8 49.4

76.5 85.5 90.5 88.5 95.5

36.4 39.0 40.1 40.2 40.5

36.6 66.6 68.7 67.6 67.6

48.7 59.6 61.9 61.3 63.3

Table 9: The detailed comparisons with other advantage shaping strategies, including the Entropy-based (Cheng et al., 2025) and KL-based (Huang et al., 2025a) methods.

0.5 0.4 0.3 0

25

50

75

100

Training Steps

125

(a) Training rewards.

150

DAPO

0.6 0.5 0.4 0

25

50

75

100

Training Steps

125

150

(b) Validation accuracy.

1.0 0.9 0.9 0.8 0.8 0.7 0.7 0.6

VGPO (Ours)

0

25

50

75

DAPO

100

Training Steps

125

(a) Training rewards.

Average Accuracy

0.6

VGPO (Ours) 0.7

Average Reward

Average Reward

DAPO

Average Accuracy

VGPO (Ours)

0.7

150

VGPO (Ours)

0.9

DAPO

0.8 0.7 0.6 0.5

0

25

50

75

100

Training Steps

125

150

(b) Validation accuracy.

Figure 7: Training dynamics of Qwen2.5-VL-3B: (a) training rewards and (b) validation accuracy on MMK12 (Meng et al., 2025) across DAPO (Yu et al., 2025), and our VGPO.

Figure 8: Training dynamics of Qwen2.5-VL-32B: (a) training rewards and (b) validation accuracy on MMK12 (Meng et al., 2025) across DAPO (Yu et al., 2025), and our VGPO.

more effective for complex multimodal tasks than relying on implicit regularization techniques.

suggests that our visual-guided strategy effectively reduces variance and stabilizes the optimization process for large-scale models. Consistent with the rewards, Figure 8(b) shows that VGPO consistently outperforms DAPO in validation accuracy, establishing a significant margin by the end of training.

B.3

Analysis of Training Dynamics

To further investigate the convergence, stability, and scalability of our VGPO, we visualize the training dynamics of training rewards and validation accuracy on MMK12 dataset for both Qwen2.5-VL3B (Figure 7) and Qwen2.5-VL-32B (Figure 8). Dynamics on Qwen2.5-VL-3B. As shown in Figure 7, VGPO demonstrates a clear performance advantage. In terms of training rewards, VGPO maintains a consistently higher reward trajectory compared to DAPO throughout the training steps. More importantly, regarding validation accuracy, VGPO achieves a higher final accuracy, which indicates superior generalization capabilities. Dynamics on Qwen2.5-VL-32B. The advantages of VGPO are even more pronounced on the larger 32B model, particularly regarding stability. Figure 8(a) highlights that DAPO suffers from severe oscillation during training, whereas VGPO exhibits a smooth and steady increase in rewards. This

B.4

Detailed Results of Different Compensation Schedules

Table 10 provides the detailed performance breakdown across all datasets, complementing the aggregated results in the main text. We adopt the Linear Compensation Strategy since it aligns well with the progressive, continuous, and nearly linear decay of visual attention observed in our empirical analysis across four visualdependent tasks (LogicVista, CLEVR Counting, MMMU-Pro, MathVerse-V). To explore potential alternatives, we compare our Linear strategy against Step-Function and Exponential schedules. As shown in Table 10, all these schedules exhibit higher performance compared to the baseline. However, the Linear strategy consistently achieves

General Mathematical & Geometric Reasoning

Compensation Schedules Baseline (DAPO) w/ Step-Function w/ Exponential w/ Linear (Ours)

68.7 69.1 70.6 74.1

69.6 69.2 69.9 71.6

70.8 69.4 69.3 72.5

77.0 80.3 81.4 81.5

51.3 52.3 54.2 54.3

Vision-dependent Multimodal Reasoning

Avg-Math

MathVista MathVerse WeMath MMK12 GeoMath Geo3k

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV

45.6 47.6 44.9 45.8

63.8 64.7 65.1 66.6

47.4 47.5 49.0 49.4

85.5 91.5 90.5 95.5

39.0 39.2 38.8 40.5

66.6 64.7 65.5 67.6

59.6 60.7 61.0 63.3

Table 10: Results of General Mathematical & Geometric Reasoning and Vision-dependent Multimodal Reasoning tasks with different compensation schedules. Note: For the Exponential schedule, we test powers of 1.0 and 2.0 (reporting the best result from 2.0). We test 1.0 for the Step-wise schedule. General Mathematical & Geometric Reasoning

Models

Avg-Math

MathVista MathVerse WeMath MMK12 GeoMath Geo3k

Vision-dependent Multimodal Reasoning

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV

Baseline (DAPO, Qwen2.5-VL-7B) +VGPO (Ours)

68.7 74.1

69.6 71.6

70.8 72.5

77.0 81.5

51.3 54.3

45.6 45.8

63.8 66.6△4.4%

47.4 49.4

85.5 95.5

39.0 40.5

66.6 67.6

59.6 63.3△6.2%

Baseline (DAPO, Qwen2.5-VL-3B) +VGPO (Ours)

63.9 65.0

57.1 61.4

62.9 62.1

64.1 71.6

47.6 50.2

36.3 36.1

55.3 57.7△4.3%

43.0 45.9

65.5 78.0

31.5 32.5

53.1 58.1

48.3 53.6△11.0%

Baseline (DAPO, Qwen2-VL-2B) +VGPO (Ours)

50.8 51.1

33.2 37.8

36.3 42.5

42.3 44.2

34.8 37.0

8.7 34.4 16.8 38.2△11.1%

30.0 30.2

66.5 83.5

19.8 21.9

31.6 36.4

37.0 43.0△16.2%

Table 11: Performance comparison of stronger to weaker visual encoders on General Mathematical & Geometric Reasoning and Vision-dependent Multimodal Reasoning tasks.

the best performance. We attribute these results to the following reasons: • Exponential schedule (over-correction): This schedule puts most of the bonus on the very last tokens, which often correspond to calculation or formatting rather than visual grounding. Overweighting those tokens dilutes the benefit of compensating truly visual-dependent steps and can hurt accuracy compared to the more balanced linear schedule. • Step-Function schedule (training instability): This schedule makes position compensation jump from zero to full compensation at a single threshold. However, it does not match our observed progressive forgetting: visual attention decays gradually over the trajectory, so we use a linear schedule that increases compensation smoothly rather than a single step. B.5

Robustness on Weaker Visual Encoders

To investigate the robustness of VGPO across different visual encoder capacities, we evaluate our method using Qwen2-VL-2B-Instruct (Bai et al., 2025b), which employs a significantly smaller and weaker visual encoder compared to the Qwen2.5VL series. As shown in Table 11, despite the inherently constrained perceptual capacity of the 2B baseline, applying VGPO achieves even better relative improvements (+11.1% on Math and +16.2% on vision-dependent tasks). This empirical evidence confirms that VGPO demonstrates strong

generalizability and robustness, effectively resolving the policy-level bottleneck of temporal visual forgetting regardless of the base visual encoder’s strength. B.6

Detailed Results of Full-trajectory Compensation

Table 12 provides the detailed performance breakdown across all datasets, complementing the aggregated results in the main text. As shown in Table 12, we find that the full-trajectory compensation strategy significantly drops performance on the majority of benchmarks. Moreover, the Late/Early Visual Activation Ratio exhibits a lower value (0.8591) than our VGPO (1.0696). This indicates that excessive elevation of early-trajectory compensation often overemphasizes early visual attention and leads to higher visual forgetting with lower performance. B.7

Prompt Template

Template for Multi-modal Reasoning SYSTEM You are a helpful assistant. USER <question> You first think through the reasoning process as an internal monologue, enclosed within <think> </think> tags. Then, provide your final answer enclosed within \boxed{}.

Compensation Schedules Baseline (DAPO) Full-trajectory Late-trajectory (Ours)

General Mathematical & Geometric Reasoning

Avg-Math

MathVista MathVerse WeMath MMK12 GeoMath Geo3k 68.7 71.4 74.1

69.6 50.3 71.6

70.8 53.8 72.5

77.0 60.1 81.5

51.3 44.9 54.3

45.6 37.3 45.8

Vision-dependent Multimodal Reasoning

Avg-Vision

LogicVista Counting MMMU-Pro MathVerseV 63.8 53.0 66.6

47.4 40.5 49.4

85.5 92.0 95.5

39.0 36.4 40.5

66.6 47.8 67.6

59.6 54.2 63.3

Table 12: Performance comparison of Full-trajectory compensation versus our Late-trajectory compensation strategy.

C

Discussion about Visual Focus Score vs. Attention Weights

To validate the theoretical grounding of our proposed Visual Focus Score, we conduct an empirical correlation analysis between the Visual Focus Score and actual attention weights. Empirical validation of correlation with Attention Weights. We select the validation dataset (MMK12-val (Meng et al., 2025)) for this analysis. At each generation step of each sample: • For the Visual Focus Score, we compute the cosine similarity between the final hidden state of the t-th generated token and the mean-pooled visual prototype. • For the Attention Weights, we compute the sum of the attention probabilities from the t-th generated token to all image token positions, averaged across all attention heads in the final layer. We then calculate the Pearson correlation coefficient between these two step-wise sequences (t = 1, 2, . . . , T ). We observe a strong positive correlation (r ≈ 0.67), confirming that our similaritybased Visual Focus Score closely aligns with the actual visual attention mechanism. Why use hidden-state similarity rather than actual attention weights? During implementation, extracting actual attention weights requires setting output_attentions=True. This causes the model to fall back to the less efficient eager attention implementation rather than Flash Attention (Shah et al., 2024), losing Flash Attention’s speed and memory benefits (typically incurring a 20%–30% additional computational cost). In contrast, our hidden-state-based implementation only requires output_hidden_states=True, which adds approximately 5%–10% overhead and remains fully compatible with Flash Attention, making it highly efficient for large-scale training.

D

The Use of LLMs

This manuscript utilizes LLMs exclusively for the purpose of linguistic refinement and stylistic improvement to enhance clarity. The LLMs serve

strictly as a writing aid and contribute no intellectual content to the research conceptualization, methodological design, data execution, analysis, or interpretation. These core scientific components remain the sole work of the authors. All suggestions provided by the LLMs are critically reviewed and selectively adopted to ensure the accuracy and integrity of the text. The authors retain full accountability for the content and findings presented herein, and the LLM does not qualify for authorship or contributorship.

E

Failure Case Analysis

To provide deeper insights into the boundaries of our proposed VGPO, we present both qualitative and quantitative analyses of failure cases where increased visual focus does not translate to improved accuracy. Qualitative insight (hallucination): Our error analysis reveals that if the base visual encoder fundamentally misinterprets a feature (e.g., misrepresenting a blurry “6” as an “8” due to low resolution), VGPO can actually reinforce this error. By incentivizing the LLM to strongly attend to a flawed visual embedding, the model produces confidently incorrect answers, rather than falling back on common-sense priors that might have occasionally guessed correctly. Quantitative insight: As shown in our original analysis in Figure 2(c), while correct samples generally show higher visual accumulation, there is a subset of incorrect samples that also exhibit high visual focus scores. This quantitatively demonstrates that “looking closely” is a necessary but not sufficient condition for correctness. The model must also possess the correct world knowledge and reasoning logic to interpret the visual cues properly. Here we provide two typical failure cases to illustrate these points.

Failure Case 1: Visual Perception Error (GeoMath, id: geomath-2884)

Question: <image> How many vertical bars are visible in the image, and what is the total number of elements (cow and bars) present in the image? Ground Truth: 8 Model Prediction: To determine the number of vertical bars and the total number of elements in the image, let’s break it down step by step. 1. Count the vertical bars: The image shows a series of vertical bars that form a fence. By visually inspecting the image, it appears there are at least 10 vertical bars. (visual counting error) 2. Count the elements: There is one cow visible in the image. There are 10 vertical bars. 3. Calculate the total number of elements: Total elements = 10 + 1 = 11. Therefore, the total number of elements is: \boxed{11}. Failure Analysis: The model’s prediction fails due to an Object Counting Error. It correctly identifies the task of counting bars and adding the cow, but falters on the initial visual perception step, incorrectly counting 10 vertical bars instead of the actual visible bars (7, not been occluded by other bars or the cow). While its mathematical logic (10 + 1 = 11) is internally consistent, the entire reasoning is built on a flawed premise. The root cause is a failure in basic visual perception, not a flaw in logical deduction.

Failure Case 2: Logical Reasoning Error (MathVerse-V, id: 2)

Question: <image>. Ground Truth: D Model Prediction: To determine the size of angle CED, we start by analyzing the given information. First, we know that ∠B = 60◦ and ∠A = 80◦ . Therefore, we can find ∠C using the fact that the sum of the angles in a triangle is 180◦ : ∠C = 180◦ − ∠A − ∠B = 180◦ − 80◦ − 60◦ = 40◦ Since DE ∥ BC, by the corresponding angles property (wrong geometric theorem), we have: ∠CED = ∠C = 40◦ Therefore, the size of angle CED is 40◦ . Thus, the correct answer is: \boxed{A}. Failure Analysis: The model correctly perceives the geometric setup and calculates ∠C as 40◦ . The error stems from flawed knowledge: it misapplies the parallel line theorem, incorrectly equating ∠CED with ∠C. It confuses supplementary consecutive interior angles with equal corresponding angles, showing that while the model “looks” correct, it lacks the proper logic to interpret what it sees.

Record · ID 6011 · SHA-256 6cf7a917bf100cbd
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.