1 of 18

1

DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO

Jinyoung Park1, Jeehye Na1, Jinyoung Kim2,  Hyunwoo J. Kim1

1Korea Advanced Institute of Science and Technology

2Korea University

MLV Lab

2 of 18

2

Preliminary

Reinforcement fine-tuning

  • SFT learns to imitate
    • It mimics outputs without exploring better ones

  • No feedback in SFT
    • Fixed targets prevent learning from other high quality of reasoning paths.

  • RL enables active learning
    • It explore multiple reasoning paths with feedback

Luong, Trung Quoc, et al. "Reft: Reasoning with reinforced fine-tuning." ACL 2024.

MLV Lab

3 of 18

3

Preliminary

Group Relative Preference Optimization

Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).

MLV Lab

4 of 18

4

Preliminary

Group Relative Preference Optimization

  1. Generate multiple responses

Sample multiple responses from the policy for the same input

Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).

MLV Lab

5 of 18

5

Preliminary

Group Relative Preference Optimization

Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).

2. Evaluate each response

Use a reward model or rule-based reward functions �(e.g. format, task-specific metrics)

MLV Lab

6 of 18

6

Preliminary

Group Relative Preference Optimization

Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).

3. Normalize Rewards into advantage

Standardize rewards to have zero mean and unit variance

MLV Lab

7 of 18

7

Preliminary

Group Relative Preference Optimization

Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).

4. Optimize the model by maximizing the objective function below

Update the policy model to increase the probability of high reward responses

MLV Lab

8 of 18

8

Preliminary

Group Relative Preference Optimization

Limitations in GRPO

1. Reliance on Safeguards (Clipping, Min)

MLV Lab

9 of 18

9

Motivation

1. Reliance on Safeguards (Clipping, Min)

Safeguards ensure stable learning through conservative updates, but they vanish the gradients if their conditions are satisfied.

Schulman, John, et al. "Proximal policy optimization algorithms." arXiv preprint arXiv:1707.06347 (2017).

  • Uses heuristic clipping and min operations to prevent drastic change

MLV Lab

10 of 18

10

Preliminary

Group Relative Preference Optimization

Limitations in GRPO

2. Vanishing Advantage

MLV Lab

11 of 18

11

Motivation

2. Vanishing Advantage

Wang, Haozhe, et al. "Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning." arXiv preprint arXiv:2504.08837 (2025).

  • Advantage becomes zero when rewards are same within the group
  • Results in no learning signal

Too hard!

Too easy!

#Effective queries decrease!!

MLV Lab

12 of 18

12

Method

Sol 1. Reg-GRPO (Regressive GRPO)

  • Reframes advantage maximization as advantage regression
  • GRPO: Maximizes the probability of good response

  • Reg-GRPO: Regresses to align the implicit reward with advantages derived from actual reward

MLV Lab

13 of 18

13

Method

Sol 2. Difficulty-aware Data Augmentation

 

MLV Lab

14 of 18

14

Method

Difficulty-aware Data Augmentation

 

MLV Lab

15 of 18

15

Method

Difficulty-aware Data Augmentation

 

MLV Lab

16 of 18

16

Experiment

Quantitative Results

MLV Lab

17 of 18

Qualitative Results

Question: Considering the progress shown in the video and my current observation shown in the image, what action should I take next in order to clean the house?� �A. collect dirt B. put down broom C. roll cable around vacuum cleaner D. remove charger� �Output the thinking process in <think> </think> and final answer in <answer> </answer> tags, i.e., <think> reasoning process here </think><answer> answer here </answer>.

GRPO: <think> The person in the video is holding a broom, which suggests that they are in the process of cleaning the house. The presence of the broom indicates that the next logical step would be to continue using it to collect dirt and debris. The other options do not seem relevant to the current activity shown in the video.</think> <answer>A. collect dirt</answer>

DeepVideo-R1(Ours): <think> The person in the video is currently cleaning the floor with a broom. The next logical step would be to put down the broom after finishing cleaning the floor. </think> <answer> B. put down broom</answer>

MLV Lab

18 of 18

18

Summary

  • Regressive GRPO for stable learning without heuristic safeguards
    • Avoids heuristic safeguards by directly predicting advantage values

  • Difficulty-aware data augmentation
    • Dynamically adjusts sample difficulty to boost learning signal

  • Improved video reasoning
    • Achieves stronger in/out-of-distribution performance

MLV Lab