1
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
Jinyoung Park1, Jeehye Na1, Jinyoung Kim2, Hyunwoo J. Kim1
1Korea Advanced Institute of Science and Technology
2Korea University
MLV Lab
2
Preliminary
Reinforcement fine-tuning
Luong, Trung Quoc, et al. "Reft: Reasoning with reinforced fine-tuning." ACL 2024.
MLV Lab
3
Preliminary
Group Relative Preference Optimization
Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).
MLV Lab
4
Preliminary
Group Relative Preference Optimization
Sample multiple responses from the policy for the same input
Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).
MLV Lab
5
Preliminary
Group Relative Preference Optimization
Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).
2. Evaluate each response
Use a reward model or rule-based reward functions �(e.g. format, task-specific metrics)
MLV Lab
6
Preliminary
Group Relative Preference Optimization
Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).
ㅍ
3. Normalize Rewards into advantage
Standardize rewards to have zero mean and unit variance
MLV Lab
7
Preliminary
Group Relative Preference Optimization
Shao, Zhihong, et al. "Deepseekmath: Pushing the limits of mathematical reasoning in open language models." arXiv preprint arXiv:2402.03300 (2024).
4. Optimize the model by maximizing the objective function below
Update the policy model to increase the probability of high reward responses
MLV Lab
8
Preliminary
Group Relative Preference Optimization
Limitations in GRPO
1. Reliance on Safeguards (Clipping, Min)
MLV Lab
9
Motivation
1. Reliance on Safeguards (Clipping, Min)
Safeguards ensure stable learning through conservative updates, but they vanish the gradients if their conditions are satisfied.
Schulman, John, et al. "Proximal policy optimization algorithms." arXiv preprint arXiv:1707.06347 (2017).
MLV Lab
10
Preliminary
Group Relative Preference Optimization
Limitations in GRPO
2. Vanishing Advantage
MLV Lab
11
Motivation
2. Vanishing Advantage
Wang, Haozhe, et al. "Vl-rethinker: Incentivizing self-reflection of vision-language models with reinforcement learning." arXiv preprint arXiv:2504.08837 (2025).
Too hard!
Too easy!
#Effective queries decrease!!
MLV Lab
12
Method
Sol 1. Reg-GRPO (Regressive GRPO)
�
MLV Lab
13
Method
Sol 2. Difficulty-aware Data Augmentation
MLV Lab
14
Method
Difficulty-aware Data Augmentation
MLV Lab
15
Method
Difficulty-aware Data Augmentation
MLV Lab
16
Experiment
Quantitative Results
MLV Lab
Qualitative Results
Question: Considering the progress shown in the video and my current observation shown in the image, what action should I take next in order to clean the house?� �A. collect dirt B. put down broom C. roll cable around vacuum cleaner D. remove charger� �Output the thinking process in <think> </think> and final answer in <answer> </answer> tags, i.e., <think> reasoning process here </think><answer> answer here </answer>.
GRPO: <think> The person in the video is holding a broom, which suggests that they are in the process of cleaning the house. The presence of the broom indicates that the next logical step would be to continue using it to collect dirt and debris. The other options do not seem relevant to the current activity shown in the video.</think> <answer>A. collect dirt</answer>
DeepVideo-R1(Ours): <think> The person in the video is currently cleaning the floor with a broom. The next logical step would be to put down the broom after finishing cleaning the floor. </think> <answer> B. put down broom</answer>
MLV Lab
18
Summary
MLV Lab