GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Shih-Yang Liu et al.
arXiv:2601.05242
In short
When RL for LLMs combines several rewards (say accuracy, format and length), the popular GRPO method normalises them together and can collapse distinct combinations into the same training signal. GDPO normalises each reward separately, and beats GRPO on tool calling, maths and coding.
Why it matters
Multi-objective RL is the norm in post-training; this fixes a subtle flaw in the default algorithm.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 10 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Reinforcement Learning · read first ✓ understood
Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.
- Reward · read first ✓ understood
A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.
- Agent ✓ understood
In RL, the learner or decision-maker that takes actions in an environment to maximize cumulative reward.
- Policy ✓ understood
A strategy or mapping from states to actions that defines the agent's behavior in reinforcement learning.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Policy Gradient · read first ✓ understood
RL methods that directly optimize the policy by computing gradients of expected reward with respect to policy parameters.
In the frontier
- Rank
- #43 of 100
- Citations
- 114
- as of Aug 9, 2026
- Published
- Jan 2026
Topics: RL for reasoning
Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026