Training Jan 2026 · #43 most cited · 114 citations

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Shih-Yang Liu et al.

arXiv:2601.05242

In short

When RL for LLMs combines several rewards (say accuracy, format and length), the popular GRPO method normalises them together and can collapse distinct combinations into the same training signal. GDPO normalises each reward separately, and beats GRPO on tool calling, maths and coding.

Why it matters

Multi-objective RL is the norm in post-training; this fixes a subtle flaw in the default algorithm.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 10 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Reinforcement Learning · read first ✓ understood

    Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.

  3. Reward · read first ✓ understood

    A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.

  4. Agent ✓ understood

    In RL, the learner or decision-maker that takes actions in an environment to maximize cumulative reward.

  5. Policy ✓ understood

    A strategy or mapping from states to actions that defines the agent's behavior in reinforcement learning.

  6. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  7. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  8. Loss Function ✓ understood

    A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

  9. Gradient Descent ✓ understood

    An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.

  10. Policy Gradient · read first ✓ understood

    RL methods that directly optimize the policy by computing gradients of expected reward with respect to policy parameters.

In the frontier

Rank
#43 of 100
Citations
114
as of Aug 9, 2026
Published
Jan 2026

Topics: RL for reasoning

Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026

Nearby papers

Summary in our own words; read the paper for the details. ← All papers