Training Aug 2025 · #46 most cited · 112 citations

On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

Yongliang Wu et al.

arXiv:2508.05629

In short

The authors show that ordinary supervised fine-tuning is secretly RL with a badly shaped reward, which explains why it generalises worse. Rescaling each token’s loss by its probability, a one-line change they call Dynamic Fine-Tuning, improves generalisation on maths, code and multimodal tasks.

Why it matters

A tiny, theory-backed change to the most common training step in the stack.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 9 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  5. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  6. Pre-training ✓ understood

    Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

  7. Fine-Tuning · read first ✓ understood

    The process of further training a pre-trained model on a specific dataset to adapt it for a particular task or domain.

  8. Reinforcement Learning · read first ✓ understood

    Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.

  9. Reward · read first ✓ understood

    A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.

In the frontier

Rank
#46 of 100
Citations
112
as of Aug 9, 2026
Published
Aug 2025

Topics: RL for reasoning , Reasoning methods

Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026

Nearby papers

Summary in our own words; read the paper for the details. ← All papers