Landmark 2 stops to get here · leads to 3

Reward

A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.

Your route here

2 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Reinforcement Learning ✓ understood

    Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.

  3. Reward · you are here ✓ understood

Picture it

  1. 01 Agent acts Takes action a in state s
  2. 02 Scalar reward r One number: how good that step was
  3. 03 Return G Sum of future rewards, discounted by γ
  4. 04 Policy update Shift toward actions with higher return
Notice that the reward is just one number per step; the agent learns by maximizing the sum of rewards over time, not any single one.

Where it sits

Explore nearby

In the research

All papers →

5 papers that build on Reward .

Frontier · Jan 2026 · 114 citations GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Multi-objective RL is the norm in post-training; this fixes a subtle flaw in the default algorithm. Frontier · Aug 2025 · 112 citations On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification A tiny, theory-backed change to the most common training step in the stack. Frontier · Aug 2025 · 90 citations Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning More stable RL for image generators, plus a better yardstick for them. Frontier · Nov 2025 · 75 citations DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research It shows how to use RL where there is no verifiable answer. Frontier · Nov 2025 · 60 citations DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning Self-verification is a path to reasoning where no answer key exists.