Landmark 2 stops to get here · leads to 3
Reward
A scalar feedback signal indicating how good an action was, used to train reinforcement learning agents.
Your route here
2 stops · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Reinforcement Learning ✓ understood
Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.
- Reward · you are here ✓ understood
Picture it
- 01 Agent acts Takes action a in state s
- 02 Scalar reward r One number: how good that step was
- 03 Return G Sum of future rewards, discounted by γ
- 04 Policy update Shift toward actions with higher return
Where it sits
Before this
Reinforcement Learning Reward
Explore nearby
Agents & RL Value Function A function estimating expected cumulative reward from a state (state-value) or state-action pair (action-value/Q-value). Agents & RL Policy A strategy or mapping from states to actions that defines the agent's behavior in reinforcement learning. Agents & RL Environment In RL, the world the agent interacts with, providing states, accepting actions, and returning rewards. Language & LLMs RLHF Reinforcement Learning from Human Feedback - training models using human preferences to align behavior with human values. Agents & RL Inverse Reinforcement Learning Learning reward functions from expert demonstrations, inferring what is being optimized.
In the research
All papers →5 papers that build on Reward .