Landmark 1 stop to get here · leads to 15

Reinforcement Learning

Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.

Your route here

1 stop · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Reinforcement Learning · you are here ✓ understood

Picture it

Observe stateThe agent sees the environment's…Choose actionIts policy π picks an action aEnvironment respondsReturns a reward r and the next…Update policyFavor actions that led to more…Maximize total reward
  1. 01 Observe state The agent sees the environment's current state s
  2. 02 Choose action Its policy π picks an action a
  3. 03 Environment responds Returns a reward r and the next state s′
  4. 04 Update policy Favor actions that led to more total reward

↺ back to 01 · Maximize total reward

Each pass around the loop is one interaction; no one labels the right answer, the agent learns it from rewards over many passes.

Reinforcement learning (RL) is learning from consequences. Nobody tells the learner the correct answer. An agent tries actions, the environment responds, and a reward signal says how well things went. Over many attempts, the agent works out which actions pay off in which situations.

That sets it apart from supervised learning, where every example comes with the right label. In RL the feedback is a score, not an answer, and it often arrives late: in chess, you only learn whether a move was good when the game ends dozens of moves later. Deciding which earlier choices deserve the credit is a large part of the problem.

The core ideas

  • Policy: the agent’s strategy, a mapping from situations to actions. It’s what training produces.
  • Value function: an estimate of how much total future reward a situation or action is worth. It lets the agent accept a small cost now for a bigger payoff later.
  • Exploration vs exploitation: repeat what has worked, or try something new that might work better. Too little exploration and the agent settles into a mediocre habit.

Sutton and Barto’s textbook is the standard account of these ideas, formalized as a Markov decision process: states, actions, rewards and transitions between states.

Where it shows up

Games made RL famous because they come with a built-in score and unlimited practice. In 2013, DeepMind trained a single convolutional network with RL to play Atari 2600 games straight from the screen pixels; with no game-specific changes, it beat all previous methods on six of seven games and a human expert on three.

RL is also a late stage in training most chat models. In RLHF, people rank a model’s answers, a reward model learns to predict those rankings, and RL tunes the language model to score well on it. Robotics, chip layout and recommendation systems use RL too, wherever success can be scored but the right move can’t be written down in advance.

The catch

RL is hungry for experience. Agents can need millions of attempts, which is cheap in a simulator and slow, expensive or dangerous in the physical world. And an agent optimizes the reward you wrote, not the goal you meant: if a loophole scores points, it will find and exploit it. Designing a reward that can’t be gamed is often harder than the learning itself.

Where it sits

Explore nearby

In the research

All papers →

45 papers that build on Reinforcement Learning ; showing 5, canon first.

Canon · 2013 Playing Atari with Deep Reinforcement Learning It started deep reinforcement learning: one learner, raw perception, many tasks. Canon · 2016 Mastering the Game of Go with Deep Neural Networks and Tree Search It proved that learned intuition plus search can master a problem long thought a decade away. Canon · 2017 Proximal Policy Optimization Algorithms PPO became the default RL algorithm, including for RLHF on the first ChatGPT-era models. Canon · 2022 Training Language Models to Follow Instructions with Human Feedback This recipe, RLHF, is what turned raw language models into assistants like ChatGPT. Frontier · Aug 2025 · 1.2K citations InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency It narrows the gap between open and commercial multimodal models on reasoning and agent tasks.

Sources

  1. Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction . 2nd edition, MIT Press, 2018
  2. Mnih et al., "Playing Atari with Deep Reinforcement Learning" . DeepMind, 2013
  3. Ouyang et al., "Training language models to follow instructions with human feedback" . The InstructGPT paper, 2022