Reinforcement learning (RL) is learning from consequences. Nobody tells the learner the correct answer. An agent tries actions, the environment responds, and a reward signal says how well things went. Over many attempts, the agent works out which actions pay off in which situations.
That sets it apart from supervised learning, where every example comes with the right label. In RL the feedback is a score, not an answer, and it often arrives late: in chess, you only learn whether a move was good when the game ends dozens of moves later. Deciding which earlier choices deserve the credit is a large part of the problem.
The core ideas
- Policy: the agent’s strategy, a mapping from situations to actions. It’s what training produces.
- Value function: an estimate of how much total future reward a situation or action is worth. It lets the agent accept a small cost now for a bigger payoff later.
- Exploration vs exploitation: repeat what has worked, or try something new that might work better. Too little exploration and the agent settles into a mediocre habit.
Sutton and Barto’s textbook is the standard account of these ideas, formalized as a Markov decision process: states, actions, rewards and transitions between states.
Where it shows up
Games made RL famous because they come with a built-in score and unlimited practice. In 2013, DeepMind trained a single convolutional network with RL to play Atari 2600 games straight from the screen pixels; with no game-specific changes, it beat all previous methods on six of seven games and a human expert on three.
RL is also a late stage in training most chat models. In RLHF, people rank a model’s answers, a reward model learns to predict those rankings, and RL tunes the language model to score well on it. Robotics, chip layout and recommendation systems use RL too, wherever success can be scored but the right move can’t be written down in advance.
The catch
RL is hungry for experience. Agents can need millions of attempts, which is cheap in a simulator and slow, expensive or dangerous in the physical world. And an agent optimizes the reward you wrote, not the goal you meant: if a loophole scores points, it will find and exploit it. Designing a reward that can’t be gamed is often harder than the learning itself.