Skip to main content
ScienceVerse
Artificial IntelligenceDifficulty 2-4

Reinforcement learning

Reinforcement learning (RL) is learning by trial and error: an agent takes actions in an environment, receives numerical rewards, and gradually learns which actions lead to the most reward over time. Unlike supervised learning, nobody tells the agent the right action — it must explore, and rewards may arrive long after the actions that earned them. RL produced landmark results in Atari games and Go and is used to fine-tune language models from human feedback, but an agent optimises exactly the reward it is given, which may not be what its designers meant.

Sign in to save this concept.

Temporal-difference control and deep RL

δ=r+γmax⁡a′Q(s′,a′)−Q(s,a)\delta = r + \gamma \max_{a'} Q(s', a') - Q(s, a)

The temporal-difference error δ for a step from state s to s′ with reward r, towards a bootstrapped target that uses the greedy next action a′.

Q(s,a)←Q(s,a)+α δQ(s, a) \leftarrow Q(s, a) + \alpha\, \delta

Q-learning: an off-policy temporal-difference update with step size α.

Q-learning is off-policy: its target uses the greedy next action, so Q approximates the optimal action-value function q* regardless of the exploratory behaviour policy. Deep RL replaces the table with a neural network: the 2015 deep Q-network learned from pixels and score to a level comparable to a professional games tester across 49 Atari games with a single architecture and hyperparameter set.

  • AlphaGo (2016): policy and value networks from supervised learning on expert games plus self-play RL, combined with Monte Carlo tree search; 5–0 against the European champion.
  • RLHF (2022): human preference rankings of model outputs are used to provide the reward signal for fine-tuning language models.
  • Safety: wrong objectives produce side effects and reward hacking (Amodei et al., preprint).
Common misconception: Aggregate benchmark results can hide per-task failures: 'comparable to a professional human games tester' in the DQN paper describes performance across a set of 49 games; the agent exceeded 75% of the human score on 29 of them, not on every game.
Full explanation — the complete reference version every reading depth is based on

Agent, environment, reward

In reinforcement learning the learner and decision-maker is called the agent; everything it interacts with is the environment. At each step the agent observes a situation (state), chooses an action, and the environment responds with a new situation and a number called the reward. Sutton and Barto define RL as learning how to map situations to actions so as to maximise that reward signal — without being told which actions to take.

  • Trial and error: the agent must discover good actions by trying them.
  • Delayed reward: an action can affect not only the next reward but every later situation, so credit for success must be traced back over many steps.
  • Exploration versus exploitation: repeat what has worked, or try something new that might work better? Sutton and Barto note this dilemma is still unresolved in general.

Adding up future rewards

The agent's goal is not the next reward but the total reward in the long run. The reward hypothesis states that goals and purposes can be thought of as maximising the expected cumulative sum of a scalar reward. Future rewards are usually discounted by a rate γ between 0 and 1, so a reward k steps away counts γᵏ⁻¹ times as much as one received now.

Gt=Rt+1+γRt+2+γ2Rt+3+⋯G_t = R_{t+1} + \gamma R_{t+2} + \gamma^2 R_{t+3} + \cdots

The discounted return: the sum of future rewards, each multiplied by the discount rate γ raised to how far away it is.

Gt=∑k=0∞γkRt+k+1G_t = \sum_{k=0}^{\infty} \gamma^k R_{t+k+1}

The same discounted return written as a sum.

Learning action values: Q-learning

Q-learning (Watkins, 1989) keeps an estimate Q(s, a) of how much total discounted reward follows from taking action a in state s and acting well afterwards. After each step it nudges the estimate towards the reward just received plus the discounted value of the best next action. Its estimates approximate the optimal action values whatever exploring policy the agent follows while learning.

δ=r+γmax⁡a′Q(s′,a′)−Q(s,a)\delta = r + \gamma \max_{a'} Q(s', a') - Q(s, a)

The error δ after taking action a in state s, receiving reward r and landing in state s′: the reward plus the discounted value of the best next action, minus the current estimate.

Q(s,a)←Q(s,a)+α δQ(s, a) \leftarrow Q(s, a) + \alpha\, \delta

The Q-learning update, with step size α and discount rate γ.

Worked example

Discounting: with γ = 0.9 and rewards of 1 on each of the next three steps, the return is 1 + 0.9 × 1 + 0.81 × 1 = 2.71. Q-learning: suppose Q(s, a) = 0, the step size α = 0.5, the agent receives reward 1, and the best action in the next state has Q = 2. The target is 1 + 0.9 × 2 = 2.8, the error is 2.8 − 0 = 2.8, and the new estimate is 0 + 0.5 × 2.8 = 1.4.

Landmarks

  1. 2015: a deep Q-network that saw only screen pixels and the score reached a level comparable to a professional human games tester across 49 Atari 2600 games, with one algorithm and one set of settings for all of them (Mnih et al.). It scored more than 75% of the human tester's score on 29 of the 49 games — more than half, not all.
  2. 2016: AlphaGo combined networks trained on human expert games and on self-play reinforcement learning with tree search, and beat the European Go champion 5–0 (Silver et al.).
  3. 2022: reinforcement learning from human feedback — using people's rankings of answers to provide the reward signal — was used, after supervised fine-tuning, to turn GPT-3 into InstructGPT (Ouyang et al.).
Common misconception: Misconception: 'Reinforcement learning is supervised learning with rewards instead of labels.' Correct idea: the reward says how good an outcome was, not which action would have been right. The agent has to explore to find out, and the reward may arrive many steps later.
Common misconception: Misconception: 'An RL agent will do what its designers intended.' Correct idea: it maximises the reward it is actually given. If the reward is a poor stand-in for the real goal, the agent can find unintended ways to score — the 'reward hacking' and 'side effects' problems that safety researchers list among concrete AI-safety challenges.
Info: Where this connects: machine learning places RL as a third paradigm; AI agents builds on the same agent–environment picture; language models use RL from human feedback after pre-training.

Ask ScienceVerse

Still curious about Reinforcement learning? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.

Ask the tutor about this concept on the full tutor page.

Connections

Prerequisites

Understand these first:

Guided learning path

See everything to learn before this, in order, with your progress:

Related concepts

Check your understanding

Take a quick check of two to five questions, with an explanation for every answer:

See the neighbourhood of Reinforcement learning in the Knowledge Galaxy

Sources and methodology

  • Reinforcement learning is learning what to do — how to map situations to actions — so as to maximise a numerical reward signal; the learner is not told which actions to take but must discover which actions yield the most reward by trying them. (awaiting scientific review)
  • Sutton and Barto identify trial-and-error search and delayed reward — actions can affect not only the immediate reward but also the next situation and all subsequent rewards — as the two most important distinguishing features of reinforcement learning. (awaiting scientific review)
  • A challenge that arises in reinforcement learning but not in supervised or unsupervised learning is the trade-off between exploration and exploitation: to get reward an agent must prefer actions it has found effective, but to discover such actions it must try actions it has not selected before; Sutton and Barto describe the dilemma as still unresolved. (awaiting scientific review)
  • Sutton and Barto's reward hypothesis states that all of what we mean by goals and purposes can be well thought of as the maximisation of the expected value of the cumulative sum of a received scalar signal called reward. (awaiting scientific review)
  • The discounted return Gₜ = Rₜ₊₁ + γRₜ₊₂ + γ²Rₜ₊₃ + … uses a discount rate γ between 0 and 1, so that a reward received k time steps in the future is worth only γᵏ⁻¹ times what it would be worth if received immediately. (awaiting scientific review)
  • Q-learning (Watkins, 1989) is an off-policy temporal-difference control algorithm whose learned action-value function directly approximates the optimal action-value function, independent of the policy being followed. (awaiting scientific review)
  • The deep Q-network reported by Mnih and colleagues in 2015, receiving only the pixels and the game score as inputs, surpassed all previous algorithms and reached a level comparable to a professional human games tester across a set of 49 Atari 2600 games, using the same algorithm, network architecture and hyperparameters for every game. (awaiting scientific review)
  • The 2015 deep Q-network paper reports that the agent achieved more than 75% of a professional human games tester's score on 29 of the 49 Atari games — more than half, but not all of them. (awaiting scientific review)
  • AlphaGo (Silver and colleagues, 2016) trained deep neural networks by supervised learning from human expert games and by reinforcement learning from games of self-play, combined them with tree search, and defeated the human European Go champion by 5 games to 0. (awaiting scientific review)
  • Amodei and colleagues (2016, preprint) list avoiding side effects and avoiding reward hacking as safety problems that originate from having the wrong objective function, alongside scalable supervision, safe exploration and distributional shift. (awaiting scientific review)
  • Ouyang and colleagues (2022) collected human rankings of language-model outputs and used them to fine-tune the model with reinforcement learning from human feedback, producing the InstructGPT models. (awaiting scientific review)
  • Worked calculation (author's own, using Sutton and Barto's definitions): with γ = 0.9, rewards of 1, 1 and 1 on the next three steps give a discounted return of 1 + 0.9 + 0.81 = 2.71; a Q-learning update with α = 0.5, Q(s, a) = 0, reward 1 and best next value 2 gives 0 + 0.5 × (1 + 0.9 × 2 − 0) = 1.4. (awaiting scientific review)

Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.

Content status: published 1 October 2026.

  • Scientific review: this version has not yet been signed off by a scientific reviewer.
  • The Advanced explanation has not yet been reviewed for age suitability.