Learning path: Reinforcement learning
The concepts to learn before Reinforcement learning, in order. Each step links to its page and, where one exists, its quiz.
Sign in to see which steps you have already opened and which quizzes you have completed.
Your path
- Machine learningDirect prerequisite
Machine learning is the part of AI in which a program improves at a task through experience — data — instead of following rules written in advance. Supervised learning learns from labelled examples, unsupervised learning finds structure in unlabelled data, and reinforcement learning learns which actions earn reward by trial and error. The real test of a learned model is not how well it fits its training data but how well it performs on new data it has never seen.
- Reinforcement learningYour goal
Reinforcement learning (RL) is learning by trial and error: an agent takes actions in an environment, receives numerical rewards, and gradually learns which actions lead to the most reward over time. Unlike supervised learning, nobody tells the agent the right action — it must explore, and rewards may arrive long after the actions that earned them. RL produced landmark results in Atari games and Go and is used to fine-tune language models from human feedback, but an agent optimises exactly the reward it is given, which may not be what its designers meant.