Skip to main content
ScienceVerse

Learning path: Training data

The concepts to learn before Training data, in order. Each step links to its page and, where one exists, its quiz.

Sign in to see which steps you have already opened and which quizzes you have completed.

Your path

  1. Machine learningDirect prerequisite

    Machine learning is the part of AI in which a program improves at a task through experience — data — instead of following rules written in advance. Supervised learning learns from labelled examples, unsupervised learning finds structure in unlabelled data, and reinforcement learning learns which actions earn reward by trial and error. The real test of a learned model is not how well it fits its training data but how well it performs on new data it has never seen.

    Take the Machine learning quiz

  2. Training dataYour goal

    Training data is the set of examples a machine-learning model learns from, and it shapes everything the model can and cannot do. Data is usually split into training, validation and test sets so that the model is judged on examples it never learned from. Gaps, mislabelled examples and skewed samples in the data become gaps, errors and biases in the model, which is why datasets need documenting and checking like any other scientific instrument.

    Take the Training data quiz

Where to go next

Concepts that build directly on Training data:

  • Computer vision: Computer vision is the field of getting machines to extract useful information from images and video — recognising objects, reading text, finding faces. To a computer an image is a grid of numbers; convolutional neural networks learn small filters that detect edges and textures and combine them into larger patterns. The organisers of the ImageNet challenge describe 2012 as a turning point, after which the vast majority of entries used deep convolutional networks. High benchmark scores still hide real failure modes: tiny engineered perturbations can fool models, and systems trained or tested on skewed data can be far less accurate for some groups of people.
  • Language models: A language model is a system that assigns probabilities to sequences of words — in practice, it repeatedly predicts a probability for each possible next token (a word or piece of a word) given the text so far. Large language models are neural networks with billions of parameters trained on vast amounts of text, and they can carry out many tasks described in plain language. Because they generate statistically likely text rather than looking facts up, they can state false things confidently ('hallucination'), and they reproduce biases in their training data.