Skip to main content
ScienceVerse
Artificial IntelligenceDifficulty 1-3

Training data

Training data is the set of examples a machine-learning model learns from, and it shapes everything the model can and cannot do. Data is usually split into training, validation and test sets so that the model is judged on examples it never learned from. Gaps, mislabelled examples and skewed samples in the data become gaps, errors and biases in the model, which is why datasets need documenting and checking like any other scientific instrument.

Sign in to save this concept.

Data quality, provenance and documentation

Supervised learning estimates a mapping from the empirical distribution of training pairs. Any mismatch between that distribution and the deployment distribution — through sampling, labelling or time — becomes generalisation error that held-out testing on the same flawed distribution cannot detect.

  • NIST's three bias categories: systemic, computational and statistical, human-cognitive — none requires intent.
  • Label noise in test sets can reverse model rankings: Northcutt et al. report that on ImageNet with corrected labels, ResNet-18 outperforms ResNet-50 if the prevalence of originally mislabelled test examples rises by just 6%.
  • Contamination: web-scale corpora can contain benchmark items, so GPT-3's authors ran dedicated contamination studies.
  • Documentation: datasheets record motivation, composition, collection process and recommended uses.
Common misconception: Mitigating one measured bias does not certify a system as fair: NIST states that systems in which harmful biases are mitigated are not necessarily fair — for example, they may still be inaccessible to people with disabilities or exacerbate existing disparities.
Full explanation — the complete reference version every reading depth is based on

What training data is

A training dataset is a collection of examples. In supervised learning each example pairs an input (an image, a sentence, a row of measurements) with a label giving the correct output. The model adjusts its internal parameters until its outputs match the labels well. Whatever patterns are in the examples — including accidental ones — are what the model learns.

Splitting the data

  • Training set: the examples the model's parameters are fitted to.
  • Validation set: held back from fitting and used to make design choices, such as how large the model should be.
  • Test set: used once, at the end, to estimate performance on new data. It must not influence any choice about the model.

A classic example is the MNIST database of handwritten digits described by LeCun and colleagues in 1998: they trained on 60,000 examples and tested on 10,000 images. For one of the two collections the digits came from, the first 250 writers went into training and the other 250 into testing, so the test checked handwriting the model had never seen.

How data goes wrong

  • Unrepresentative samples: in 2018, Buolamwini and Gebru found two facial-analysis benchmarks were 79.6% (IJB-A) and 86.2% (Adience) lighter-skinned subjects, so performance on darker-skinned faces was measured on comparatively few examples.
  • Wrong labels: a 2021 study estimated at least 3.3% label errors on average across the test sets of 10 popular datasets, and at least 6% in the ImageNet validation set.
  • Out-of-date data: NIST warns training data can become stale or detached from the context where the system is later used.
  • Contamination: test questions can leak into web-scale training data simply because they are published online, which can make scores look better than they are. The GPT-3 authors measured this: the effect was small on most benchmarks, but on a few it could be inflating results, so they flagged or dropped those.

NIST groups AI bias into three categories — systemic, computational and statistical, and human-cognitive — and stresses that each can arise without anyone intending to discriminate. A skewed dataset is an example of how bias can enter quietly, through who and what happened to be collected. The 'Biased Training Cards' activity shows the effect with paper cards: a learner who is only shown red apples tends to write a rule such as 'red and round', which then rejects green apples and accepts a red ball.

Worked example: auditing a dataset

Suppose a face dataset has 1,000 images and 862 of them are of lighter-skinned people. The share is 862 / 1,000 = 86.2% — the same proportion reported for the Adience benchmark. Only 138 images (13.8%) would then represent everyone else. A model can show a high overall accuracy while being far less accurate for the under-represented group, because that group contributes little to the overall score. The fix starts with measuring composition and reporting accuracy separately for each group, not just in aggregate.

Common misconception: Misconception: 'Data is neutral, so a model trained on lots of data is objective.' Correct idea: data records choices about what was collected, from whom and how it was labelled. More data of the same skew reproduces the skew; size does not cancel bias.
Common misconception: Misconception: 'Benchmark labels are ground truth.' Correct idea: labels are made by people and contain errors — measurably so in widely used test sets — so small differences between models' scores may not be meaningful.

Documenting data

Gebru and colleagues proposed 'datasheets for datasets': just as an electronic component comes with a datasheet, every dataset should record its motivation, composition, collection process and recommended uses. NIST's framework similarly says that maintaining the provenance of training data supports transparency and accountability, and notes that training data may be subject to copyright.

Info: Where this connects: machine learning explains why the train/test split matters; computer vision and language models show what happens when web-scale data meets real-world use.

Ask ScienceVerse

Still curious about Training data? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.

Ask the tutor about this concept on the full tutor page.

Connections

Prerequisites

Understand these first:

Guided learning path

See everything to learn before this, in order, with your progress:

Leads to

This concept is a building block for:

Try the experiment

Put this concept into practice with a hands-on activity (each shows its supervision requirement first):

Check your understanding

Take a quick check of two to five questions, with an explanation for every answer:

See the neighbourhood of Training data in the Knowledge Galaxy

Sources and methodology

  • In the experiments LeCun and colleagues reported in 1998, the MNIST handwritten-digit database supplied 60,000 training examples and a test set of 10,000 images, and one of its two source collections was split by writer so that its training and test digits came from different people. (awaiting scientific review)
    • Gradient-based learning applied to document recognition — Peer-reviewed paper
  • Buolamwini and Gebru (2018) found that two widely used facial-analysis benchmark datasets were overwhelmingly composed of lighter-skinned subjects: 79.6% for IJB-A and 86.2% for Adience. (awaiting scientific review)
  • A 2021 study estimated an average of at least 3.3% label errors across the test sets of 10 commonly used computer-vision, language and audio datasets, including at least 6% of the ImageNet validation set. (awaiting scientific review)
  • NIST identifies three major categories of AI bias — systemic, computational and statistical, and human-cognitive — and notes that each can occur in the absence of prejudice, partiality or discriminatory intent. (awaiting scientific review)
  • NIST warns that datasets used to train AI systems may become detached from their original and intended context, or may become stale or outdated relative to the context in which the system is deployed. (awaiting scientific review)
  • The GPT-3 authors described 'data contamination' — test-set content ending up in web-scale training data such as Common Crawl because it already exists on the web — as a growing problem when training high-capacity models. (awaiting scientific review)
  • Gebru and colleagues proposed that every dataset be accompanied by a 'datasheet' documenting its motivation, composition, collection process and recommended uses, by analogy with the datasheets that accompany electronic components. (awaiting scientific review)

Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.

Content status: published 1 October 2026.

  • Scientific review: this version has not yet been signed off by a scientific reviewer.
  • The Advanced explanation has not yet been reviewed for age suitability.