Training data
Training data is the set of examples a machine-learning model learns from, and it shapes everything the model can and cannot do. Data is usually split into training, validation and test sets so that the model is judged on examples it never learned from. Gaps, mislabelled examples and skewed samples in the data become gaps, errors and biases in the model, which is why datasets need documenting and checking like any other scientific instrument.
Data quality, provenance and documentation
Supervised learning estimates a mapping from the empirical distribution of training pairs. Any mismatch between that distribution and the deployment distribution — through sampling, labelling or time — becomes generalisation error that held-out testing on the same flawed distribution cannot detect.
- NIST's three bias categories: systemic, computational and statistical, human-cognitive — none requires intent.
- Label noise in test sets can reverse model rankings: Northcutt et al. report that on ImageNet with corrected labels, ResNet-18 outperforms ResNet-50 if the prevalence of originally mislabelled test examples rises by just 6%.
- Contamination: web-scale corpora can contain benchmark items, so GPT-3's authors ran dedicated contamination studies.
- Documentation: datasheets record motivation, composition, collection process and recommended uses.
Full explanation — the complete reference version every reading depth is based on
What training data is
A training dataset is a collection of examples. In supervised learning each example pairs an input (an image, a sentence, a row of measurements) with a label giving the correct output. The model adjusts its internal parameters until its outputs match the labels well. Whatever patterns are in the examples — including accidental ones — are what the model learns.
Splitting the data
- Training set: the examples the model's parameters are fitted to.
- Validation set: held back from fitting and used to make design choices, such as how large the model should be.
- Test set: used once, at the end, to estimate performance on new data. It must not influence any choice about the model.
A classic example is the MNIST database of handwritten digits described by LeCun and colleagues in 1998: they trained on 60,000 examples and tested on 10,000 images. For one of the two collections the digits came from, the first 250 writers went into training and the other 250 into testing, so the test checked handwriting the model had never seen.
How data goes wrong
- Unrepresentative samples: in 2018, Buolamwini and Gebru found two facial-analysis benchmarks were 79.6% (IJB-A) and 86.2% (Adience) lighter-skinned subjects, so performance on darker-skinned faces was measured on comparatively few examples.
- Wrong labels: a 2021 study estimated at least 3.3% label errors on average across the test sets of 10 popular datasets, and at least 6% in the ImageNet validation set.
- Out-of-date data: NIST warns training data can become stale or detached from the context where the system is later used.
- Contamination: test questions can leak into web-scale training data simply because they are published online, which can make scores look better than they are. The GPT-3 authors measured this: the effect was small on most benchmarks, but on a few it could be inflating results, so they flagged or dropped those.
NIST groups AI bias into three categories — systemic, computational and statistical, and human-cognitive — and stresses that each can arise without anyone intending to discriminate. A skewed dataset is an example of how bias can enter quietly, through who and what happened to be collected. The 'Biased Training Cards' activity shows the effect with paper cards: a learner who is only shown red apples tends to write a rule such as 'red and round', which then rejects green apples and accepts a red ball.
Worked example: auditing a dataset
Suppose a face dataset has 1,000 images and 862 of them are of lighter-skinned people. The share is 862 / 1,000 = 86.2% — the same proportion reported for the Adience benchmark. Only 138 images (13.8%) would then represent everyone else. A model can show a high overall accuracy while being far less accurate for the under-represented group, because that group contributes little to the overall score. The fix starts with measuring composition and reporting accuracy separately for each group, not just in aggregate.
Documenting data
Gebru and colleagues proposed 'datasheets for datasets': just as an electronic component comes with a datasheet, every dataset should record its motivation, composition, collection process and recommended uses. NIST's framework similarly says that maintaining the provenance of training data supports transparency and accountability, and notes that training data may be subject to copyright.
Ask ScienceVerse
Still curious about Training data? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.
Ask the tutor about this concept on the full tutor page.
Connections
Guided learning path
See everything to learn before this, in order, with your progress:
Try the experiment
Put this concept into practice with a hands-on activity (each shows its supervision requirement first):
Check your understanding
Take a quick check of two to five questions, with an explanation for every answer:
See the neighbourhood of Training data in the Knowledge Galaxy
Sources and methodology
- In the experiments LeCun and colleagues reported in 1998, the MNIST handwritten-digit database supplied 60,000 training examples and a test set of 10,000 images, and one of its two source collections was split by writer so that its training and test digits came from different people. (awaiting scientific review)
- Gradient-based learning applied to document recognition — Peer-reviewed paper
- Buolamwini and Gebru (2018) found that two widely used facial-analysis benchmark datasets were overwhelmingly composed of lighter-skinned subjects: 79.6% for IJB-A and 86.2% for Adience. (awaiting scientific review)
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification — Peer-reviewed paper
- A 2021 study estimated an average of at least 3.3% label errors across the test sets of 10 commonly used computer-vision, language and audio datasets, including at least 6% of the ImageNet validation set. (awaiting scientific review)
- Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks — Peer-reviewed paper
- NIST identifies three major categories of AI bias — systemic, computational and statistical, and human-cognitive — and notes that each can occur in the absence of prejudice, partiality or discriminatory intent. (awaiting scientific review)
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023) — Government or standards body
- NIST warns that datasets used to train AI systems may become detached from their original and intended context, or may become stale or outdated relative to the context in which the system is deployed. (awaiting scientific review)
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023) — Government or standards body
- The GPT-3 authors described 'data contamination' — test-set content ending up in web-scale training data such as Common Crawl because it already exists on the web — as a growing problem when training high-capacity models. (awaiting scientific review)
- Language Models are Few-Shot Learners — Peer-reviewed paper
- Gebru and colleagues proposed that every dataset be accompanied by a 'datasheet' documenting its motivation, composition, collection process and recommended uses, by analogy with the datasheets that accompany electronic components. (awaiting scientific review)
- Datasheets for datasets — Peer-reviewed paper
Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.
Content status: published 1 October 2026.
- Scientific review: this version has not yet been signed off by a scientific reviewer.
- The Advanced explanation has not yet been reviewed for age suitability.