Skip to main content
ScienceVerse
Artificial IntelligenceDifficulty 2-4

Computer vision

Computer vision is the field of getting machines to extract useful information from images and video — recognising objects, reading text, finding faces. To a computer an image is a grid of numbers; convolutional neural networks learn small filters that detect edges and textures and combine them into larger patterns. The organisers of the ImageNet challenge describe 2012 as a turning point, after which the vast majority of entries used deep convolutional networks. High benchmark scores still hide real failure modes: tiny engineered perturbations can fool models, and systems trained or tested on skewed data can be far less accurate for some groups of people.

Sign in to save this concept.

Architectures and benchmarks

LeNet-5 (1998) established the convolutional template — local receptive fields, weight sharing, sub-sampling — reaching 0.95% MNIST test error (0.8% with distortion-based augmentation). AlexNet (2012) scaled the template with ReLU units, GPU training and dropout, winning ILSVRC-2012 at 15.3% top-5 error versus 26.2%. Vision Transformers (2020) replaced convolution with self-attention over image patches, performing very well given large-scale pre-training.

  • Human baselines are fragile: ILSVRC's annotators scored about 5.1% and about 12.0% top-5 error on different samples.
  • Disaggregated evaluation exposes subgroup failure (Gender Shades: up to 34.7% versus at most 0.8%).
  • Benchmarks carry label noise and sampling choices inherited from how they were built.
Common misconception: Convolutional networks are not fully invariant to transformations: sub-sampling gives 'some degree' of tolerance to shifts, scale and distortion, as LeCun et al. put it, which is why augmentation with distorted examples still improves results.
Full explanation — the complete reference version every reading depth is based on

An image is a grid of numbers

A digital image is a grid of pixels. Each pixel's value comes from an image sensor measuring how much light reached that spot, and it is stored as numbers — one brightness value for a greyscale image, or three (red, green, blue) for colour. The MNIST handwritten digits, for example, are centred in 28 × 28 pixel fields, so each digit is 784 numbers. The challenge of computer vision is to go from those raw numbers to an answer like 'this is a 7' or 'there is a cyclist on the left'.

Convolutional neural networks

A convolutional network slides small filters across the image. Each filter is a little grid of weights; at each position it multiplies the pixels underneath by its weights and adds them up, producing a 'feature map' that lights up wherever its pattern appears. LeCun and colleagues (1998) describe three ideas that make this work: local receptive fields (each unit looks at a small patch), shared weights (the same filter is reused everywhere) and sub-sampling (shrinking feature maps to tolerate small shifts and distortions). They note the local-receptive-field idea arose almost at the same time as Hubel and Wiesel's discovery of orientation-selective neurons in the cat's visual system.

Worked example: an edge detector

Take a 3 × 3 filter whose left column is −1, middle column 0 and right column +1. Place it on a patch whose left two columns are dark (0) and right column bright (9): the sum is 3 × 9 = 27, a strong response. On a patch that is the same brightness everywhere, the −1s and +1s cancel and the response is 0. That filter is a vertical-edge detector. A trained network learns thousands of filters like this by itself. Goodfellow and colleagues illustrate the layers that follow: edges are combined into corners and contours, those into object parts, and parts into a whole object. The 'Guess the Picture from Its Features' activity is a hands-on analogy for that build-up — an analogy only, because the network learns its features from data.

Milestones

  1. 1998: the convolutional network LeNet-5 reached 0.95% test error on MNIST digits (0.8% with distorted training examples).
  2. 2012: in the ImageNet challenge (1000 object classes), Krizhevsky, Sutskever and Hinton's deep convolutional network won with a top-5 error of 15.3%, against 26.2% for the runner-up.
  3. 2020: Vision Transformers showed that a transformer applied to sequences of image patches can also classify images very well when pre-trained on large datasets.

'Top-5 error' counts an answer as wrong only if the correct label is missing from the model's five best guesses. Comparisons with people need care: in the ImageNet analysis, one trained annotator had about 5.1% top-5 error on a 1,500-image sample, while a second annotator had about 12.0% on a smaller sample.

Failure modes

  • Adversarial examples: Szegedy et al. (2013) showed an imperceptible, deliberately computed change to an image can make a network misclassify it — and the same change can fool a different network. Goodfellow et al. (2014) argued the main cause is the models' linear nature, and that training on such examples helps.
  • Unequal accuracy: a 2018 audit of three commercial gender classifiers found error rates up to 34.7% for darker-skinned women, against at most 0.8% for lighter-skinned men.
  • Labels and data: benchmark images were labelled by crowdworkers, and training sets can under-represent people, places and lighting conditions the system later meets.
Common misconception: Misconception: 'A vision model that beats people on a benchmark sees the world the way people do.' Correct idea: models can be fooled by changes people cannot even notice, and a benchmark measures one dataset's images and labels — not vision in general.
Common misconception: Misconception: 'A high overall accuracy means the system works well for everyone.' Correct idea: accuracy must be checked for each group separately. The 2018 Gender Shades audit found a gap from 0.8% to 34.7% error between groups.
Info: Where this connects: neural networks supply the building blocks; training data explains why skewed or mislabelled datasets matter; transformers now compete with convolutional networks in vision.

Ask ScienceVerse

Still curious about Computer vision? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.

Ask the tutor about this concept on the full tutor page.

Connections

Prerequisites

Understand these first:

Guided learning path

See everything to learn before this, in order, with your progress:

Related concepts

Try the experiment

Put this concept into practice with a hands-on activity (each shows its supervision requirement first):

Check your understanding

Take a quick check of two to five questions, with an explanation for every answer:

See the neighbourhood of Computer vision in the Knowledge Galaxy

Sources and methodology

  • Convolutional networks combine three architectural ideas — local receptive fields, shared weights and spatial or temporal sub-sampling — to give some degree of invariance to shifts, changes of scale and distortions. (awaiting scientific review)
    • Gradient-based learning applied to document recognition — Peer-reviewed paper
  • LeCun and colleagues note that the idea of connecting units to local receptive fields was almost simultaneous with Hubel and Wiesel's discovery of locally sensitive, orientation-selective neurons in the cat's visual system. (awaiting scientific review)
    • Gradient-based learning applied to document recognition — Peer-reviewed paper
  • Goodfellow and colleagues illustrate how a deep network breaks image recognition into nested steps: given the pixels, the first hidden layer can identify edges by comparing the brightness of neighbouring pixels, later layers combine edges into corners and contours and then into object parts, and the output identifies the object. (awaiting scientific review)
  • The convolutional network LeNet-5 reached a test error rate of 0.95% on the MNIST handwritten-digit test set, falling to 0.8% when it was trained with artificially distorted examples. (awaiting scientific review)
    • Gradient-based learning applied to document recognition — Peer-reviewed paper
  • In the ILSVRC-2012 competition, a variant of the deep convolutional network described by Krizhevsky, Sutskever and Hinton achieved a winning top-5 test error rate of 15.3%, compared with 26.2% for the second-best entry. (awaiting scientific review)
  • The image-classification task of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) uses 1000 object classes, and the challenge relied extensively on crowdworkers recruited through Amazon Mechanical Turk to annotate its images. (awaiting scientific review)
  • The organisers of the ImageNet challenge describe 2012 as a turning point brought by large-scale convolutional neural networks: following the success of a deep-learning method that year, the vast majority of entries in 2013 used deep convolutional networks. (awaiting scientific review)
  • In the ILSVRC analysis, one trained human annotator's estimated top-5 classification error on a 1,500-image sample was 5.1%, compared with 6.8% for the GoogLeNet model on the same sample; a second annotator, on a smaller sample, had an estimated 12.0% error. (awaiting scientific review)
  • Szegedy et al. (2013) showed that a neural network can be made to misclassify an image by applying a certain imperceptible perturbation, and that the same perturbation can cause a different network, trained on a different subset of the data, to misclassify the same input. (awaiting scientific review)
  • Goodfellow, Shlens and Szegedy (2014) argued that the primary cause of neural networks' vulnerability to adversarial perturbations is their linear nature, and used adversarial examples for training to reduce the test-set error of a maxout network on MNIST. (awaiting scientific review)
  • In a 2018 audit of three commercial gender-classification systems, darker-skinned females were the most misclassified group, with error rates of up to 34.7%, while the maximum error rate for lighter-skinned males was 0.8%. (awaiting scientific review)
  • Dosovitskiy and colleagues (2020) showed that a pure transformer applied directly to sequences of image patches can perform very well on image classification when pre-trained on large amounts of data. (awaiting scientific review)
  • Worked calculation (author's own illustration of a convolution): a 3×3 filter with columns −1, 0, +1 applied to a 3×3 patch whose left two columns are 0 and right column is 9 gives 3 × 9 = 27, while the same filter on a uniform patch gives 0, so the filter responds to a vertical dark-to-bright edge. (awaiting scientific review)
    • Gradient-based learning applied to document recognition — Peer-reviewed paper

Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.

Content status: published 1 October 2026.

  • Scientific review: this version has not yet been signed off by a scientific reviewer.
  • The Advanced explanation has not yet been reviewed for age suitability.