Skip to main content
ScienceVerse

Learning path: Neural networks

The concepts to learn before Neural networks, in order. Each step links to its page and, where one exists, its quiz.

Sign in to see which steps you have already opened and which quizzes you have completed.

Your path

  1. Machine learningDirect prerequisite

    Machine learning is the part of AI in which a program improves at a task through experience — data — instead of following rules written in advance. Supervised learning learns from labelled examples, unsupervised learning finds structure in unlabelled data, and reinforcement learning learns which actions earn reward by trial and error. The real test of a learned model is not how well it fits its training data but how well it performs on new data it has never seen.

    Take the Machine learning quiz

  2. An artificial neural network is a mathematical function built from many simple units arranged in layers: each unit multiplies its inputs by weights, adds a bias and passes the total through an activation function. Stacking layers lets the network represent patterns — such as XOR — that no single linear unit can. The weights are not written by hand but learned from data, usually by back-propagation and gradient descent; the networks are inspired by the brain but are not realistic models of it.

    Take the Neural networks quiz

Where to go next

Concepts that build directly on Neural networks:

  • Computer vision: Computer vision is the field of getting machines to extract useful information from images and video — recognising objects, reading text, finding faces. To a computer an image is a grid of numbers; convolutional neural networks learn small filters that detect edges and textures and combine them into larger patterns. The organisers of the ImageNet challenge describe 2012 as a turning point, after which the vast majority of entries used deep convolutional networks. High benchmark scores still hide real failure modes: tiny engineered perturbations can fool models, and systems trained or tested on skewed data can be far less accurate for some groups of people.
  • Language models: A language model is a system that assigns probabilities to sequences of words — in practice, it repeatedly predicts a probability for each possible next token (a word or piece of a word) given the text so far. Large language models are neural networks with billions of parameters trained on vast amounts of text, and they can carry out many tasks described in plain language. Because they generate statistically likely text rather than looking facts up, they can state false things confidently ('hallucination'), and they reproduce biases in their training data.
  • Transformers: A transformer is a neural-network architecture, introduced in 2017, that processes a whole sequence at once using attention: every position computes how relevant every other position is and takes a weighted mix of their information. Because it drops the step-by-step recurrence of earlier language models, it trains far more in parallel; it became the de facto standard architecture for language processing, underlies large language models such as GPT-3, and has been applied to images too. Its main cost is that attention compares every pair of positions, so computation grows with the square of the sequence length.