Skip to main content
ScienceVerse

Learning path: Transformers

The concepts to learn before Transformers, in order. Each step links to its page and, where one exists, its quiz.

Sign in to see which steps you have already opened and which quizzes you have completed.

Your path

  1. Neural networksDirect prerequisite

    An artificial neural network is a mathematical function built from many simple units arranged in layers: each unit multiplies its inputs by weights, adds a bias and passes the total through an activation function. Stacking layers lets the network represent patterns — such as XOR — that no single linear unit can. The weights are not written by hand but learned from data, usually by back-propagation and gradient descent; the networks are inspired by the brain but are not realistic models of it.

    Take the Neural networks quiz

  2. TransformersYour goal

    A transformer is a neural-network architecture, introduced in 2017, that processes a whole sequence at once using attention: every position computes how relevant every other position is and takes a weighted mix of their information. Because it drops the step-by-step recurrence of earlier language models, it trains far more in parallel; it became the de facto standard architecture for language processing, underlies large language models such as GPT-3, and has been applied to images too. Its main cost is that attention compares every pair of positions, so computation grows with the square of the sequence length.

    Take the Transformers quiz