Transformers
A transformer is a neural-network architecture, introduced in 2017, that processes a whole sequence at once using attention: every position computes how relevant every other position is and takes a weighted mix of their information. Because it drops the step-by-step recurrence of earlier language models, it trains far more in parallel; it became the de facto standard architecture for language processing, underlies large language models such as GPT-3, and has been applied to images too. Its main cost is that attention compares every pair of positions, so computation grows with the square of the sequence length.
Architecture and complexity
The 2017 model was an encoder–decoder with N = 6 layers each side, model dimension 512 and h = 8 heads (key and value dimension 64 each), residual connections and layer normalisation around each sublayer, and sinusoidal positional encodings. Decoder self-attention is masked to preserve the autoregressive property.
- Self-attention: O(n²·d) per layer, O(1) sequential operations, O(1) maximum path length.
- Recurrent layer: O(n·d²) per layer, O(n) sequential operations, O(n) path length.
- Result: 28.4 BLEU (WMT14 En–De) and 41.8 BLEU (WMT14 En–Fr, single model) after 3.5 days on eight GPUs.
Descendants specialise the template: encoder-only bidirectional models such as BERT, decoder-only autoregressive language models, and Vision Transformers operating on image patches.
Full explanation — the complete reference version every reading depth is based on
The problem transformers solved
Earlier neural translation models read a sentence one word at a time with recurrent networks and squeezed it into a single fixed-length vector. Bahdanau and colleagues (2014) conjectured that this vector was a bottleneck and let the model 'search' back over the source sentence for the parts relevant to each output word — an early attention mechanism. In 2017 Vaswani and colleagues went further: their Transformer is based solely on attention, dispensing with recurrence and convolutions entirely.
Attention: queries, keys and values
Each token is turned into three vectors: a query (what am I looking for?), a key (what do I contain?) and a value (what information do I pass on?). Attention compares a token's query with every key, turns the comparison scores into weights with a softmax, and outputs a weighted sum of the values. Tokens that are relevant to each other end up sharing information, however far apart they are.
Attention weights A: compare queries Q with keys K, scale by the square root of the key dimension dₖ, and apply softmax so each row of weights adds up to 1.
Scaled dot-product attention: the weights A mix the values V.
The division by √dₖ matters: without it, large dot products push the softmax into regions where its gradients are extremely small, which makes training hard. The original model ran 8 attention 'heads' in parallel, each with its own learned projections and dimension 64, so different heads can attend to different kinds of relationship.
Worked example: one attention step
Let dₖ = 4, so √dₖ = 2. The query is [1, 0, 1, 0]. Key 1 is [1, 0, 1, 0] (dot product 2) and key 2 is [0, 1, 0, 1] (dot product 0). Scaled scores: 1 and 0. Softmax: e¹/(e¹ + e⁰) ≈ 2.718/3.718 ≈ 0.731 and ≈ 0.269. If value 1 is [10, 0] and value 2 is [0, 10], the output is 0.731 × [10, 0] + 0.269 × [0, 10] ≈ [7.31, 2.69] — mostly the information from the token whose key matched.
Putting the architecture together
- Layers: the original Transformer stacked 6 identical layers in the encoder and 6 in the decoder (model width 512), each combining multi-head attention with a small feed-forward network, plus residual connections and layer normalisation.
- Order: attention by itself ignores word order, so positional encodings (sine and cosine waves of different frequencies in the original paper) are added to the token embeddings.
- Masking: when generating text, the decoder's attention is masked so position i can only look at earlier positions — it cannot peek at the answer.
- Variants: BERT (2018) pre-trains bidirectional transformer encoders that read left and right context; GPT-style models are decoder-only and predict the next token; Vision Transformers apply the same idea to image patches. By 2020 the Transformer had become the de facto standard for natural-language processing.
Strengths and costs
A self-attention layer connects any two positions in a single step and needs only O(1) sequential operations, whereas a recurrent layer needs O(n) sequential steps — so transformers parallelise well on GPUs. On translation, the 2017 model reached 28.4 BLEU (English–German) and 41.8 BLEU (English–French) after 3.5 days on eight GPUs, a small fraction of the training cost of earlier best models. The price is that attention compares every pair of positions: per-layer cost grows as O(n²·d), so very long inputs are expensive.
Ask ScienceVerse
Still curious about Transformers? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.
Ask the tutor about this concept on the full tutor page.
Connections
Guided learning path
See everything to learn before this, in order, with your progress:
Related concepts
- Language models — Applied in
- Computer vision — Applied in
- CPU — Related to
Check your understanding
Take a quick check of two to five questions, with an explanation for every answer:
See the neighbourhood of Transformers in the Knowledge Galaxy
Sources and methodology
- The Transformer, proposed by Vaswani and colleagues in 2017, is a sequence-transduction architecture based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- Vaswani and colleagues describe an attention function as mapping a query and a set of key–value pairs to an output computed as a weighted sum of the values, where each value's weight comes from a compatibility function of the query with the corresponding key. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- In scaled dot-product attention, the dot products of a query with all keys are divided by √dₖ (the square root of the key dimension) and passed through a softmax to obtain the weights on the values; the authors introduced the scaling to counteract what they suspected was the effect of large dot products pushing the softmax into regions with extremely small gradients. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- The original Transformer used h = 8 parallel attention heads, each with key and value dimension 64 (the model dimension of 512 divided by 8), and stacks of N = 6 identical layers in both its encoder and its decoder. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- Because the Transformer contains no recurrence and no convolution, positional encodings — sine and cosine functions of different frequencies in the original paper — are added to the input embeddings so the model can use the order of the sequence. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- A self-attention layer has per-layer complexity O(n²·d) for sequence length n and representation dimension d, with O(1) sequential operations and an O(1) maximum path length between positions, whereas a recurrent layer needs O(n) sequential operations. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- In the Transformer decoder, self-attention is masked so that the prediction for position i can depend only on the known outputs at positions before i. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- The original Transformer reached 28.4 BLEU on the WMT 2014 English-to-German translation task and a single-model 41.8 BLEU on WMT 2014 English-to-French after training for 3.5 days on eight GPUs. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
- Bahdanau and colleagues (2014) conjectured that encoding a whole source sentence into one fixed-length vector is a bottleneck, and let a translation model (soft-)search for the parts of the source sentence relevant to predicting each target word. (awaiting scientific review)
- Neural Machine Translation by Jointly Learning to Align and Translate — Peer-reviewed paper
- BERT (Devlin and colleagues, 2018) pre-trains deep bidirectional Transformer representations from unlabelled text by jointly conditioning on both left and right context in all layers. (awaiting scientific review)
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Peer-reviewed paper
- Dosovitskiy and colleagues (2020) describe the Transformer architecture as having become the de facto standard for natural language processing tasks. (awaiting scientific review)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Peer-reviewed paper
- Worked calculation (author's own): with key dimension dₖ = 4, a query [1, 0, 1, 0] and keys [1, 0, 1, 0] and [0, 1, 0, 1] give dot products 2 and 0, scaled scores 1 and 0, and softmax weights of about 0.731 and 0.269; with values [10, 0] and [0, 10] the attention output is about [7.31, 2.69]. (awaiting scientific review)
- Attention Is All You Need — Peer-reviewed paper
Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.
Content status: published 1 October 2026.
- Scientific review: this version has not yet been signed off by a scientific reviewer.
- The Advanced explanation has not yet been reviewed for age suitability.