Skip to main content
ScienceVerse
Artificial IntelligenceDifficulty 3-5

Transformers

A transformer is a neural-network architecture, introduced in 2017, that processes a whole sequence at once using attention: every position computes how relevant every other position is and takes a weighted mix of their information. Because it drops the step-by-step recurrence of earlier language models, it trains far more in parallel; it became the de facto standard architecture for language processing, underlies large language models such as GPT-3, and has been applied to images too. Its main cost is that attention compares every pair of positions, so computation grows with the square of the sequence length.

Sign in to save this concept.

Architecture and complexity

The 2017 model was an encoder–decoder with N = 6 layers each side, model dimension 512 and h = 8 heads (key and value dimension 64 each), residual connections and layer normalisation around each sublayer, and sinusoidal positional encodings. Decoder self-attention is masked to preserve the autoregressive property.

  • Self-attention: O(n²·d) per layer, O(1) sequential operations, O(1) maximum path length.
  • Recurrent layer: O(n·d²) per layer, O(n) sequential operations, O(n) path length.
  • Result: 28.4 BLEU (WMT14 En–De) and 41.8 BLEU (WMT14 En–Fr, single model) after 3.5 days on eight GPUs.

Descendants specialise the template: encoder-only bidirectional models such as BERT, decoder-only autoregressive language models, and Vision Transformers operating on image patches.

Common misconception: Quadratic cost is per layer in sequence length n; the constant factors and the d-dependence mean that for short sequences self-attention can be cheaper than recurrence, which is the regime Vaswani et al. highlighted (n smaller than d).
Full explanation — the complete reference version every reading depth is based on

The problem transformers solved

Earlier neural translation models read a sentence one word at a time with recurrent networks and squeezed it into a single fixed-length vector. Bahdanau and colleagues (2014) conjectured that this vector was a bottleneck and let the model 'search' back over the source sentence for the parts relevant to each output word — an early attention mechanism. In 2017 Vaswani and colleagues went further: their Transformer is based solely on attention, dispensing with recurrence and convolutions entirely.

Attention: queries, keys and values

Each token is turned into three vectors: a query (what am I looking for?), a key (what do I contain?) and a value (what information do I pass on?). Attention compares a token's query with every key, turns the comparison scores into weights with a softmax, and outputs a weighted sum of the values. Tokens that are relevant to each other end up sharing information, however far apart they are.

A=softmax ⁣(QK⊤dk)A = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)

Attention weights A: compare queries Q with keys K, scale by the square root of the key dimension dₖ, and apply softmax so each row of weights adds up to 1.

Attention(Q,K,V)=A V\mathrm{Attention}(Q, K, V) = A\, V

Scaled dot-product attention: the weights A mix the values V.

The division by √dₖ matters: without it, large dot products push the softmax into regions where its gradients are extremely small, which makes training hard. The original model ran 8 attention 'heads' in parallel, each with its own learned projections and dimension 64, so different heads can attend to different kinds of relationship.

Worked example: one attention step

Let dₖ = 4, so √dₖ = 2. The query is [1, 0, 1, 0]. Key 1 is [1, 0, 1, 0] (dot product 2) and key 2 is [0, 1, 0, 1] (dot product 0). Scaled scores: 1 and 0. Softmax: e¹/(e¹ + e⁰) ≈ 2.718/3.718 ≈ 0.731 and ≈ 0.269. If value 1 is [10, 0] and value 2 is [0, 10], the output is 0.731 × [10, 0] + 0.269 × [0, 10] ≈ [7.31, 2.69] — mostly the information from the token whose key matched.

Putting the architecture together

  • Layers: the original Transformer stacked 6 identical layers in the encoder and 6 in the decoder (model width 512), each combining multi-head attention with a small feed-forward network, plus residual connections and layer normalisation.
  • Order: attention by itself ignores word order, so positional encodings (sine and cosine waves of different frequencies in the original paper) are added to the token embeddings.
  • Masking: when generating text, the decoder's attention is masked so position i can only look at earlier positions — it cannot peek at the answer.
  • Variants: BERT (2018) pre-trains bidirectional transformer encoders that read left and right context; GPT-style models are decoder-only and predict the next token; Vision Transformers apply the same idea to image patches. By 2020 the Transformer had become the de facto standard for natural-language processing.

Strengths and costs

A self-attention layer connects any two positions in a single step and needs only O(1) sequential operations, whereas a recurrent layer needs O(n) sequential steps — so transformers parallelise well on GPUs. On translation, the 2017 model reached 28.4 BLEU (English–German) and 41.8 BLEU (English–French) after 3.5 days on eight GPUs, a small fraction of the training cost of earlier best models. The price is that attention compares every pair of positions: per-layer cost grows as O(n²·d), so very long inputs are expensive.

Common misconception: Misconception: 'Attention means the model pays attention and understands like a person.' Correct idea: attention is a specific calculation — softmax-weighted averaging of value vectors, with weights set by query–key dot products. The name describes the arithmetic, not awareness.
Common misconception: Misconception: 'A transformer reads words in order, like a recurrent network.' Correct idea: self-attention treats the input as a set; without the added positional encodings it could not tell 'dog bites man' from 'man bites dog'.
Info: Where this connects: neural networks supply the layers and training method; language models are the biggest application; computer vision now uses transformers on image patches.

Ask ScienceVerse

Still curious about Transformers? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.

Ask the tutor about this concept on the full tutor page.

Connections

Prerequisites

Understand these first:

Guided learning path

See everything to learn before this, in order, with your progress:

Related concepts

Check your understanding

Take a quick check of two to five questions, with an explanation for every answer:

See the neighbourhood of Transformers in the Knowledge Galaxy

Sources and methodology

  • The Transformer, proposed by Vaswani and colleagues in 2017, is a sequence-transduction architecture based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. (awaiting scientific review)
  • Vaswani and colleagues describe an attention function as mapping a query and a set of key–value pairs to an output computed as a weighted sum of the values, where each value's weight comes from a compatibility function of the query with the corresponding key. (awaiting scientific review)
  • In scaled dot-product attention, the dot products of a query with all keys are divided by √dₖ (the square root of the key dimension) and passed through a softmax to obtain the weights on the values; the authors introduced the scaling to counteract what they suspected was the effect of large dot products pushing the softmax into regions with extremely small gradients. (awaiting scientific review)
  • The original Transformer used h = 8 parallel attention heads, each with key and value dimension 64 (the model dimension of 512 divided by 8), and stacks of N = 6 identical layers in both its encoder and its decoder. (awaiting scientific review)
  • Because the Transformer contains no recurrence and no convolution, positional encodings — sine and cosine functions of different frequencies in the original paper — are added to the input embeddings so the model can use the order of the sequence. (awaiting scientific review)
  • A self-attention layer has per-layer complexity O(n²·d) for sequence length n and representation dimension d, with O(1) sequential operations and an O(1) maximum path length between positions, whereas a recurrent layer needs O(n) sequential operations. (awaiting scientific review)
  • In the Transformer decoder, self-attention is masked so that the prediction for position i can depend only on the known outputs at positions before i. (awaiting scientific review)
  • The original Transformer reached 28.4 BLEU on the WMT 2014 English-to-German translation task and a single-model 41.8 BLEU on WMT 2014 English-to-French after training for 3.5 days on eight GPUs. (awaiting scientific review)
  • Bahdanau and colleagues (2014) conjectured that encoding a whole source sentence into one fixed-length vector is a bottleneck, and let a translation model (soft-)search for the parts of the source sentence relevant to predicting each target word. (awaiting scientific review)
  • BERT (Devlin and colleagues, 2018) pre-trains deep bidirectional Transformer representations from unlabelled text by jointly conditioning on both left and right context in all layers. (awaiting scientific review)
  • Dosovitskiy and colleagues (2020) describe the Transformer architecture as having become the de facto standard for natural language processing tasks. (awaiting scientific review)
  • Worked calculation (author's own): with key dimension dₖ = 4, a query [1, 0, 1, 0] and keys [1, 0, 1, 0] and [0, 1, 0, 1] give dot products 2 and 0, scaled scores 1 and 0, and softmax weights of about 0.731 and 0.269; with values [10, 0] and [0, 10] the attention output is about [7.31, 2.69]. (awaiting scientific review)

Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.

Content status: published 1 October 2026.

  • Scientific review: this version has not yet been signed off by a scientific reviewer.
  • The Advanced explanation has not yet been reviewed for age suitability.