← Back to Transformers
All 11 questions. After you submit, each question shows the right answer and why. Back to the quick check
The Transformer architecture from 2017 works by passing information through attention rather than by reading words strictly one at a time.
In attention, what is the output for a token?
By 2020, researchers described the Transformer as the standard architecture for natural-language processing tasks.
Why does a transformer need positional encodings?
In scaled dot-product attention, what are the query–key dot products divided by before the softmax, and why?
Two keys give scaled attention scores of 1 and 0. What softmax weight goes to the first key? (3 decimal places)
In the Transformer decoder, the prediction for a position may attend to later positions in the output so it can see what comes next.
How does the per-layer computational cost of self-attention grow with sequence length n?
What distinguishes BERT's pre-training from a left-to-right language model?
In the original Transformer, the model dimension was 512 and there were h = 8 heads. What was the per-head key dimension dₖ?
What bottleneck did Bahdanau et al. (2014) identify in encoder–decoder translation models, and how did their attention mechanism address it?