Skip to main content
ScienceVerse

← Back to Transformers

Transformers quiz

All 11 questions. After you submit, each question shows the right answer and why. Back to the quick check

ExplorerQuestion 1 of 11

The Transformer architecture from 2017 works by passing information through attention rather than by reading words strictly one at a time.

InvestigatorQuestion 2 of 11

In attention, what is the output for a token?

InvestigatorQuestion 3 of 11

By 2020, researchers described the Transformer as the standard architecture for natural-language processing tasks.

InvestigatorQuestion 4 of 11

Why does a transformer need positional encodings?

ScientistQuestion 5 of 11

In scaled dot-product attention, what are the query–key dot products divided by before the softmax, and why?

ScientistQuestion 6 of 11

Two keys give scaled attention scores of 1 and 0. What softmax weight goes to the first key? (3 decimal places)

AdvancedQuestion 7 of 11

In the Transformer decoder, the prediction for a position may attend to later positions in the output so it can see what comes next.

AdvancedQuestion 8 of 11

How does the per-layer computational cost of self-attention grow with sequence length n?

AdvancedQuestion 9 of 11

What distinguishes BERT's pre-training from a left-to-right language model?

ExpertQuestion 10 of 11

In the original Transformer, the model dimension was 512 and there were h = 8 heads. What was the per-head key dimension dₖ?

ExpertQuestion 11 of 11

What bottleneck did Bahdanau et al. (2014) identify in encoder–decoder translation models, and how did their attention mechanism address it?