Screen 1 of 3 - Read before you experiment

Attention constructs a new representation for every token

Each token is projected into Query, Key, and Value vectors. Query-Key dot products create compatibility scores, scaling controls their magnitude, and softmax converts them into attention weights.

A causal mask removes future positions before softmax. The resulting weights combine Value vectors into a context vector that represents the selected token in its current sequence.

Evidence levelExact deterministic educational forward pass, not pretrained GPT weights
Screen 2 of 3 - Experiment

Run, inspect, and compare

Follow the three guided moves above. Change one variable at a time so every visual change has a clear cause.

Continue to the attention explanation
Screen 3 of 3 - Consolidate

By the end of this lesson, you will be able to:

  • Follow tokens through embedding and Query, Key, and Value projections.
  • Explain scaled dot-product attention and why a causal mask blocks future tokens.
  • Read attention weights and construct the weighted context vector for a selected token.

Explains shift from RNN to self-attention. Introduces Query, Key, Value, and weighted sums. Outlines Transformer encoder-decoder and multi-head attention.