Screen 1 of 3 - Read before you experiment
Attention constructs a new representation for every token
Each token is projected into Query, Key, and Value vectors. Query-Key dot products create compatibility scores, scaling controls their magnitude, and softmax converts them into attention weights.
A causal mask removes future positions before softmax. The resulting weights combine Value vectors into a context vector that represents the selected token in its current sequence.
Evidence levelExact deterministic educational forward pass, not pretrained GPT weights
Screen 2 of 3 - Experiment
Run, inspect, and compare
Follow the three guided moves above. Change one variable at a time so every visual change has a clear cause.
Screen 3 of 3 - Consolidate
By the end of this lesson, you will be able to:
- Follow tokens through embedding and Query, Key, and Value projections.
- Explain scaled dot-product attention and why a causal mask blocks future tokens.
- Read attention weights and construct the weighted context vector for a selected token.
Explains shift from RNN to self-attention. Introduces Query, Key, Value, and weighted sums. Outlines Transformer encoder-decoder and multi-head attention.