Large Language Models
Attention
Every token looks at every other token and decides what matters.
You can jump straight in, but this star assumes Embeddings and Training & Optimization. Not recommended as a first stop.
Four layers of depth
Each layer ends with a quiz. Finish layer 4 and you own this concept.
- L1IntuitionThe mechanism that lets every token look at every other token and decide what matters — the core idea underneath every transformer.6m
40 XP - L2MechanicsThe exact scaled dot-product formula, why the scaling factor exists, and how masking and multi-head splitting actually work mechanically.10m
70 XP - L3CodeImplement scaled dot-product and multi-head attention from scratch in PyTorch, and see how FlashAttention changes only the implementation, not the math.15m
110 XP - L4FoundationsThe full derivation of the scaling term, attention's exact FLOP and memory cost, and why FlashAttention's arithmetic-intensity argument works on real hardware.13m
180 XP
Where this leads
33 stars in the atlas.