Atlas

Large Language Models

Attention

Every token looks at every other token and decides what matters.

You can jump straight in, but this star assumes Embeddings and Training & Optimization. Not recommended as a first stop.

Four layers of depth

Each layer ends with a quiz. Finish layer 4 and you own this concept.

  1. L1IntuitionThe mechanism that lets every token look at every other token and decide what matters — the core idea underneath every transformer.6m
    40 XP
  2. L2MechanicsThe exact scaled dot-product formula, why the scaling factor exists, and how masking and multi-head splitting actually work mechanically.10m
    70 XP
  3. L3CodeImplement scaled dot-product and multi-head attention from scratch in PyTorch, and see how FlashAttention changes only the implementation, not the math.15m
    110 XP
  4. L4FoundationsThe full derivation of the scaling term, attention's exact FLOP and memory cost, and why FlashAttention's arithmetic-intensity argument works on real hardware.13m
    180 XP

Where this leads

33 stars in the atlas.