Large Language Models
Positional Encoding
Attention has no sense of order. This gives it one.
You can jump straight in, but this star assumes The Transformer. Not recommended as a first stop.
Four layers of depth
Each layer ends with a quiz. Finish layer 4 and you own this concept.
- L1IntuitionAttention has no idea what order words came in unless you tell it. This is the story of how you tell it.6m
40 XP - L2MechanicsThe four schemes, mechanically: how each one actually computes a position signal and injects it into attention.10m
70 XP - L3CodeImplementing sinusoidal, learned, ALiBi, and full rotary position embeddings from scratch.14m
110 XP - L4FoundationsThe rotation matrix behind RoPE, why it produces a genuinely relative dot product, and the arithmetic of extending context.14m
180 XP
33 stars in the atlas.