Inference Optimization
The KV Cache
Why generating token 500 is as cheap as token 5.
You can jump straight in, but this star assumes The Transformer. Not recommended as a first stop.
Four layers of depth
Each layer ends with a quiz. Finish layer 4 and you own this concept.
- L1IntuitionWhy generating text one token at a time is wasteful without a cache, and what the KV cache saves.5m
40 XP - L2MechanicsPrefill vs decode bottlenecks, cache growth mechanics, and why GQA and paging exist.8m
70 XP - L3CodeImplement KV cache sizing math and a minimal cached-attention forward pass.12m
110 XP - L4FoundationsArithmetic intensity, the roofline model, and deriving why decode is bandwidth-bound from first principles.15m
180 XP
Where this leads
33 stars in the atlas.