Atlas

Inference Optimization

The KV Cache

Why generating token 500 is as cheap as token 5.

You can jump straight in, but this star assumes The Transformer. Not recommended as a first stop.

Four layers of depth

Each layer ends with a quiz. Finish layer 4 and you own this concept.

  1. L1IntuitionWhy generating text one token at a time is wasteful without a cache, and what the KV cache saves.5m
    40 XP
  2. L2MechanicsPrefill vs decode bottlenecks, cache growth mechanics, and why GQA and paging exist.8m
    70 XP
  3. L3CodeImplement KV cache sizing math and a minimal cached-attention forward pass.12m
    110 XP
  4. L4FoundationsArithmetic intensity, the roofline model, and deriving why decode is bandwidth-bound from first principles.15m
    180 XP

Where this leads

33 stars in the atlas.