The Transformer

Layer 4 · Foundations

The Transformer

The full parameter and FLOP accounting for a dense transformer, and why arithmetic intensity makes autoregressive decoding fundamentally different from training.

13 min read180 XP

Fix a dense decoder-only transformer with layers, hidden size , vocabulary size , and (for this derivation) an MLP expansion ratio of 4. The full parameter count, ignoring the small normalization weight vectors: