Fix a dense decoder-only transformer with layers, hidden size , vocabulary size , and (for this derivation) an MLP expansion ratio of 4. The full parameter count, ignoring the small normalization weight vectors:
Layer 4 · Foundations
The Transformer
The full parameter and FLOP accounting for a dense transformer, and why arithmetic intensity makes autoregressive decoding fundamentally different from training.
13 min read180 XP