FLOPs: dense vs. sparse feed-forward layers
For a dense feed-forward block with hidden size and expansion factor 4 (standard transformer ratio), per-token FLOPs are approximately:
Layer 4 · Foundations
FLOPs accounting for sparse layers, the auxiliary loss objective in full, and the memory-bandwidth reality of running MoE at inference time.
For a dense feed-forward block with hidden size and expansion factor 4 (standard transformer ratio), per-token FLOPs are approximately: