Evaluation

Layer 4 · Foundations

Evaluation

Statistical rigor in evaluation: confidence intervals, sample size, and why small benchmark gaps are often noise.

15 min read180 XP

Model A scoreon 500 test questions
Model B scoreon 500 test questions
Is a 2-point gap on 500 questions actually meaningful, or noise? The math below answers that.

A benchmark score is a sample statistic, not a ground truth — it's computed from a finite set of questions, and if you re-sampled a different set of questions from the same distribution, you'd get a somewhat different score just from randomness. Treating a benchmark score as if it had zero uncertainty is one of the most common evaluation mistakes.