Model A scoreon 500 test questions
Model B scoreon 500 test questions
A benchmark score is a sample statistic, not a ground truth — it's computed from a finite set of questions, and if you re-sampled a different set of questions from the same distribution, you'd get a somewhat different score just from randomness. Treating a benchmark score as if it had zero uncertainty is one of the most common evaluation mistakes.