Atlas

Harness Engineering

Evaluation

Benchmarks lie. Building evals that don't.

You can jump straight in, but this star assumes Fine-tuning. Not recommended as a first stop.

Four layers of depth

Each layer ends with a quiz. Finish layer 4 and you own this concept.

  1. L1IntuitionWhy 'is this model good' is a much harder question than it sounds, and how the field tries to answer it.5m
    40 XP
  2. L2MechanicsHow benchmarks, LLM-as-judge, and preference-based leaderboards are actually constructed and scored.8m
    70 XP
  3. L3CodeImplement a reference-based scorer, an LLM-as-judge pairwise comparator, and a simple Elo update.12m
    110 XP
  4. L4FoundationsStatistical rigor in evaluation: confidence intervals, sample size, and why small benchmark gaps are often noise.15m
    180 XP

33 stars in the atlas.