WORLD 05 EVALUATION
Turn model behavior into evidence you can trust.
From benchmark design to regression gates across 8 modules: the measurement loop, contamination, scoring, judges, agreement, statistics, and production monitoring.
JOURNEY · 8 MODULES · 24 SCENES
From a claim about a model to evidence that holds.
Eval Foundations
The measurement loop: define the claim, build the instrument, measure, diagnose, repeat.
01 · Loop
The Evaluation Loop
Measure → diagnose → change → re-measure. Evaluation is not a final report; it is the control loop that tells you whether the system moved.
OPEN EXHIBIT →02 · Axis
Capability vs Alignment Evals
Can it do the task, and does it do it the way we intended? Two axes that fail independently and need separate instruments.
OPEN EXHIBIT →03 · Suite
Static Suites vs Dynamic Arenas
Frozen item banks are reproducible but leak; live arenas resist contamination but drift with the crowd. Serious reports carry both.
OPEN EXHIBIT →Benchmark Design
Anatomy of items, difficulty calibration, discrimination, and designed capability coverage.
04 · Item
Anatomy of a Benchmark Item
Prompt, reference, scoring rule, metadata, and source. Every field is a design decision that shapes what the score means.
OPEN EXHIBIT →05 · Calibrate
Difficulty & Discrimination
Item difficulty p and discrimination d: why floor and ceiling items waste budget and what a useful spread looks like.
OPEN EXHIBIT →06 · Coverage
Capability Coverage & Taxonomy
Map items to capability leaves before writing them. Balanced coverage prevents a single easy skill from carrying the score.
OPEN EXHIBIT →Contamination & Integrity
Leakage detection, canary traps, private and rotating sets that keep the signal honest.
07 · Leakage
Contamination Detection
n-gram overlap, embedding proximity, and perplexity tells that reveal whether eval items leaked into training data.
OPEN EXHIBIT →08 · Canary
Canary Strings & Trap Items
Plant uniquely identifiable markers and impossible items so memorization and gaming show up as anomalies, not silent score inflation.
OPEN EXHIBIT →09 · Hygiene
Private & Rotating Sets
Hold out private splits, refresh on a cadence, and version eval sets so today's number is comparable to yesterday's.
OPEN EXHIBIT →Automatic Scoring
Exact match, token F1, BLEU/ROUGE limits, and programmatic grading with verifiers.
10 · String
Exact Match & Token F1
Normalization first, then exact match or token-level F1 — cheap, deterministic, and blind to meaning it was not told to check.
OPEN EXHIBIT →11 · Overlap
BLEU/ROUGE and Their Limits
n-gram precision and recall reward surface overlap; they punish correct paraphrase just as hard as incorrect content.
OPEN EXHIBIT →12 · Verifier
Execution & Programmatic Grading
Unit tests, regexes, and checkers turn open-ended output into a pass/fail oracle — the backbone of code and math eval.
OPEN EXHIBIT →LLM-as-Judge
Rubrics, pairwise win rates, and the positional, length, and self-preference biases of the referee.
13 · Rubric
Rubric Design & Pointwise Scoring
Criteria, anchors, and output schema. A judge is only as good as the rubric that tells it what a 3 means versus a 4.
OPEN EXHIBIT →14 · Pairwise
Pairwise Comparison & Win Rates
Compare two answers instead of scoring one. Lower variance, harder aggregation, and the Bradley-Terry bridge to Elo.
OPEN EXHIBIT →15 · Bias
Judge Bias: Position, Length, Self
The referee has preferences. Swap order, control for length, avoid same-family judges — then measure what is left.
OPEN EXHIBIT →Human & Arena Evaluation
Annotation protocols, inter-annotator agreement, and Elo leaderboards from live battles.
16 · Protocol
Annotation Protocols & Guidelines
Instructions, gold items, adjudication, and rater calibration decide whether human labels are a measurement or an opinion pool.
OPEN EXHIBIT →17 · Agreement
Inter-Annotator Agreement
Percent agreement is inflated by chance; Cohen's kappa subtracts it. If raters cannot agree, the task itself may be ill-defined.
OPEN EXHIBIT →18 · Elo
Arena Elo & Leaderboards
Live battles, Elo updates, and confidence around a rank — why an Elo gap is a distribution, not a fact.
OPEN EXHIBIT →Statistical Rigor
Confidence intervals, paired tests, bootstrap resampling, and the multiple-comparisons tax.
19 · CI
Confidence Intervals for Scores
Every accuracy is a draw from a sampling distribution. Report the interval and the n that produced it, or the score is a rumor.
OPEN EXHIBIT →20 · Paired
Paired Tests & McNemar
Same items, two models: compare the discordant pairs. Paired tests are sharper than two independent confidence intervals.
OPEN EXHIBIT →21 · Resample
Bootstrap & Multiple Comparisons
Resample to build intervals without distributional assumptions — then pay the tax for every additional comparison you make.
OPEN EXHIBIT →Regression Gates & Monitoring
Eval-driven CI, prompt and model regression, and detecting drift in production.
22 · Gate
Eval-Driven CI Gates
Per-capability regression budgets turn evaluation into a release gate: no ship when a tested axis falls past its allowance.
OPEN EXHIBIT →23 · Golden
Prompt & Model Regression Testing
A frozen golden set catches the quiet breakage a prompt tweak or fine-tune causes before users do.
OPEN EXHIBIT →24 · Drift
Online Monitoring & Drift Detection
Score distributions move as inputs and populations shift. Sample production, track drift, and gate deploys on live evidence.
OPEN EXHIBIT →