AI ENGVisual Encyclopedia

WORLD 05  EVALUATION

Turn model behavior into evidence you can trust.

From benchmark design to regression gates across 8 modules: the measurement loop, contamination, scoring, judges, agreement, statistics, and production monitoring.

JOURNEY · 8 MODULES · 24 SCENES

From a claim about a model to evidence that holds.

MODULE 01

Eval Foundations

The measurement loop: define the claim, build the instrument, measure, diagnose, repeat.

MODULE 02

Benchmark Design

Anatomy of items, difficulty calibration, discrimination, and designed capability coverage.

MODULE 03

Contamination & Integrity

Leakage detection, canary traps, private and rotating sets that keep the signal honest.

MODULE 04

Automatic Scoring

Exact match, token F1, BLEU/ROUGE limits, and programmatic grading with verifiers.

MODULE 05

LLM-as-Judge

Rubrics, pairwise win rates, and the positional, length, and self-preference biases of the referee.

MODULE 06

Human & Arena Evaluation

Annotation protocols, inter-annotator agreement, and Elo leaderboards from live battles.

MODULE 07

Statistical Rigor

Confidence intervals, paired tests, bootstrap resampling, and the multiple-comparisons tax.

MODULE 08

Regression Gates & Monitoring

Eval-driven CI, prompt and model regression, and detecting drift in production.