AI ENGVisual Encyclopedia

MODULE 6: HUMAN & ARENA EVALUATION · SCENE 16

Annotation Protocols & Guidelines

Instructions, gold items, adjudication, and rater calibration decide whether human labels are a measurement or an opinion pool.

GOLD-ITEM GATEKEEP ≥ 0.803/6 RATERS · 50%
rater-010.98 rater-020.92 rater-030.85 rater-040.72 rater-050.61 rater-060.55 GOLD LINE = QUALIFICATION · FAILING RATERS PRODUCE OPINIONS, NOT LABELS

Raise the gate and the pool shrinks but the labels harden. Adjudicate what remains instead of averaging it away.

TECHNICAL BREAKDOWNModule 6: Human & Arena Evaluation

Human labels are an instrument, not an opinion pool

Human evaluation quality is determined before any labeling starts — by the guideline, the task framing, the rater training, and the quality-control stream. A vague guideline produces labels that cannot be aggregated no matter how many raters you add.

Guidelines are the spec

Every decision the rater must make should be documented with examples, including edge cases and tie handling.

Gold items

Seed known-answer items into the stream to detect raters who are not reading or not following the guide.

Adjudication

Resolve disagreements through a defined process rather than majority-voting them into noise.

MATHEMATICAL FORMULATION · RATER RELIABILITY GATE
keep_rater if gold_accuracy ≥ threshold

A rater answering seeded gold items with 60% accuracy is not producing usable labels regardless of how many items they label.

REAL-WORLD PRODUCTION ENGINEERING
  • Annotation platforms run qualification rounds and embed gold items continuously, not just at onboarding.
  • Safety data labeling uses detailed policy rubrics with escalation paths for ambiguous cases.