Human labels are an instrument, not an opinion pool
Human evaluation quality is determined before any labeling starts — by the guideline, the task framing, the rater training, and the quality-control stream. A vague guideline produces labels that cannot be aggregated no matter how many raters you add.
Guidelines are the spec
Every decision the rater must make should be documented with examples, including edge cases and tie handling.
Gold items
Seed known-answer items into the stream to detect raters who are not reading or not following the guide.
Adjudication
Resolve disagreements through a defined process rather than majority-voting them into noise.
A rater answering seeded gold items with 60% accuracy is not producing usable labels regardless of how many items they label.
- Annotation platforms run qualification rounds and embed gold items continuously, not just at onboarding.
- Safety data labeling uses detailed policy rubrics with escalation paths for ambiguous cases.