AI ENGVisual Encyclopedia

MODULE 4: AUTOMATIC SCORING · SCENE 11

BLEU/ROUGE and Their Limits

n-gram precision and recall reward surface overlap; they punish correct paraphrase just as hard as incorrect content.

P10.80P20.60P30.40P40.20
CAND LEN8REF LEN10BP · 0.779BLEU · 0.345
GEOMETRIC MEAN OF PRECISIONS × BREVITY PENALTYBLEU · 0.345A PERFECT PARAPHRASE SHARES FEW N-GRAMS — AND THE SHORT CANDIDATE IS PENALIZED AGAIN

High precision on unigrams collapses by 4-grams; a single weak order dominates the geometric mean. Accuracy of the words, not of the thought.

TECHNICAL BREAKDOWNModule 4: Automatic Scoring

BLEU and ROUGE measure surface overlap, not meaning

BLEU scores n-gram precision against one or more references with a brevity penalty; ROUGE scores reference n-gram recall. Both are fast and reproducible, and both are blind to paraphrase: a correct restatement scores as badly as a wrong one. They are diagnostics for surface form, not judges of quality.

BLEU precision

Clip each candidate n-gram count by its reference count, then take the geometric mean of precisions across n.

Brevity penalty

Short candidates are penalized so a model cannot win by emitting only high-precision fragments.

ROUGE recall

Fraction of reference n-grams the candidate covers — favors long, inclusive outputs.

MATHEMATICAL FORMULATION · BLEU
BLEU = BP · exp(Σ w_n · log p_n), BP = min(1, e^(1 − ref/cand))

Precisions [0.8, 0.6, 0.4, 0.2] with equal weights give a geometric mean of about 0.44; a candidate shorter than the reference is then penalized by the brevity factor.

REAL-WORLD PRODUCTION ENGINEERING
  • BLEU remains standard for machine translation but correlates weakly with human judgment on open-ended generation.
  • ROUGE is standard for summarization; extraction-based summaries score higher than abstractive ones because they copy n-grams.