Clipped n-gram precision scores one n-gram size. BLEU combines several of them and then patches the one hole precision alone can't cover.
That hole: precision never punishes you for saying too little. Output the single word "the" against a twenty-word reference and your 1-gram precision is a perfect 1.0.
So BLEU has two parts.
The geometric mean of the precisions, for every n from 1 to max_n:
A geometric mean rather than an arithmetic one, because it is unforgiving: if any single p_n is zero, the whole product is zero. Fluency at every n-gram size, or nothing.
The brevity penalty, where c is the candidate's length and r the reference length:
BLEU is the two multiplied together.
Task: write bleu(candidate, references, max_n) returning the score, rounded to 4 decimal places.
n, exactly as in the n-gram precision problem.0, return 0.0 immediately — you can't take the log of zero. Same if the candidate is too short to have n-grams at some n.r is the closest reference length to c. If two references are equally close, use the shorter one.BP is capped at 1.Notice that BLEU scores a corpus far better than a sentence — on a single sentence one missing 4-gram zeroes the whole thing. That brutality is deliberate, and it's also why BLEU is computed over whole test sets in practice.