To score a machine translation, you could ask what fraction of its words appear in a reference translation. That falls apart immediately: a system that outputs "the the the the the" against a reference containing "the" scores a perfect 1.0.
Modified n-gram precision closes that hole with clipping. An n-gram can only be credited as many times as it actually appears in a single reference.
Task: write clipped_precision(candidate, references, n) returning that precision, rounded to 4 decimal places.
The procedure:
n consecutive tokens — and count how often each distinct one occurs.min(candidate count, ceiling) for it.candidate is a list of tokens; references is a list of such lists.n it has no n-grams at all. Return 0.0 rather than dividing by zero.Step 2 is the whole idea, and the denominator in step 4 is what keeps it honest: repeats stay in the bottom of the fraction while being clipped out of the top, so padding the output with a common word can only hurt you.