BLEU
A precision-oriented metric for evaluating machine translation quality.
BLEU (Bilingual Evaluation Understudy) scores how well generated text matches one or more human reference texts by counting overlapping n-grams — typically sequences of one to four consecutive words — and applying a brevity penalty to discourage overly short outputs. Scores range from 0 to 1, where 1.0 means an exact match. On AIF-C01 it appears when evaluating foundation model outputs, such as comparing candidate translations against references in Amazon Bedrock workflows. Remember that BLEU emphasizes precision, making it well suited to translation, whereas ROUGE emphasizes recall, making it better for summarization.
PlayPrepHQ study notes are written and reviewed against primary exam sources. How we create & review content →