Human Evaluation
Having people rate model output for quality, relevance, or safety.
Human evaluation has trained raters or domain experts directly assess foundation model outputs for fluency, factual accuracy, relevance, and safety — dimensions automated metrics like BLEU or ROUGE cannot reliably capture. In Amazon Bedrock, its model evaluation jobs support human-based reviews, which matter most during model selection and fine-tuning validation. The key exam distinction is human evaluation versus human-in-the-loop: human evaluation is a periodic quality assessment that informs model readiness, while human-in-the-loop embeds ongoing human review into live production workflows. It complements, rather than replaces, runtime oversight.
PlayPrepHQ study notes are written and reviewed against primary exam sources. How we create & review content →