Toxicity

Harmful, offensive, or abusive content in model output.

Toxicity refers to model outputs that contain harmful, offensive, hateful, or abusive language, including content that targets individuals or groups based on protected characteristics. On the AIF-C01 exam, toxicity is a core responsible AI concern because generative models trained on large web corpora can reproduce and amplify such language at scale. Amazon Bedrock addresses toxicity through Guardrails, a feature that applies content filters at configurable strength levels across categories such as hate, insults, sexual content, violence, and misconduct. These filters evaluate both user prompts (input) and model responses (output), enforcing safety policies consistently across foundation models.

PlayPrepHQ study notes are written and reviewed against primary exam sources. How we create & review content →

Related terms

Back to Guidelines for Responsible AI