Content Filtering
Also called: Content Moderation
The process of detecting, blocking, or modifying content that violates safety policies, compliance requirements, or application rules.
Explore more about Guardrails
Related terms
The process of detecting, classifying, and handling harmful, unsafe, or policy-violating content in AI inputs and outputs.
Safety FilterA mechanism designed to detect and block toxic, harmful, or policy-violating content in model inputs and outputs.
Output ValidationThe process of checking AI-generated outputs against expected formats, schemas, rules, or safety requirements before they are accepted or used.
Policy EnforcementThe process of applying defined rules and policies to control, restrict, or permit AI system inputs, outputs, and actions.
Response SanitizationThe process of filtering or removing sensitive, unsafe, or non-compliant content from model outputs before they reach users.