Moderation
The process of detecting, classifying, and handling harmful, unsafe, or policy-violating content in AI inputs and outputs.
Explore more about Guardrails
Related terms
The process of detecting, blocking, or modifying content that violates safety policies, compliance requirements, or application rules.
Safety FilterA mechanism designed to detect and block toxic, harmful, or policy-violating content in model inputs and outputs.
Input ValidationThe process of verifying and sanitizing user inputs or external data before they are processed by an AI system to ensure they conform to expected formats, constraints, and security policies, helping prevent errors, misuse, prompt injection, and malicious inputs.
Response SanitizationThe process of filtering or removing sensitive, unsafe, or non-compliant content from model outputs before they reach users.