Safety Filter
A mechanism designed to detect and block toxic, harmful, or policy-violating content in model inputs and outputs.
Explore more about Guardrails
Related terms
The process of detecting, blocking, or modifying content that violates safety policies, compliance requirements, or application rules.
Input ValidationThe process of verifying and sanitizing user inputs or external data before they are processed by an AI system to ensure they conform to expected formats, constraints, and security policies, helping prevent errors, misuse, prompt injection, and malicious inputs.
Output ValidationThe process of checking AI-generated outputs against expected formats, schemas, rules, or safety requirements before they are accepted or used.
ModerationThe process of detecting, classifying, and handling harmful, unsafe, or policy-violating content in AI inputs and outputs.
Prompt ShieldA guardrail that detects or blocks malicious prompt content and instruction-based attacks before they can influence an AI system.
Rule EngineA software component that executes predefined logical rules to validate, filter, or control AI inputs and outputs.