Prompt Shield
A guardrail that detects or blocks malicious prompt content and instruction-based attacks before they can influence an AI system.
Explore more about Guardrails
Related terms
A guardrail that detects attempts to manipulate an AI system through malicious or unintended instructions embedded in prompts or external content.
Jailbreak DetectionThe process of identifying prompts or interactions that attempt to bypass an AI system's safety policies, behavioral constraints, or security controls, enabling the system to block, flag, or mitigate unauthorized or unsafe requests.
Input ValidationThe process of verifying and sanitizing user inputs or external data before they are processed by an AI system to ensure they conform to expected formats, constraints, and security policies, helping prevent errors, misuse, prompt injection, and malicious inputs.
Content FilteringThe process of detecting, blocking, or modifying content that violates safety policies, compliance requirements, or application rules.
Safety FilterA mechanism designed to detect and block toxic, harmful, or policy-violating content in model inputs and outputs.
Policy EnforcementThe process of applying defined rules and policies to control, restrict, or permit AI system inputs, outputs, and actions.