Jailbreak Testing
Also called: Jailbreak Evaluation, Prompt Jailbreak Testing
A security testing method that evaluates whether an AI system can be manipulated into bypassing its safety policies, behavioral constraints, or security guardrails through adversarial prompts or other attack techniques.
Explore more about Testing
Related terms
The process of identifying prompts or interactions that attempt to bypass an AI system's safety policies, behavioral constraints, or security controls, enabling the system to block, flag, or mitigate unauthorized or unsafe requests.
Prompt InjectionAn attack that inserts malicious or unintended instructions into prompts or external content to manipulate an AI system's behavior.
Red TeamingA structured testing methodology where adversarial tactics are simulated to discover vulnerabilities, safety flaws, and failure modes in a system.
Agent GuardrailsPolicies, constraints, and runtime controls that keep an AI agent operating within defined safety, security, and behavioral boundaries.