Related terms
A structured testing methodology where adversarial tactics are simulated to discover vulnerabilities, safety flaws, and failure modes in a system.
Jailbreak TestingA security testing method that evaluates whether an AI system can be manipulated into bypassing its safety policies, behavioral constraints, or security guardrails through adversarial prompts or other attack techniques.
Safety TestingThe process of evaluating AI systems to identify vulnerabilities, harmful outputs, and policy violations before deployment.
RobustnessA metric measuring an AI system's ability to maintain performance and stability under varied, unseen, or adversarial conditions.
Threat ModelingA systematic process for identifying, analyzing, and prioritizing potential security threats and vulnerabilities in a system.