Related terms
The process of verifying that multi-step AI tasks and execution flows operate correctly, reliably, and as intended.
Tool TestingThe process of verifying that external tools, APIs, and function integrations behave correctly when invoked by an AI system.
Functional TestingA software testing method that verifies whether an AI application, agent, or system performs its intended functions correctly by validating its behavior against specified functional requirements and expected outcomes.
Red TeamingA structured testing methodology where adversarial tactics are simulated to discover vulnerabilities, safety flaws, and failure modes in a system.
Human EvaluationAn evaluation method in which human reviewers assess the quality of AI system outputs against defined criteria such as correctness, relevance, helpfulness, safety, or fluency, providing judgments that complement or validate automated evaluation metrics.