Related terms
The total duration required for an AI system or service to process an input query and return the completed output.
Time to First TokenTTFTThe duration of time it takes for a language model to produce its initial output token after receiving a request.
Latency MonitoringThe practice of tracking response latency over time to detect slowdowns, performance degradation, and service issues.
Latency OptimizationThe practice of reducing response latency in AI systems to improve responsiveness, throughput, and user experience.