Related terms
The time delay between an AI system receiving a request and producing a response or completing an operation.
Time to First TokenTTFTThe duration of time it takes for a language model to produce its initial output token after receiving a request.
ThroughputThe amount of data, tokens, or requests processed by a system within a specified period.
Latency MonitoringThe practice of tracking response latency over time to detect slowdowns, performance degradation, and service issues.
Latency OptimizationThe practice of reducing response latency in AI systems to improve responsiveness, throughput, and user experience.