Related terms
The time delay between an AI system receiving a request and producing a response or completing an operation.
Tokens per SecondTPSA performance metric measuring the rate at which an inference engine generates output tokens per second.
Response TimeThe total duration required for an AI system or service to process an input query and return the completed output.
Time to First TokenTTFTThe duration of time it takes for a language model to produce its initial output token after receiving a request.
Cost per RequestA metric that measures the average monetary cost incurred to process a single API call, inference, or user request.