Related terms
The time delay between an AI system receiving a request and producing a response or completing an operation.
Response TimeThe total duration required for an AI system or service to process an input query and return the completed output.
ThroughputThe amount of data, tokens, or requests processed by a system within a specified period.
Tokens per SecondTPSA performance metric measuring the rate at which an inference engine generates output tokens per second.
Cost per RequestA metric that measures the average monetary cost incurred to process a single API call, inference, or user request.