Related terms
The continuous tracking and analysis of infrastructure, API, and model usage costs to optimize spending and detect unexpected expenses.
Cost OptimizationThe process of reducing infrastructure, model, and operational costs while maintaining or improving application performance, reliability, and quality.
Token UsageThe total volume of prompt and completion tokens consumed during interactions with a language model.
Response TimeThe total duration required for an AI system or service to process an input query and return the completed output.
LatencyThe time delay between an AI system receiving a request and producing a response or completing an operation.