Related terms
The tracking and analysis of token usage and consumption patterns to manage costs and performance.
Tokens per SecondTPSA performance metric measuring the rate at which an inference engine generates output tokens per second.
Cost per RequestA metric that measures the average monetary cost incurred to process a single API call, inference, or user request.
Token OptimizationTechniques and strategies aimed at reducing prompt and completion token counts to lower latency and operational costs.
Metrics CollectionThe process of gathering quantitative measurements from AI systems and applications to track performance, reliability, usage, and operational health.