Users judge an AI application by how quickly it responds. Language models generate text token by token, so latency needs deliberate design.
Metrics That Matter
- Time to first token (TTFT): how long before any output appears. Dominates perceived speed in chat.
- Tokens per second: how fast output flows once it starts.
- Total latency: end to end, including retrieval, tool calls and post-processing.
Streaming
Streaming sends tokens to the user as they're generated rather than waiting for the complete response. It dramatically improves perceived responsiveness for chat and long answers. Structured output that must be validated as a whole may need to wait for completion.
Reducing Latency
- Use a smaller or faster model where quality allows.
- Shorten prompts and requested outputs.
- Use prompt caching for long, repeated prefixes.
- Run independent steps — retrieval, tool calls — in parallel.
- Cache responses to repeated questions.
- Choose a hosting region close to users.
- Limit the number of agent steps.
Design for Waiting
Show progress indicators, stream partial results, and for long tasks run them in the background and notify users when done.
Measure in Production
Track latency percentiles (not just averages) per feature. The slowest few percent of requests often define the user experience.