Skip to content

Latency and Streaming in LLM Applications

Why language model responses feel slow, the metrics that matter, and techniques including streaming to make applications responsive.

Editorial team 2 min read

Users judge an AI application by how quickly it responds. Language models generate text token by token, so latency needs deliberate design.

Metrics That Matter

  • Time to first token (TTFT): how long before any output appears. Dominates perceived speed in chat.
  • Tokens per second: how fast output flows once it starts.
  • Total latency: end to end, including retrieval, tool calls and post-processing.

Streaming

Streaming sends tokens to the user as they're generated rather than waiting for the complete response. It dramatically improves perceived responsiveness for chat and long answers. Structured output that must be validated as a whole may need to wait for completion.

Reducing Latency

  • Use a smaller or faster model where quality allows.
  • Shorten prompts and requested outputs.
  • Use prompt caching for long, repeated prefixes.
  • Run independent steps — retrieval, tool calls — in parallel.
  • Cache responses to repeated questions.
  • Choose a hosting region close to users.
  • Limit the number of agent steps.

Design for Waiting

Show progress indicators, stream partial results, and for long tasks run them in the background and notify users when done.

Measure in Production

Track latency percentiles (not just averages) per feature. The slowest few percent of requests often define the user experience.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026