Skip to content

Long-Context Language Models

What large context windows enable, their costs and limitations, and when retrieval is still better.

Editorial team 1 min read

Context windows have grown from a few thousand tokens to hundreds of thousands or more. This changes what's possible.

What It Enables

  • Analysing entire contracts, codebases or books at once.
  • Long conversations and agent sessions.
  • Many examples in a single prompt.
  • Comparing multiple documents directly.

Limitations

  • Cost and latency: more tokens mean higher cost and slower responses.
  • Attention to detail: performance can drop for information buried in very long contexts.
  • Reasoning across the whole: finding a fact is easier than synthesising many scattered facts.

Making It Work

  • Put long documents first and questions last.
  • Label documents clearly.
  • Ask the model to quote relevant passages before answering.
  • Use prompt caching for repeated large contexts.

Long Context Versus Retrieval

  • Long context: simpler, good for one-off analysis of a manageable document set.
  • Retrieval: better for very large collections, frequent queries and cost control.

Many systems combine both: retrieve relevant material into a generous context.

Test

Evaluate on your own documents and questions.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026