Skip to content

What is Context Window?

Generative AI & LLMs, explained by the engineers who build it. Definition, how it works, use cases and common questions.

Context Window definition

A context window is the maximum amount of text, measured in tokens, that a large language model can consider at once, including the system prompt, conversation history, supplied documents and the model's own response. Anything outside the window is invisible to the model, so its size limits how much information a single request can use.

How does a context window work?

A language model is stateless: it remembers nothing between requests. Everything it should take into account must be sent inside one request, and all of it, plus the response it generates, has to fit within the context window. Chat applications create the illusion of memory by resending the conversation history with every new message, which is why long chats eventually hit the limit and older messages are dropped or summarized.

Windows are limited because standard transformer attention compares every token with every other token, so compute and memory grow quickly with length. Engineering advances have pushed windows from about 2,000 tokens in early models such as GPT-3 to hundreds of thousands to a million tokens or more in current frontier and open models. Larger windows still cost more per request and respond more slowly when filled.

What fills a context window

  • System prompt: role, rules and output format.
  • Tool definitions: the names, descriptions and schemas of available functions.
  • Retrieved content: documents or records supplied by a RAG pipeline.
  • Conversation history: earlier user messages and model replies.
  • Tool results: data returned by API calls during an agent run.
  • The response: output tokens count against the same limit.

Long context vs RAG

Big windows tempt teams to paste in everything. That works for one long contract or a single codebase file set, but it has costs. Every token is billed on every request, latency rises, and research such as the "Lost in the Middle" study found that models tend to use information at the start and end of a long input more reliably than information buried in the middle.

Retrieval-augmented generation sends only the passages relevant to each question, which keeps requests fast and cheap and enforces document permissions. The two approaches work well together: retrieval selects the right material, and a generous window lets the system include complete sections rather than fragments.

How to manage context in applications

Managing what goes into the window, sometimes called context engineering, is one of the main levers for quality in LLM applications. The goal is the smallest set of tokens that contains everything the model needs for this particular step. Less noise also means fewer distractions for the model.

  • Summarize or trim old conversation turns instead of resending them in full.
  • Retrieve fewer, more relevant chunks and rerank them.
  • Store long-term facts in a memory store and fetch them when relevant.
  • Place critical instructions at the beginning and restate key constraints near the end.
  • Use prompt caching for long, stable prefixes such as policies or tool lists.

Example: reviewing a long contract

A legal team asks an assistant to compare a 150-page supply agreement with its standard terms. Sending both documents in full fits many modern windows, but answers about clauses deep in the middle are less reliable. Splitting the task by section, comparing each against the matching standard clause and then combining the findings gives more consistent results. Nexzem designs this kind of context strategy as part of every LLM application it builds.

Context Window: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What happens when you exceed the context window?

The API returns an error, or the application silently drops content to fit, usually the oldest messages. Either way the model cannot see what was removed. Well-built applications count tokens before sending and summarize or trim deliberately, so important instructions and facts are never cut by accident.

Does a bigger context window mean a better model?

Not by itself. A large window lets a model accept more input, but how well it uses that input varies between models. Test with your own long documents and check whether answers stay accurate for information located in the middle, and compare cost and latency at realistic input sizes.

Is the context window the same as memory?

No. The context window is short-term working space for a single request. Long-term memory in AI applications is built separately, by storing facts or conversation summaries in a database and inserting the relevant ones into the context window when they are needed.

Keep exploring the generative ai & llms glossary

Need Context Window in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.