Definition
A context window is the most text, measured in tokens, that a language model can consider at once, including instructions, documents and its reply.
Context window explained
Language models process text as tokens, which are chunks of words. Everything the model uses to produce an answer must fit in its context window: the system instructions, any documents or search results provided, the conversation history, and the answer it writes. Context windows have grown enormously, and larger models can now take in long documents at once.
A large context window doesn't remove the need for good design:
- Cost and speed grow with the amount of text sent.
- Models can pay less attention to information buried in the middle of very long inputs.
- Irrelevant material can distract the model and lower accuracy.
That is why retrieval, sending only the most relevant passages, is still standard practice. For AI search, it also explains why concise, self-contained passages are easier for assistants to use than long, rambling pages: the system has to choose what to include in a limited space.
Example
Your team tries to answer contract questions by pasting a whole contract library into an assistant, and answers get slow and vague. Retrieving only the clauses relevant to each question keeps the context window focused and answers become faster and more accurate.
Why it matters
The context window shapes what an AI tool can take into account, which affects accuracy, cost and how your content is used in AI answers.
Related service
Agentic AI Tools
Custom AI agents and automations that connect to your data and tools, with a human in the loop.
Custom agentic AI toolsPublished by Vidern, founded and led by Malhar Shah. Updated .
Have a process an AI agent could run?
Describe the workflow and we'll reply within 1 business day with what we'd build.
Related terms
- LLM (Large Language Model)An LLM (large language model) is an AI model trained on huge amounts of text to generate language, such as the models behind ChatGPT, Claude and Gemini.
- RAG (Retrieval-Augmented Generation)RAG (retrieval-augmented generation) is when an AI system first retrieves relevant documents, then has a language model write its answer from them.
- EmbeddingsEmbeddings are lists of numbers that represent the meaning of text, so software can find passages with similar meaning even when they use different words.
- Prompt engineeringPrompt engineering is designing the instructions, context and examples given to a language model so it produces accurate, consistent output for a task.