The context window

A context window is a hard limit on how many tokens one call may carry, counting the input and the generated answer together. Every part of an assembled context is a claim against that limit, and the claims arrive from different systems with no awareness of each other. Fitting them together is the last step before a call goes out, and it decides what the model never gets to see.

Window sizes went from 2,048 tokens on the model that introduced few shot prompting in 2020 to a million on several models sold today. That is a factor of roughly five hundred in five years, and it changed which failure a team meets first. A window of 2,048 tokens produced an error, at a point in the code where somebody had to deal with it. A window of a million tokens produces an invoice.

The budget is therefore an engineering artefact with an owner, a number and a metric, in the same way a latency budget is. Teams that never write one still have a budget, and it is whatever the sum of the parts happened to be this morning.

The sections below name the claims that compete for one window, work a single assistant's budget twice at six months apart, cover the room an answer needs, and finish with what a longer window does and does not fix.

The claims that compete for one window

The window, 128,000 tokens50,300 used77,700 unused, and charged for in neither money nor timeThose 50,300 tokens, expanded to full widthHistory21,000Retrieved10,800In order along the barSystem 4,800 Tools 4,400 Demonstrations 5,400 History 21,000Retrieved 10,800 Tool results 2,100 Question 300 Answer 1,500

One call on a support assistant six months after launch. Conversation history and retrieved passages are seven parts in ten between them, and the customer's actual question is 300 tokens of the 50,300. The unused three fifths of the window are the reason nothing raised an error.

Seven claims compete, and each one grows for a different reason.

  1. System instructions and policy grow after every incident, because adding a line is the fastest available response to a complaint.
  2. Tool definitions grow with every new capability, since each tool brings a name, a description and a schema that all sit in the window.
  3. Few shot demonstrations grow because adding one more example is the cheapest answer to a misclassification.
  4. Conversation history grows by use alone, with nobody deciding anything.
  5. Retrieved passages grow when somebody raises the number of chunks kept after a miss, and the raise is rarely lowered again.
  6. Tool results grow when a tool returns a list and the list gets longer.
  7. The user's message is the only one nobody controls and the only one that stays small.

Every growth in that list was a locally correct decision made by somebody solving a real problem. None of them was a decision about the total, because the total had no owner.

One budget worked twice

The assistant behind the figure runs on a window of 128,000 tokens. Its budget at launch and its budget six months later look like this.

Claim on the windowAt launchSix months later
System instructions and policy1,2004,800
Tool definitions1,600 for four tools4,400 for eleven tools
Few shot demonstrations1,800 for four5,400 for twelve
Conversation history2,40021,000 across forty turns
Retrieved passages3,600, four at 90010,800, twelve at 900
Tool results this turn7002,100
The user's message300300
Input total11,60048,800
Reserved for the answer1,5001,500
Tokens on one call13,10050,300

Nothing overflowed. The input grew by a factor of 4.2 and the window never came close to filling, since 50,300 of 128,000 is under two fifths. At ten thousand requests a day the same feature now sends 488 million input tokens a day where it used to send 116 million, on answers nobody has shown to be any better. Nothing in the system raised anything, because nothing in the system was watching a number that nobody had set.

A budget makes the number explicit and gives every claim an allowance.

# A budget with an owner, checked before every call goes out.
BUDGET = {
    "system":        5_000,
    "tools":         4_500,
    "examples":      5_500,
    "history":      12_000,
    "retrieved":    12_000,
    "tool_results":  4_000,
    "question":      1_000,
    "answer":        6_000,
}   # 50,000 of a 128,000 window, chosen rather than accumulated

def fit(parts, budget=BUDGET):
    """Cut each part to its own allowance. Fail the call outright before
    quietly shipping a context with the policy missing from it."""
    out = {}
    for name, text in parts.items():
        allowance, cost = budget[name], count_tokens(text)
        if cost <= allowance:
            out[name] = text
        elif name in ("history", "tool_results"):
            out[name] = summarise_to(text, allowance)
        elif name == "retrieved":
            out[name] = drop_lowest_ranked(text, allowance)
        else:
            raise BudgetExceeded(name, cost, allowance)
    emit_metric("context_tokens", sum(count_tokens(t) for t in out.values()))
    return out

The final branch is the one that earns the function. A system prompt over its allowance is a defect in the prompt, so the call fails loudly and somebody fixes it. History and tool results are compressible, and retrieved passages are reselectable, which is why those three have their own treatment.

The metric on the last line is what makes any of this visible later. Written out per request, it turns a doubled bill from a mystery into a line somebody can read.

One record per call, measured before fit cuts anything.

request=7f21c9  window=128000
  measured   system 4812    tools 4398     examples 5402
             history 20961  retrieved 10783
             tool_results 2094  question 297
             input 48747
  budget     history allowance 12000, over by 8961
  action     history summarised to 11840
  sent       input 39626   output 812   total 40438

The three numbers worth alerting on are in there. Input tokens per request shows the bill, the count of allowance breaches per hour shows which claim is growing, and the identity of the claim that breached, which is history on this call, names the part of the system somebody has to change.

The room the answer needs

Input and output share the window on most interfaces, so the space for the answer has to come out of the budget before anything else is assembled. A context packed to the last token leaves the model nothing to write with, and the failure is a reply that stops in the middle of a sentence.

Two settings control this and they do different jobs. The maximum output token count is a cap, and hitting it truncates mid sentence with no warning in the body of the reply. The reserve inside the budget is an allocation, and it is what stops the assembly from spending the tokens the cap was counting on. A feature that sets a cap and never reserves has two components disagreeing about the same tokens.

On some models the reserve has to cover more than the visible reply. A model that produces its working before answering spends output tokens on that working, and the working can run several times longer than the answer it precedes. So a feature that reserved 1,500 tokens for a two paragraph reply, and then moved to a model that reasons first, begins truncating on the day of the move with nothing else about it having changed.

Why a longer window moves the problem

A longer window removes the error and leaves three costs standing.

Money. Input tokens are charged on every call, and a context four times larger is four times the input charge for the same answer. Page 16 works this through with prices.

Time. The model processes the whole input before it emits the first token, and the attention work grows with the square of the sequence length, so a long context shows up directly in the time a user waits before anything appears.

Accuracy. Performance falls as input length grows, including on tasks a model handles perfectly at short lengths. The material fits. It is simply used less well, and that finding has a name, a body of evidence behind it and a page of its own two pages from here.

A budget under pressure needs somewhere to give. There are exactly four things a team can do with a context that has grown too large, and they have been named and taught as a set since the middle of 2025.

Common misconceptions

“A window of a million tokens removes the need for a budget.”

It removes the error. Every token still costs money on every call and adds to the time before the first word appears, and accuracy falls as input length grows on tasks that are otherwise trivial. A larger window converts a loud failure into a quiet and expensive one.

“The context window is the model's memory.”

Nothing persists between calls. The window is reassembled from scratch for every request, and a conversation that appears to remember is one where the application resent the earlier turns. Anything not put back in is gone, which is why persistence is a design decision somebody has to make.

Where this is examined
Prompt and Context Engineering
Engineering the Context, 18 per cent of the exam.
Related material
Concepts