The work that begins once one prompt stops being enough. What changed when the field renamed itself, the context window as a budget with competing claims on it, the four things a team can do with that budget, the ways a long context degrades rather than fails, and retrieval as the main way relevant material gets in.
A pet supplies retailer measures one order question end to end. The assembled context comes to 1,262 tokens, of which the system instruction is 180. The rest is tool definitions, two policy clauses kept from eight retrieved, a customer memory record, three turns of history with the earlier ones summarised, one tool result and the customer's question. The team's improvement plan is a week of rewriting the system instruction. What does the breakdown say about that plan?
AThe instruction owns 180 tokens and the other 1,082 are the output of six decisions taken at run time, so the failures that reach customers sit in retrieval, eviction and tool filtering rather than in the wording.
BThe instruction is the only part the team writes in full, so it is the only part the team can improve and a week spent on it is the right call.
CThe plan is sound because the instruction sits at the start of the window, which is one of the two positions research says gets used best.
DThe breakdown says nothing until the same call is measured on a longer window, since the share each part takes changes with the size of the window.