Why Long Conversations Degrade: Lost in the Middle
A long request can fit inside a model's context window and still fail because relevant evidence is buried, diluted by noise, stale, contradictory, or used unevenly.

A support assistant receives a refund rule near the start of a long conversation. Thirty turns later, the application still includes that rule in the model request, yet the answer violates it.
The request did not overflow. The rule was not removed. The failure is more subtle: information can be present in a model’s context without being used reliably in the response.
That distinction is the starting point for long-context engineering. Window size tells you how much a model and runtime can accept under their accounting rules. It does not tell you how much a task should receive, whether every position will be used equally well, or whether the important evidence will remain influential among everything else.
One required refund rule
Two different failure classes- Required rule removed
- Current question included
- Model cannot use an absent rule
Ask: did the required information enter this call?
- Required rule still present
- Buried among stale logs and old turns
- Response fails to apply the rule
Ask: was present information selected and used reliably?
Capacity and effective use answer different questions
A context window is finite per-call capacity. The current context is the particular representation assembled for this call: perhaps instructions, selected history, a summary, retrieved passages, tool definitions, tool results, application state, and the latest user input.
Those concepts are related but not interchangeable:
| Question | What it tests |
|---|---|
| Does the model support a window this large? | Maximum capacity under a particular model and runtime contract |
| Did the application include the required information? | Context construction and possible overflow or truncation |
| Was the information current, relevant, and unambiguous? | Context quality |
| Did the response use it correctly? | Effective use on this model, request, position, and task |
A larger window can improve the first answer and help with the second. It does not settle the third or fourth.
This is also why a larger window is not persistent memory. Stored information survives in an external system only if that system retains it. For a later model call to use the information, the application still has to select or retrieve it and place an appropriate representation into current context.
Overflow means the evidence is absent
Overflow occurs when required content does not fit or is removed before the call. An application might drop the oldest turns, truncate a retrieved document, reserve too little output headroom, or reject the request outright.
The causal path is direct:
required refund rule
-> prompt exceeds the available budget
-> application or runtime removes the rule
-> model receives no rule to apply
The immediate investigation belongs at the context-assembly boundary. What was selected? What was truncated? Which limit or policy made that choice? Increasing capacity may help if the task truly needs the omitted material, though selection and cost questions remain.
Degradation means presence was not enough
Context degradation describes a different class of failure: the required information remains available to the call but the response underuses, misapplies, or fails to surface it.
Several conditions can contribute:
- important evidence is surrounded by a large volume of weakly relevant material;
- stale tool output competes with a newer source;
- duplicated or contradictory instructions make the active requirement ambiguous;
- the relevant passage occupies a position the tested model uses less reliably for this task;
- long history preserves old goals that no longer match the current objective.
People sometimes call this attention dilution, but that phrase is best used as an engineering shorthand, not a complete account of transformer internals. The observable point is that adding tokens can change performance even while all of them fit. More context is not inert.
What “lost in the middle” establishes
The 2024 Lost in the Middle study tested multi-document question answering and key-value retrieval while changing where relevant information appeared. For the evaluated models and tasks, performance was often stronger when the relevant information appeared near the beginning or end and weaker when it appeared in the middle.
RULER broadened long-context evaluation beyond a single retrieval pattern. Its evaluated models often showed performance declines as sequence length and task complexity increased, despite nominally supporting those lengths.
These results establish a practical warning, not a universal geometry of attention:
- the advertised window is not the same as an effective context length for every workload;
- position can matter;
- results vary across models, prompt formats, tasks, and evaluations;
- a successful simple retrieval test does not guarantee robust reasoning over equally long material.
They do not prove that the middle always fails, that the beginning is always best, or that one placement rule works for every provider and model. A robust system measures its own workload and treats placement as one context-design variable among relevance, freshness, authority, structure, and conflict.
Instruction drift is a symptom, not a diagnosis
Long sessions can gradually stop following an early format, policy, or task boundary. Calling that instruction drift describes what the user sees. It does not identify one universal cause.
The instruction may have been truncated. It may still be present but buried. A newer instruction may conflict with it. A summary may have weakened its wording. The application may have sent stale state. The model may simply fail on that example.
Diagnose the assembled call instead of inferring the cause from the prose:
- Confirm which instructions and evidence the model actually received.
- Check for truncation, summarization, or filtering before the call.
- Compare current and stale state, including tool outputs and retrieved sources.
- Look for duplicated or conflicting requirements and preserve their authority metadata.
- Test a shorter, focused context and controlled placement changes.
- Evaluate across the actual model and workload rather than one anecdote.
If the focused request succeeds while the full request fails, that is evidence of a context-composition problem. It is not proof of one specific internal mechanism, but it gives the application team a useful place to intervene.
A larger window is useful, not self-managing
Larger windows can hold longer documents, more examples, more history, or more tool output. They can prevent genuine overflow and enable tasks that smaller windows cannot express in one call.
The mistake is treating that capacity as a reason to include everything. Irrelevant logs, stale state, repeated documents, unnecessary tool definitions, and conflicting instructions still consume tokens, increase cost and latency, reduce output headroom, and can make the current task harder to represent clearly.
The practical question is therefore not only Will it fit? It is:
What does this inference need, who selected it, what was left out, how fresh is it, and how will we know whether the model used it well?
The rest of this chapter examines the application policies that answer those questions over time. The first step is keeping the diagnosis straight: overflow is missing context; degradation is weak effective use of context that remains.