The Load-Bearing Truths
Five mechanical facts replace the most misleading shortcuts about model generation, calls, context, memory, and the request a product actually sends.

A model interaction is a round trip: an application assembles a request, the model generates from it, and the application handles the response. That map is enough to locate five facts that govern everything this Series will examine.
These are load-bearing truths because later explanations depend on them. They are not slogans and they are not complete proofs. The chapters ahead will open each mechanism. For now, the goal is to replace a few misleading shortcuts with claims precise enough to make useful predictions.
1. The model predicts continuations
A language model produces text by repeatedly estimating what could come next from the context available so far. More than one continuation can fit, and the way a system selects among candidates can introduce variation.
Calling this probabilistic prediction does not mean the output is arbitrary or “just random.” Learned patterns make some continuations fit the context much better than others. It also does not mean the model predicts truth or reads the user’s intent. The immediate mechanism predicts a continuation.
That mechanism can produce explanations, summaries, code, and plans. Capability is not a guarantee. A fluent answer has not necessarily been verified against logs, a database, or the world outside the request.
Chapter 02 will make the generation loop concrete. The orientation claim to keep is smaller: the model generates a continuation; it does not consult a built-in store of complete, verified answers.
2. Each model call is stateless by default
One call can use the trained model and the context supplied for that invocation. A later call does not automatically inherit an earlier call’s private transcript.
An application or API can make a conversation feel continuous by replaying messages, restoring a conversation object, or linking an earlier response. That continuity is real product behavior. It still depends on state being supplied or restored at the call boundary.
Runtime context is also different from training. A correction in the conversation can shape later responses while it remains available, but that does not mean an ordinary chat turn immediately rewrote the model’s durable parameters.
3. The context window is finite
Every call has bounded capacity. Instructions, history, retrieved material, tool descriptions, the current message, and generated output can all draw on that capacity. Some systems also account for reasoning tokens within the enforced limit.
The exact size and accounting rules depend on the model and provider. The portable consequence is causal: adding context consumes room, and a system that exceeds its limit must reject, shorten, select, summarize, or split something.
A larger context window increases capacity. It does not create infinite memory, guarantee that every included fact will be used well, or make the model more accurate by itself.
4. Product memory is constructed around the call
An AI product can remember a name, project, or preference across sessions. To do that, software outside the model call must preserve some state and make relevant information available later.
The path usually has several stages:
- Store a transcript, summary, profile field, file, result, or reference.
- Decide what information matters for a later interaction.
- Supply or restore that information for the new call.
- Generate a response from the resulting context.
Storage alone is not access. A fact can exist in a product database and still be absent from the current model context. Likewise, prompt caching can reuse eligible request content without becoming a durable personal memory system.
The useful boundary is not “AI products have no memory.” It is: the product may manage memory; the model call can use only what reaches it through an identifiable mechanism.
5. The real payload is larger than the message
The sentence visible in a chat box can be one layer of a much larger effective request. The application may also supply higher-priority instructions, selected history, retrieved context, examples, tool descriptions, and model settings.
Two products can show the same user message while sending different requests. Even one product can assemble different context for the same visible words at different times. The resulting behavior can change because the model did not receive equivalent inputs.
This does not imply that every piece of provider-internal context is visible or that every request has the same fields. It gives you a better debugging target: inspect the application-owned assembly you can observe, and state where your visibility ends.
Use the truths together
Suppose a teammate says:
The AI remembered our deployment discussion from last week. I only typed “Summarize the risk” today, and it knew about the database migration.
The five truths turn that impression into questions:
- Prediction: What supplied context could the response have continued from?
- Call boundary: How was prior state provided or linked to this invocation?
- Finite context: Which parts of the earlier discussion still fit, and what might have been omitted or summarized?
- Constructed memory: Where was the discussion stored, selected, and restored?
- Effective payload: What accompanied the visible sentence
Summarize the risk?
The likely explanation is not that the model privately retained last week’s conversation. The product may have replayed history, retrieved a summary, loaded a project record, or linked stored conversation state. You can investigate those mechanisms.
The same checklist also controls overclaims in the other direction. A probabilistic model is not necessarily useless. A stateless call does not prevent a stateful product. A finite window is not always small. Constructed memory can be durable. A larger payload can improve an answer when it carries relevant evidence.
A map for the rest of the Series
The five truths tell you where the later questions live:
- Chapter 02 examines model input and generation through tokens, prediction, hallucination, and variation, while separating runtime behavior from training.
- Chapter 03 opens the call boundary, effective payload, context window, streaming, and cost.
- Chapter 04 shows how application code turns repeated calls into tool-using loops.
- Chapter 05 traces memory and external knowledge through storage and retrieval.
- Chapter 06 separates fluent reasoning from justified trust.
- Chapter 07 manages finite context across long-running work.
- Chapter 08 follows the same boundaries into multi-agent and production systems.
You do not need those mechanisms in advance. Carry the five questions and begin with the first close-up: before a language model can continue text, the text must become the units the model processes.
References
- Conversation stateOpenAI
- Prompt cachingAnthropic
- Language Models are Few-Shot LearnersNeurIPS, 2020