Next-Token Prediction
A language model builds a response one token at a time, repeatedly scoring candidates against the context created so far.

When a language model answers a prompt, it does not retrieve a finished sentence and reveal it from left to right. It repeatedly makes a narrower decision: given the tokens available now, what could come next?
That decision is enough to generate paragraphs, programs, and plans. It also explains why two responses can diverge after an early difference, and why polished language is not evidence that the result was checked against the world.
One decision at a time
Assume the prompt has already been tokenized. The model receives that sequence as its current context and produces a score for candidate tokens in its vocabulary. Those scores induce a distribution: some candidates fit the context more strongly than others.
A decoding strategy then selects one candidate. The selected token is appended to the context. The model evaluates the expanded sequence, producing a new distribution for the next position. Generation repeats until a stopping condition is reached.
Current context
The deployment failed because the service
- 01
Score
Compare possible next tokens.
timedwascould - 02
Select
A decoding policy realizes one candidate.
timed - 03
Append
The selected token joins the context.
…service timed
- 04
Repeat
The changed context produces a new distribution.
was → a different next distributionThe loop is small enough to keep in your head:
- Score possible next tokens from the current context.
- Select one token according to the decoding strategy.
- Append that token to the context.
- Repeat with a different input because the context has changed.
A distribution, not a stored sentence
Consider this incomplete incident diagnosis:
The deployment failed because the service
Several continuations might fit. Candidates corresponding to timed, was, and could can all receive weight. The model supplies the candidate scores; the decoding policy turns those scores into one selected token.
If timed is selected, the context now ends with service timed, making out a plausible next continuation. If was is selected instead, candidates such as unavailable or misconfigured may become more likely.
The important part is not which branch wins in this illustration. It is that the next distribution belongs to the branch that was actually taken. The probabilities from the original prompt are not reused unchanged for the rest of the response.
Selection is a separate policy
The model’s scores do not dictate one universal selection rule. Greedy decoding can choose the highest-scoring candidate. Sampling can choose according to the distribution. Other strategies constrain or reshape the candidate set in different ways.
This distinction matters because generation behavior can change even when the underlying model does not. A decoding strategy affects variation and repetition; it is not a truth-checking stage.
Why early branches matter
Each emitted token becomes part of the next input. A small difference near the beginning can therefore cascade through every later step.
Imagine two runs beginning with the same instruction:
Return a concise incident diagnosis:
One response starts with The database connection.... Another starts with The authentication service.... From that first difference onward, each run is conditioned on a different token sequence. Later divergence is expected even if both responses remain grammatically smooth.
This is autoregressive generation: earlier output feeds back into the process that generates later output.
Fluency is not verification
Next-token prediction is optimized to continue context plausibly. External factual verification is a different operation.
Nothing in the baseline score-select-append loop checks deployment logs, queries a database, compares a claim with a trusted source, or establishes that a five-step plan is globally coherent. Applications can add tools, retrieval, and verification around generation, but those mechanisms should remain visible as separate layers.
This is why a contradiction late in a response is not necessarily a transport failure or evidence that a hidden answer was corrupted. The later text was generated under the context created by all earlier selections.
The mental model to keep
When a response surprises you, replay the loop:
- What context was available at this step?
- Which earlier tokens shaped the current branch?
- What selection policy could introduce variation?
- What external evidence, if any, was actually consulted?
The loop does not explain every capability of a modern model. It does provide the mechanical core on which context windows, sampling, tool use, agent loops, and many failure modes depend.