Next-Token Prediction

A language model builds a response one token at a time, repeatedly scoring candidates against the context created so far.

  • Explainer
  • 4 min read
Illustration of a language model predicting the next token from a sequence of candidates.

When a language model answers a prompt, it does not retrieve a finished sentence and reveal it from left to right. It repeatedly makes a narrower decision: given the tokens available now, what could come next?

That decision is enough to generate paragraphs, programs, and plans. It also explains why two responses can diverge after an early difference, and why polished language is not evidence that the result was checked against the world.

One decision at a time

Assume the prompt has already been tokenized. The model receives that sequence as its current context and produces a score for candidate tokens in its vocabulary. Those scores induce a distribution: some candidates fit the context more strongly than others.

A decoding strategy then selects one candidate. The selected token is appended to the context. The model evaluates the expanded sequence, producing a new distribution for the next position. Generation repeats until a stopping condition is reached.

The autoregressive loop.Candidate weights are illustrative, not factual confidence. Each selected token changes the context used for the next step.

The loop is small enough to keep in your head:

  1. Score possible next tokens from the current context.
  2. Select one token according to the decoding strategy.
  3. Append that token to the context.
  4. Repeat with a different input because the context has changed.

A distribution, not a stored sentence

Consider this incomplete incident diagnosis:

The deployment failed because the service

Several continuations might fit. Candidates corresponding to timed, was, and could can all receive weight. The model supplies the candidate scores; the decoding policy turns those scores into one selected token.

If timed is selected, the context now ends with service timed, making out a plausible next continuation. If was is selected instead, candidates such as unavailable or misconfigured may become more likely.

The important part is not which branch wins in this illustration. It is that the next distribution belongs to the branch that was actually taken. The probabilities from the original prompt are not reused unchanged for the rest of the response.

Selection is a separate policy

The model’s scores do not dictate one universal selection rule. Greedy decoding can choose the highest-scoring candidate. Sampling can choose according to the distribution. Other strategies constrain or reshape the candidate set in different ways.

This distinction matters because generation behavior can change even when the underlying model does not. A decoding strategy affects variation and repetition; it is not a truth-checking stage.

Why early branches matter

Each emitted token becomes part of the next input. A small difference near the beginning can therefore cascade through every later step.

Imagine two runs beginning with the same instruction:

Return a concise incident diagnosis:

One response starts with The database connection.... Another starts with The authentication service.... From that first difference onward, each run is conditioned on a different token sequence. Later divergence is expected even if both responses remain grammatically smooth.

This is autoregressive generation: earlier output feeds back into the process that generates later output.

Fluency is not verification

Next-token prediction is optimized to continue context plausibly. External factual verification is a different operation.

Nothing in the baseline score-select-append loop checks deployment logs, queries a database, compares a claim with a trusted source, or establishes that a five-step plan is globally coherent. Applications can add tools, retrieval, and verification around generation, but those mechanisms should remain visible as separate layers.

This is why a contradiction late in a response is not necessarily a transport failure or evidence that a hidden answer was corrupted. The later text was generated under the context created by all earlier selections.

The mental model to keep

When a response surprises you, replay the loop:

  • What context was available at this step?
  • Which earlier tokens shaped the current branch?
  • What selection policy could introduce variation?
  • What external evidence, if any, was actually consulted?

The loop does not explain every capability of a modern model. It does provide the mechanical core on which context windows, sampling, tool use, agent loops, and many failure modes depend.

References

  1. Language Models are Few-Shot LearnersNeurIPS
  2. The Curious Case of Neural Text DegenerationICLR
  3. Generation strategiesHugging Face
  4. SelfCheckGPTEMNLP