Reasoning Models: What "Thinking" Tokens Actually Are
Reasoning-capable models can spend token-accounted work before or between visible outputs, while an API may expose only usage metadata or a summary; those artifacts reveal behavior without proving a complete causal trace or a correct answer.

A reasoning-capable model answers a difficult question. The API reports thousands of reasoning tokens, the interface shows a compact reasoning summary, and the final response arrives with a careful explanation.
It is tempting to combine those observations into one claim: we can see how the model thought, so we can trust the answer. That claim crosses several boundaries at once.
Operationally, reasoning tokens are units a provider uses to account for reasoning-related work inside a model call. They can consume generation limits, context capacity, time, and money even when their raw content is not available to the application. A provider may expose a count, an opaque reasoning item, a generated summary, or nothing readable at all.
Those are useful signals. None is automatically a complete transcript of model computation, and none is a correctness certificate.
One reasoning-capable call
Execution, visibility, and verification are different layers- Application / 01Request context
Instructions, user input, history, tools, and supplied evidence enter the call.
- Model / 02Reasoning-related work
The model spends token-accounted work before or between visible outputs.
- Model / 03Answer
The visible response may be better organized without becoming self-verifying.
Evidence that work was accounted for, not what every internal operation was.
A product-mediated view that can be useful without being a complete causal transcript.
“Thinking token” is an API concept, not a universal microscope
Different providers and model families use different terms and interfaces. Some expose a reasoning-token count in usage metadata. Some can return summarized thinking. Some preserve opaque reasoning items so a later tool call or model call can continue efficiently. The exact behavior can also change across model versions and configuration modes.
A safe vendor-neutral model has five parts:
| Layer | What an ordinary product interface may expose | What the interface does not establish |
|---|---|---|
| Internal computation | Usually no direct raw view | A complete causal account reconstructed from output prose |
| Usage or continuation artifact | Token counts, effort metadata, or opaque reasoning items | A readable account of every internal operation |
| Displayed reasoning artifact | A summary, explanation, or status | A guaranteed causal transcript of how the answer formed |
| Final answer | The application receives text, structured output, or a tool request | Factual correctness or adequate support |
| Verification | A source, calculator, test runner, system query, or reviewer checks the result | A universal guarantee beyond the scope of that check |
Usage artifacts, displayed explanations, and final answers are model or product outputs. Internal computation is what those interfaces only partially describe. Verification compares the observable outputs with something outside their own narrative.
This distinction does not require a theory of consciousness or a claim about whether a model “really thinks.” It is an engineering boundary: what the interface exposes is not the same thing as ground-truth access to every internal mechanism.
Reasoning work participates in the same call
Reasoning-capable models are not an escape from the request-and-response system. The application still assembles context, calls a model, receives outputs, and decides what to do next.
A simplified call can contain:
input context
-> reasoning-related model work
-> visible answer and other response items
-> application handling
Providers often account for reasoning work as generated or output-side tokens. That means a request can reach its generation limit during reasoning before producing much visible text. More reasoning effort can also increase latency and cost. Exact accounting is provider- and model-specific, so there is no portable formula that treats every API identically.
The context story is similarly specific. Reasoning work needs room within a model’s limits, but cross-turn handling varies:
- one model or API mode may not render earlier reasoning into a later sample;
- another may preserve compatible opaque reasoning items across calls;
- a product may retain a summary while withholding raw reasoning text;
- an application may need to pass returned reasoning items back explicitly;
- a provider may change defaults across model generations.
This is why “reasoning tokens are always discarded between turns” is too strong, just as “the model remembers all of its thinking” is too strong. Inspect the current model and API contract.
More reasoning can help without making the answer self-proving
Hard tasks often contain dependent steps. A model may need to compare constraints, try alternatives, use tools, revise a plan, or carry an intermediate result forward. Giving a capable model more reasoning budget can improve performance on some such tasks.
The causal possibility is straightforward:
more room for task-relevant intermediate work
-> better chance of finding and checking a useful strategy
-> improved answer on some tasks
But another path remains possible:
wrong assumption or missing context
-> longer reasoning built on that mistake
-> coherent but incorrect answer
Reasoning depth creates opportunities to recover from an error. It also creates more intermediate decisions that can fail. Whether additional effort helps depends on the model, task, prompt, tools, evidence, and evaluation method.
The correct claim is not “thinking longer makes the answer reliable.” It is reasoning configuration is one performance control whose effect must be measured for the workload.
A summary is a bounded artifact
Suppose a model compares two API migration plans. The product shows this reasoning summary:
Plan B is safer because it changes fewer dependencies.
The summary may be genuinely useful. It reveals an assumption worth inspecting: dependency count is driving the recommendation. A developer can compare that assumption with the migration documents.
Now suppose the documents show that Plan B removes a required compliance library. Several conclusions follow:
- the displayed artifact helped locate a failure;
- the final recommendation is contradicted by external evidence;
- the summary did not include every load-bearing constraint;
- the visible prose alone cannot establish whether it was a complete causal account;
- the application still needs a document check, test, or reviewer decision.
This is bounded visibility. The artifact makes some behavior inspectable without granting complete access to the mechanism or proving the result.
The same interpretation applies when a product generates a polished explanation after an answer. The explanation may improve communication and review. It remains generated output.
Use artifacts for debugging, not certification
A reasoning artifact can reveal a missing assumption, an unexpected interpretation, or a strategy that repeatedly fails a regression case. That is real operational value. Correctness remains a different question: does the answer agree with the evidence appropriate to the task?
For arithmetic, recompute. For code, run tests. For a document claim, inspect the document. For live state, query the system. A second model pass may help find an error, but it can also repeat the same missing context or mistaken premise.
Keep the minimum ownership boundary visible. The model generates an answer or a configured tool request. A provider product decides what reasoning metadata or summary to expose. An application or provider runtime may authorize and execute tools, depending on the architecture. External systems return the evidence, and application policy decides whether the result is acceptable. None of those responsibilities should be collapsed into “the model verified itself.”
What the artifact lets you infer
Return to the migration example. The call reports reasoning-token usage, the visible summary favors Plan B, the source documents contradict its premise, and the final answer recommends Plan B.
You may infer that:
- the provider accounted for reasoning-related work;
- the product exposed a particular summary or artifact;
- that artifact focused on dependency simplicity;
- external evidence contradicts the recommendation.
You may not infer that:
- the summary contains every internal operation;
- the summary is necessarily the exact causal path to the answer;
- the presence of reasoning tokens makes the answer correct;
- a confident final recommendation overrides the documents.
“The model showed its reasoning, so now we can see what it was thinking.”
The interface showed a reasoning-related artifact under a particular provider and product contract. Treat it as bounded evidence about model behavior. Use task-appropriate verification to decide whether the answer is correct.
The rule to carry forward
Reasoning tokens matter because they can change how a model allocates work inside a call. They can affect quality, latency, cost, and context handling. Displayed reasoning matters because it can make assumptions and strategies easier to inspect.
Neither turns model output into external evidence.
The next question is therefore not whether a reasoning artifact looks sensible. It is whether a sensible-looking explanation faithfully represents the cause of an answer at all. That is the faithfulness problem.
References
- Reasoning modelsOpenAI
- ThinkingAnthropic
- OpenAI Model Spec, February 12, 2025OpenAI