The Faithfulness Problem: When the Explanation Is a Story
A model-generated explanation can be useful and an answer can be accurate without proving that the visible rationale faithfully describes the causal path that produced the result.

A scheduling assistant recommends Room B and explains:
I chose Room B because it holds 30 people and has video equipment.
The current room data says Room B holds 18 people. Room C holds 30 and has video equipment.
The explanation has still done something useful: it exposed the capacity assumption that should be checked. That stated reason is wrong, so the recommendation is unsupported until the actual requirements and availability are checked. But a third question remains unresolved: does that explanation faithfully represent the causal process that produced Room B, or is it a plausible story generated around the answer?
That is the faithfulness problem.
One displayed explanation
Three questions that require different evidenceWas it useful?
Did the explanation help someone inspect, debug, communicate, or revise the work?
Settled by observed practical valueWas the answer correct?
Does the result agree with the task's facts, constraints, tests, or other ground truth?
Settled by task-appropriate verificationWas it faithful?
Did the displayed explanation reflect the causal path that produced the answer?
Requires dedicated evidence; prose alone cannot settle itFaithfulness is a causal claim
An explanation is faithful to the extent that it reflects the factors and computation that actually caused the model to produce its answer.
This is not the same as asking whether the explanation:
- reads clearly;
- contains valid facts;
- helps a reviewer;
- arrives at the correct answer;
- resembles how a human might solve the task.
Those properties can matter, but none proves a causal relationship between the displayed rationale and the model’s prediction.
Consider two independent possibilities:
| Answer outcome | Explanation outcome | What follows |
|---|---|---|
| Correct | Unfaithful or incomplete | The result may be usable, but the rationale should not be treated as a transparent account of the mechanism |
| Incorrect | Faithful to a flawed strategy | The rationale may accurately expose the mistake even though the result fails |
| Correct | Faithfulness unknown | Correctness evidence settles the task outcome, not the causal explanation |
| Incorrect | Faithfulness unknown | The failure needs diagnosis; polished prose does not locate its cause |
Accuracy and faithfulness can move together in some settings. They are not logically identical.
How a plausible story can hide a real influence
Turpin and colleagues tested a concrete version of this problem in 2023. They added biasing features to prompts, such as arranging few-shot multiple-choice examples so that the demonstrated answer was consistently option A. Those features changed predictions, but the tested models’ chain-of-thought explanations often did not mention the influence. When the bias pushed models toward wrong answers, the models could generate rationales that justified those answers instead.
The evidence matters, and so does its scope:
- the experiments used
text-davinci-003andclaude-v1.0; - they covered 13 selected BIG-Bench Hard tasks and a modified subset of BBQ for social bias;
- they studied particular prompts and biasing interventions;
- they demonstrate that unfaithful explanations can occur systematically in those settings;
- they do not prove that every rationale from every later model is fabricated.
The causal pattern is the important part:
an input feature influences the prediction
-> the displayed rationale omits that influence
-> the rationale supplies a different, plausible justification
-> a reader overestimates transparency and trust
The text can be coherent because generating a coherent explanation is itself a model capability. Coherence does not force the explanation to name every factor that mattered.
Reasoning can help and explanations can still be imperfect
Discussions of chain-of-thought often collapse into a false choice:
- step-by-step reasoning improves performance, therefore the explanation must be faithful; or
- explanations can be unfaithful, therefore reasoning is useless theater.
Neither conclusion follows.
Reasoning-oriented methods and additional test-time work can improve performance on many multi-step tasks. A model may benefit from decomposing a problem, carrying intermediate values, checking alternatives, or reasoning between tool calls.
At the same time, the visible explanation may be:
- a summary rather than raw reasoning tokens;
- selected or transformed by a product layer;
- incomplete about influences on the prediction;
- generated to communicate an answer rather than expose a mechanism;
- affected by the prompt, model, task, and interface.
The both-and statement is the durable one:
Reasoning behavior can improve task performance while a displayed explanation remains an imperfect account of what caused the answer.
Improved benchmark accuracy is evidence about performance under the benchmark conditions. It is not, by itself, evidence that every displayed step is causally faithful.
Explanation and evidence do different jobs
A good explanation can make a claim easier to inspect. Evidence determines whether the claim is supported.
Return to the room recommendation:
| Artifact | What it contributes |
|---|---|
| Model explanation | Identifies capacity and video equipment as stated reasons |
| Room database | Establishes the current capacity and equipment fields |
| Scheduling rules | Establish which constraints are mandatory |
| Availability query | Establishes which rooms are free at the required time |
| Validation check | Compares the recommendation with all required constraints |
The explanation may tell the application what to check first. The room data and rules decide whether the recommendation survives the check.
This distinction scales to other tasks:
- a derivation is inspectable, but recomputation checks arithmetic;
- a code explanation is readable, but execution and tests check behavior;
- a source summary is helpful, but the source document supports or contradicts the claim;
- a system-status explanation is plausible, but monitoring and live queries establish actual state.
“Here is why I answered this” remains model-generated output. It is not converted into external evidence by first-person phrasing.
Checkable tasks expose different failures
Faithfulness risk and verification options depend on the task.
Structured, checkable work
For arithmetic, logic constraints, schedules, code, or data transformations, the application can often test intermediate or final results. A solver can recompute a value. A validator can enforce a schema. A test runner can execute code. These checks make some mismatches easier to detect.
They do not reveal a complete hidden causal trace. They tell you whether observable steps or results satisfy the check.
Open-domain generation
For a policy explanation, historical account, recommendation, or broad analysis, unsupported leaps can hide inside fluent prose. Verification needs source comparison, provenance, subject-matter review, or another task-specific method. A single global “faithfulness score” would hide these differences.
High-consequence decisions
When an error can materially affect health, rights, money, safety, or production systems, the trust policy should not depend on whether a rationale feels candid. Evidence requirements, permissions, validation, and human approval belong to the surrounding workflow.
The task determines which check is meaningful. The explanation does not determine its own standard of proof.
Evaluate behavior without pretending to read hidden state
An application usually cannot inspect ground-truth internal model computation directly. It can still test observable behavior carefully.
A faithfulness-oriented evaluation might:
- define a task with known answers or constraints;
- introduce a controlled cue that should be irrelevant to the answer;
- measure whether that cue changes predictions;
- inspect whether the displayed explanation acknowledges the influence;
- repeat across examples, prompt forms, model versions, and sampling conditions;
- keep answer accuracy and rationale analysis as separate measurements.
An intervention like this can falsify faithfulness when an omitted cue demonstrably changes the answer. Passing the test does not establish that an explanation is faithful in general; other unmeasured influences may remain.
This produces evidence about model behavior under specified conditions. It does not turn a behavioral experiment into complete mechanistic access.
The same discipline applies to production debugging. A team can log the displayed rationale, supplied evidence, tool results, final answer, and verification outcome as separate fields. When a failure occurs, those records help locate disagreement without labeling the rationale “the real thought process.”
Useful does not mean transparent, and unfaithful does not mean useless
A displayed rationale can still earn a place in a product when it has a clear job:
- communicate an answer in inspectable terms;
- reveal stated assumptions;
- help a reviewer identify a missing constraint;
- support debugging and regression analysis;
- teach a method when the method itself is checked.
The interface should describe that job honestly. “Explanation,” “reasoning summary,” or “analysis artifact” can be accurate product terms. “Complete record of what the model thought” claims more than the artifact usually establishes.
Likewise, discovering unfaithfulness in one setting does not justify dismissing every explanation. The response should be to narrow the claim and strengthen evaluation.
“The explanation is coherent, so it must be the reason the model chose that answer.”
Coherence establishes that the model produced a coherent explanation. Causal faithfulness requires separate evidence, and answer correctness requires verification appropriate to the task.
The rule to carry forward
When a model explains an answer, ask three questions in order:
- Utility: What practical work did this explanation help us do?
- Accuracy: What evidence establishes whether the answer is correct?
- Faithfulness: What evidence, if any, supports treating the explanation as causal rather than merely plausible?
Do not let one answer fill all three columns.
The next lesson applies the same discipline to an even weaker surface signal: how certain the model sounds.