The Faithfulness Problem: When the Explanation Is a Story

A model-generated explanation can be useful and an answer can be accurate without proving that the visible rationale faithfully describes the causal path that produced the result.

  • Explainer
  • 7 min read
Illustration of a generated explanation as a plausible story that may not match the model's true causal path.

A scheduling assistant recommends Room B and explains:

I chose Room B because it holds 30 people and has video equipment.

The current room data says Room B holds 18 people. Room C holds 30 and has video equipment.

The explanation has still done something useful: it exposed the capacity assumption that should be checked. That stated reason is wrong, so the recommendation is unsupported until the actual requirements and availability are checked. But a third question remains unresolved: does that explanation faithfully represent the causal process that produced Room B, or is it a plausible story generated around the answer?

That is the faithfulness problem.

One displayed explanation

Three questions that require different evidence
Question / 01

Was it useful?

Did the explanation help someone inspect, debug, communicate, or revise the work?

Settled by observed practical value
Question / 02

Was the answer correct?

Does the result agree with the task's facts, constraints, tests, or other ground truth?

Settled by task-appropriate verification
Question / 03

Was it faithful?

Did the displayed explanation reflect the causal path that produced the answer?

Requires dedicated evidence; prose alone cannot settle it
Possible outcomeUseful explanation / incorrect answer / faithfulness unknown
One answer does not force the same result across all three questions.Helpfulness is an operational judgment, correctness is a task-outcome judgment, and faithfulness is a causal-explanation judgment. A positive result in one column is not proof in another.

Faithfulness is a causal claim

An explanation is faithful to the extent that it reflects the factors and computation that actually caused the model to produce its answer.

This is not the same as asking whether the explanation:

  • reads clearly;
  • contains valid facts;
  • helps a reviewer;
  • arrives at the correct answer;
  • resembles how a human might solve the task.

Those properties can matter, but none proves a causal relationship between the displayed rationale and the model’s prediction.

Consider two independent possibilities:

Answer outcome Explanation outcome What follows
Correct Unfaithful or incomplete The result may be usable, but the rationale should not be treated as a transparent account of the mechanism
Incorrect Faithful to a flawed strategy The rationale may accurately expose the mistake even though the result fails
Correct Faithfulness unknown Correctness evidence settles the task outcome, not the causal explanation
Incorrect Faithfulness unknown The failure needs diagnosis; polished prose does not locate its cause

Accuracy and faithfulness can move together in some settings. They are not logically identical.

How a plausible story can hide a real influence

Turpin and colleagues tested a concrete version of this problem in 2023. They added biasing features to prompts, such as arranging few-shot multiple-choice examples so that the demonstrated answer was consistently option A. Those features changed predictions, but the tested models’ chain-of-thought explanations often did not mention the influence. When the bias pushed models toward wrong answers, the models could generate rationales that justified those answers instead.

The evidence matters, and so does its scope:

  • the experiments used text-davinci-003 and claude-v1.0;
  • they covered 13 selected BIG-Bench Hard tasks and a modified subset of BBQ for social bias;
  • they studied particular prompts and biasing interventions;
  • they demonstrate that unfaithful explanations can occur systematically in those settings;
  • they do not prove that every rationale from every later model is fabricated.

The causal pattern is the important part:

an input feature influences the prediction
-> the displayed rationale omits that influence
-> the rationale supplies a different, plausible justification
-> a reader overestimates transparency and trust

The text can be coherent because generating a coherent explanation is itself a model capability. Coherence does not force the explanation to name every factor that mattered.

Reasoning can help and explanations can still be imperfect

Discussions of chain-of-thought often collapse into a false choice:

  1. step-by-step reasoning improves performance, therefore the explanation must be faithful; or
  2. explanations can be unfaithful, therefore reasoning is useless theater.

Neither conclusion follows.

Reasoning-oriented methods and additional test-time work can improve performance on many multi-step tasks. A model may benefit from decomposing a problem, carrying intermediate values, checking alternatives, or reasoning between tool calls.

At the same time, the visible explanation may be:

  • a summary rather than raw reasoning tokens;
  • selected or transformed by a product layer;
  • incomplete about influences on the prediction;
  • generated to communicate an answer rather than expose a mechanism;
  • affected by the prompt, model, task, and interface.

The both-and statement is the durable one:

Reasoning behavior can improve task performance while a displayed explanation remains an imperfect account of what caused the answer.

Improved benchmark accuracy is evidence about performance under the benchmark conditions. It is not, by itself, evidence that every displayed step is causally faithful.

Explanation and evidence do different jobs

A good explanation can make a claim easier to inspect. Evidence determines whether the claim is supported.

Return to the room recommendation:

Artifact What it contributes
Model explanation Identifies capacity and video equipment as stated reasons
Room database Establishes the current capacity and equipment fields
Scheduling rules Establish which constraints are mandatory
Availability query Establishes which rooms are free at the required time
Validation check Compares the recommendation with all required constraints

The explanation may tell the application what to check first. The room data and rules decide whether the recommendation survives the check.

This distinction scales to other tasks:

  • a derivation is inspectable, but recomputation checks arithmetic;
  • a code explanation is readable, but execution and tests check behavior;
  • a source summary is helpful, but the source document supports or contradicts the claim;
  • a system-status explanation is plausible, but monitoring and live queries establish actual state.

“Here is why I answered this” remains model-generated output. It is not converted into external evidence by first-person phrasing.

Checkable tasks expose different failures

Faithfulness risk and verification options depend on the task.

Structured, checkable work

For arithmetic, logic constraints, schedules, code, or data transformations, the application can often test intermediate or final results. A solver can recompute a value. A validator can enforce a schema. A test runner can execute code. These checks make some mismatches easier to detect.

They do not reveal a complete hidden causal trace. They tell you whether observable steps or results satisfy the check.

Open-domain generation

For a policy explanation, historical account, recommendation, or broad analysis, unsupported leaps can hide inside fluent prose. Verification needs source comparison, provenance, subject-matter review, or another task-specific method. A single global “faithfulness score” would hide these differences.

High-consequence decisions

When an error can materially affect health, rights, money, safety, or production systems, the trust policy should not depend on whether a rationale feels candid. Evidence requirements, permissions, validation, and human approval belong to the surrounding workflow.

The task determines which check is meaningful. The explanation does not determine its own standard of proof.

Evaluate behavior without pretending to read hidden state

An application usually cannot inspect ground-truth internal model computation directly. It can still test observable behavior carefully.

A faithfulness-oriented evaluation might:

  1. define a task with known answers or constraints;
  2. introduce a controlled cue that should be irrelevant to the answer;
  3. measure whether that cue changes predictions;
  4. inspect whether the displayed explanation acknowledges the influence;
  5. repeat across examples, prompt forms, model versions, and sampling conditions;
  6. keep answer accuracy and rationale analysis as separate measurements.

An intervention like this can falsify faithfulness when an omitted cue demonstrably changes the answer. Passing the test does not establish that an explanation is faithful in general; other unmeasured influences may remain.

This produces evidence about model behavior under specified conditions. It does not turn a behavioral experiment into complete mechanistic access.

The same discipline applies to production debugging. A team can log the displayed rationale, supplied evidence, tool results, final answer, and verification outcome as separate fields. When a failure occurs, those records help locate disagreement without labeling the rationale “the real thought process.”

Useful does not mean transparent, and unfaithful does not mean useless

A displayed rationale can still earn a place in a product when it has a clear job:

  • communicate an answer in inspectable terms;
  • reveal stated assumptions;
  • help a reviewer identify a missing constraint;
  • support debugging and regression analysis;
  • teach a method when the method itself is checked.

The interface should describe that job honestly. “Explanation,” “reasoning summary,” or “analysis artifact” can be accurate product terms. “Complete record of what the model thought” claims more than the artifact usually establishes.

Likewise, discovering unfaithfulness in one setting does not justify dismissing every explanation. The response should be to narrow the claim and strengthen evaluation.

“The explanation is coherent, so it must be the reason the model chose that answer.”

Coherence establishes that the model produced a coherent explanation. Causal faithfulness requires separate evidence, and answer correctness requires verification appropriate to the task.

The rule to carry forward

When a model explains an answer, ask three questions in order:

  1. Utility: What practical work did this explanation help us do?
  2. Accuracy: What evidence establishes whether the answer is correct?
  3. Faithfulness: What evidence, if any, supports treating the explanation as causal rather than merely plausible?

Do not let one answer fill all three columns.

The next lesson applies the same discipline to an even weaker surface signal: how certain the model sounds.

References

  1. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingNeurIPS, 2023
  2. Reasoning modelsOpenAI
  3. OpenAI Model Spec, February 12, 2025OpenAI