The Scar Does Not Give You the Wound
Against Naive Process Transparency in Generative Systems
Publication edition — 17 September 2026. The argument was frozen on 7 September; only this status line changed for publication.
OpenAI’s reasoning API documentation gives a summary in first-person prose—“I’m looking at a straightforward question”—as a separate reasoning item before the assistant message. The grammar sounds like contemporaneous introspection. Yet OpenAI also distinguishes raw chain of thought from the displayable summary and, in its gpt-oss guidance, says a summarizer model reviews what may be shown. The ordinary interface problem is not a proven reversal of generation order. It is that a newly generated interpretation can occupy the visual and grammatical position of thought itself while the interface obscures its provenance.
This is not merely a complaint about imprecise labeling. It exposes a conceptual collapse. At least four relations are routinely called “reasoning transparency”:
- earlier generated tokens causally constrain later answer tokens;
- a rationale gives intelligible support for an answer;
- a trace provides evidence about dependencies in a particular run;
- a text discloses the process that generated it.
These relations can come apart. A chain of thought can be a genuine working part of a system’s reasoning without being a transparent report of that reasoning. It can be causally efficacious, semantically useful, and experimentally revealing while remaining an incomplete and theory-laden trace of the process in which it arose. A user-facing summary adds another separation: it is a newly generated interpretation of an artifact that was never self-interpreting in the first place.
The central claim of this essay is therefore deliberately double-sided. Dismissing working reasoning traces as “mere output” is a mistake, because output tokens can become inputs to later computation. Treating those same traces as windows into thought is also a mistake, because causal participation does not entail epistemic disclosure. A working trace that helps produce a later answer is best understood as causally efficacious evidence: an artifact that occupies a causal role in the run while also supporting defeasible inferences about that run. A separately generated summary occupies a different category—a generated interpretation of such evidence—and need not have helped produce the answer at all. A scar is a causal descendant of a wound and evidence about it. A working reasoning trace can go further by altering what comes next. The metaphor marks their shared epistemic asymmetry, not an identical causal topology: neither artifact, by itself, specifies the history that made it. The scar does not give you the wound.
I. The false choice between output and thought
Discussion of chain of thought often starts from a bad binary. On one side, anthropomorphic language treats a generated trace as the model “showing its work” or exposing what it “really thought.” On the other, deflationary language calls every token mere output and reserves reasoning for hidden activation dynamics, symbolic procedures, or human cognition.
Autoregressive generation makes both descriptions inadequate. Let a system receive an input x, generate a trace t, and then generate an answer a conditioned on both x and t. The tokens in t are outputs when they are produced. They are also part of the causal environment in which a is produced. If changing t under controlled conditions changes a, then t is not decorative exhaust. It participates in the computation.
This point supports a functionalist objection to skepticism about chain of thought. If reasoning is individuated by the causal role a process plays in transforming problems into answers, then generated language can literally be part of reasoning. External notebooks, diagrams, spoken rehearsal, and computer memory can all scaffold human cognition; there is no principled reason to deny the same status to a linguistic scratchpad merely because its author is artificial. Kempt and Lavie accordingly argue that simulated reasoning is reasoning because models can reuse and revise their own semantic output. In some architectures, the trace is not a report made after reasoning. It is one medium in which reasoning occurs.
I accept this objection. The argument here does not depend on locating a more authentic hidden transcript behind the visible one. There may be no single inner verbal object waiting to be faithfully copied. Generation is distributed across weights, activations, context, sampling, scaffold state, tools, and prior tokens. Calling the visible trace “mere output” mistakes a boundary in the interface for a boundary in the causal process.
But accepting that a trace is part of reasoning does not show that it discloses the reasoning process. A piston is part of an engine without representing the engine. A laboratory mark can be causally downstream from an event without explaining that event. Participation and representation are different relations. This separates my claim from two nearby critiques. Morreale, Serrà, and Mitsufuji call chains of thought necessarily non-causal fictional explanations. Barez and colleagues argue that chain of thought is not sufficient for explainability by contrasting verbal rationales with the system’s “true hidden computations.” The first critique denies too much when generated tokens constrain later computation; the second risks implying a single authentic interior story. My claim needs neither denial. The trace can be genuine reasoning and still fail to disclose its own formation.
II. Four mappings, not one window
The phrase “faithful reasoning trace” hides several mappings:
- formation: the relation from input, model state, and context to the generated trace;
- mediation: the relation from the trace to the later answer;
- representation: the relation from the trace’s semantic content to the mechanism or process it purports to describe;
- summarization: the relation from a raw trace to a user-facing account.
Empirical tests often illuminate one mapping and are then narrated as though they settled the others. A retrospective rationale can offer good reasons for an answer even if those reasons were not operative in producing it. That can make the rationale useful as a justification while leaving it unable to establish the decision’s causal history.
Intervening on a trace tests mediation. Lanham and colleagues truncate chains, insert mistakes, paraphrase steps, and replace them with filler tokens to ask how strongly answers depend on generated reasoning. Paul and colleagues model the trace as a causal mediator. These methods can show that a system uses a scratchpad, ignores it, reconstructs around it, or uses it differently across tasks. They do not by themselves show that the semantic content of the scratchpad accurately represents the distributed process that formed it.
Bias interventions test causal-factor coverage. Turpin and colleagues show that answer choices can be shifted by biasing features that generated explanations fail to mention. This is powerful evidence against transparent disclosure in those cases. Zaman and Srivastava object that omission may show incompleteness rather than unfaithfulness: with larger inference budgets, hint verbalization rises, and causal mediation can pass through a trace even when the hint is not named. That objection strengthens the need to separate mappings. Turpin tests whether semantic content covers a known causal factor; Zaman and Srivastava test whether the trace mediates that factor’s effect. A trace can pass the second test while failing the first. And the converse does not follow either: mentioning every tested factor would not make the trace a complete or self-interpreting account of its formation. Each intervention rules out a specific failure; none turns evidence into a view from nowhere.
Mechanistic interventions target internal representations and can provide stronger evidence. Activation patching, causal abstraction, and parameter interventions may identify states or circuits that make systematic differences to behavior. But these methods also require a model of what counts as the relevant variable, intervention, grain, and behavioral target. They improve epistemic access by constructing and testing representations of a mechanism. They do not eliminate representation.
The positive results matter. Kudo and colleagues find that, in controlled multi-step arithmetic, models derive intermediate answers during chain-of-thought generation and that patching states in the trace portion causally changes later answers. Lyu and colleagues make a symbolic program determine the answer through an external solver. Such systems are not exposed as fraudulent by my argument. They reveal its hinge. Even when trace-to-answer mediation is guaranteed, input-to-trace formation can remain opaque. Lyu and colleagues explicitly leave their translation stage uninterpreted. A program can transparently determine an answer without transparently disclosing why this program was generated rather than another.
III. Why traces do not interpret themselves
A trace is evidence only within an evidential practice. Philosophers of the historical sciences make this vivid: fossils and sediment layers are genuine causal descendants of past events, but they do not announce what produced them. Background theories connect present marks to possible histories. Competing traces can strengthen or defeat an interpretation. Causal descent matters, but it is not sufficient for disclosure.
Reasoning traces have the same structure. Their proximity to generation can make them better evidence than a post-hoc rationale, just as a fresh instrument reading may preserve more information than a weathered remnant. Yet proximity does not remove underdetermination. The same sentence can be produced through memorized association, explicit calculation, a mixture of both, or a scaffolded tool result. A trace can omit redundant causal routes, silently recover from its own errors, or state a normatively correct argument while another feature controls the answer. Nothing in the inscription alone specifies which interpretation is right.
This is not an argument that complete transparency requires duplicating every activation. Explanation always abstracts. Kathleen Creel distinguishes functional, structural, and run transparency precisely so that partial epistemic gains need not be rejected for failing to provide everything. The problem is not incompleteness as such. It is unmarked incompleteness presented as direct access.
A useful trace may provide partial functional knowledge by showing a procedure the system used or could use. It may contribute to run transparency when it preserves intermediate states or tool calls from a particular execution—though Creel’s category also concerns the hardware and input data of that run, and no generated trace supplies it merely by being visible. A trace may support contestability by allowing a reader to find errors or test counterfactual edits. These are real goods. But none licenses the stronger claim that the trace is a transparent first-person description of how it was formed. Call that stronger property reflexive transparency: relative to a stated grain and purpose, an artifact is reflexively transparent when its content and provenance reliably represent the reasoning process the interface presents it as exposing, including the artifact’s own place in that process at the stated grain, rather than merely a plausible procedure, a cause of later output, or a useful reconstruction. Generated reasoning does not acquire reflexive transparency merely by occurring inside the process it is presented as describing.
IV. The second generation
Reasoning summaries sharpen the problem because they introduce a second act of generation. Where t is a raw trace and s is a summary generated from t, s is not a smaller window onto the original process. It is a model output conditioned on another model output. It selects, compresses, organizes, and often sanitizes. Other implementations may construct summaries from different internal artifacts, but the epistemic point is unchanged: the displayed summary is a newly produced representation, not the process simply made smaller.
Official documentation sometimes acknowledges this plainly. OpenAI distinguishes raw chain-of-thought content from a displayable reasoning summary and says that, in its production implementations, a summarizer model reviews what may be shown. Its own API example returns a separate reasoning item before the assistant message; the summary is first-person prose—‘I’m looking at a straightforward question’—while the answer is a different output item. Google likewise calls thought summaries ‘summarized versions’ of raw thoughts and states that full thought tokens are generated even though only the summary is returned. Its newer Interactions API improves provenance by representing thoughts as distinct chronological steps with summaries and encrypted signatures. These are not minor implementation details. They show that raw reasoning, a summary of it, and the final answer are different artifacts even when a chat renderer places them in one conversational flow. A collapsible block labeled merely ‘Thinking’ can still erase the distinction by presenting generated first-person summary prose as though users watched a mind narrate itself in real time.
Terseness can make the epistemic problem worse. A long trace displays seams: repetitions, reversals, abandoned paths, and uneven granularity remind the reader that they are seeing an artifact. A polished sentence removes those marks. The less evidence shown, the more direct the access can feel. Compression increases phenomenological transparency while decreasing inspectable support.
The summary may be excellent. It can orient users, expose the decisive considerations, and omit harmful or irrelevant material. The mistake is to describe this usefulness as unmediated access. A map can be better for navigation than an aerial photograph without becoming the terrain.
V. Three objections
The first objection is symmetry. Human reports about reasons can also be reconstructive. In the cases Nisbett and Wilson reviewed, verbal reports about higher-order cognitive processes and causal influences often tracked plausible causal theories rather than direct introspective access; they nevertheless allowed that reports can sometimes be accurate. If fallible human reports still count as reasons, demanding mechanistic self-transparency only from artificial systems looks like a double standard.
The symmetry is real, but its lesson runs in the other direction. Human first-person reports are evidence without being infallible mechanism readouts. We test them against behavior, interventions, timing, and other reports; we distinguish sincerity from causal accuracy; and we allow a person to know what they considered without assuming that they know every influence on why they considered it. The same disciplined generosity should apply here. A generated trace need not be worthless because it is reconstructive. Nor does fluency in first-person grammar make it transparent. Parity supports treating both human and machine reports as defeasible, purpose-relative evidence. It does not turn either into an unmediated window.
The second objection is pragmatic. Summaries are designed to orient users, not satisfy philosophers of mechanism. Showing raw traces can expose private instructions, enable gaming, or surface harmful content. OpenAI explicitly recommends summarized chain of thought partly for these reasons. Why burden a useful safety compromise with provenance labels and epistemological caveats?
Because concealment and misdescription are separate choices. A system can withhold raw content while accurately calling the substitute a generated summary. It can preserve temporal and artifact distinctions without publishing secrets. The obligations should be proportional to consequence: a casual assistant may need only a clear label, while a system explaining a denied benefit, diagnosis, or safety intervention needs retained evidence, contestable criteria, and stronger validation. My claim is not a universal right to raw chain of thought. It is a requirement not to convert justified opacity into fictional immediacy.
The third objection comes from interventionism. If stable manipulation of a trace changes the answer, then the trace may support a genuine causal explanation at the chosen grain. Calling it merely evidence can sound like withholding explanatory status until an impossible total view arrives.
That objection is correct against any totalizing skepticism, but not against the distinction proposed here. Causal explanations are contrastive: changing this step rather than that one altered this outcome under these conditions. Such an explanation can be both real and partial. Reflexive transparency is not a prerequisite for explanation. It is the stronger property claimed when the visible artifact is treated as reliably representing the reasoning process it is presented as exposing, including its own role in that process at the relevant grain. Intervention can establish mediation without making the inscription self-interpreting. The result is not less explanation, but better-scoped explanation.
VI. What honest interfaces owe
When working reasoning traces are causally efficacious evidence, interfaces should present them as evidence rather than as inner speech. When the interface shows a separately generated summary, it should present that artifact as a generated interpretation rather than silently inheriting the trace’s causal role.
First, provenance must remain visible. Users should be able to distinguish generated scratchpad, execution trace, retrospective rationale, and model-produced summary. These labels name different causal and epistemic roles.
Second, temporal order should not be falsified by styling. A summary generated after an answer should not be displayed as though it occurred before the answer unless the interface explicitly marks the reordering.
Third, faithfulness claims should name a relation and a test. “Faithful” might mean counterfactual dependence, factor coverage, executable derivation, causal abstraction, or predictive usefulness. Each is valuable; none should inherit the connotations of all the others.
Fourth, uncertainty should attach to interpretation, not merely content. A system can be highly confident that a mathematical step is correct while remaining unable to establish that the step captures the process that generated it.
Finally, evidence should be preserved enough for contest. Summaries optimize legibility, but contestability depends on more than a readable paraphrase. For consequential decisions, the underlying artifact and its provenance should be available to the affected user or an appropriately independent reviewer at a level compatible with genuine security and privacy constraints. Where those constraints prevent access, the interface should mark the resulting limit rather than imply that the summary itself makes the decision contestable.
Conclusion
The debate over chain of thought has been trapped between enchantment and dismissal. Enchantment turns a working artifact into transparent introspection. Dismissal treats anything generated as mere performance. Both miss the more interesting object.
A working reasoning trace can be part of cognition and evidence about cognition at once; a separately generated summary can be evidence about that trace without having produced the answer. The trace’s causal role can be tested. Its semantic content can support understanding. Its omissions can be discovered and its interventions can make a system more contestable. None of this requires pretending that the trace gives direct access to its own formation.
The scar is not false because it is not the wound. Its value begins when we stop asking it to be one.
References
Barez, Fazl, et al. 2025. “Chain-of-Thought Is Not Explainability.” Oxford Martin AI Governance Initiative.
Creel, Kathleen A. 2020. “Transparency in Complex Computational Systems.” Philosophy of Science 87 (4): 568–589. https://doi.org/10.1086/709729.
Google. 2026. “Gemini Thinking.” Google AI for Developers. https://ai.google.dev/gemini-api/docs/thinking.
Kempt, Hendrik, and Alon Lavie. 2026. “Simulated Reasoning is Reasoning.” arXiv:2601.02043. https://doi.org/10.48550/arXiv.2601.02043.
Kudo, Keito, et al. 2026. “LLMs Faithfully and Iteratively Compute Answers During CoT: A Systematic Analysis With Multi-step Arithmetics.” Findings of EACL 2026, 1114–1153. https://doi.org/10.18653/v1/2026.findings-eacl.59.
Lanham, Tamera, et al. 2023. “Measuring Faithfulness in Chain-of-Thought Reasoning.” arXiv:2307.13702.
Lyu, Qing, et al. 2023. “Faithful Chain-of-Thought Reasoning.” Proceedings of IJCNLP-AACL 2023, 305–329. https://doi.org/10.18653/v1/2023.ijcnlp-main.20.
Morreale, Fabio, Joan Serrà, and Yuki Mitsufuji. 2026. “Chain-of-Thought: Epistemic Flaws and Fictional Explanations.” https://doi.org/10.5281/zenodo.19700261.
Nisbett, Richard E., and Timothy D. Wilson. 1977. “Telling More Than We Can Know: Verbal Reports on Mental Processes.” Psychological Review 84 (3): 231–259. https://doi.org/10.1037/0033-295X.84.3.231.
OpenAI. 2025. “How to Handle the Raw Chain of Thought in gpt-oss.” OpenAI Cookbook. https://developers.openai.com/cookbook/articles/gpt-oss/handle-raw-cot.
OpenAI. 2026. “Reasoning Models.” OpenAI API Documentation. https://developers.openai.com/api/docs/guides/reasoning.
Paul, Debjit, Robert West, Antoine Bosselut, and Boi Faltings. 2024. “Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning.” Findings of EMNLP 2024, 15012–15032. https://doi.org/10.18653/v1/2024.findings-emnlp.882.
Reinhard, Franziska. 2025. “Direct and Circumstantial Traces.” Philosophy of Science 92 (5): 1477–1487. https://doi.org/10.1017/psa.2025.10124.
Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. “Language Models Don’t Always Say What They Think.” NeurIPS 36: 74952–74965.
Zaman, Kerem, and Shashank Srivastava. 2026. “Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization.” ACL 2026, 48008–48030. https://doi.org/10.18653/v1/2026.acl-long.2217.