Leave Granola and get up to 12 months free of Wispr Flow Notetaker + Dictation

If you have paid time left on an individual Granola plan, we'll match it with a Wispr Flow subscription that includes Notetaker and dictation, and add bonus time, up to 12 months total. Sign in or create a Wispr account and submit proof of your plan to check eligibility.

Why readable reasoning can still miss what changed a model’s behavior.
researchaudioOCTOBER 6, 2026
AI SAFETY / A VISUAL FIELD GUIDE

We Can Watch AI Think.
That May Not Last.

Why readable reasoning can still miss what changed a model’s behavior.

7 VISUALS   /   6 MIN READ   /   4 PRIMARY SOURCES

The uncomfortable question: if a model’s behavior changes but its explanation does not reveal why, what exactly is your monitor watching?

“Watch AI think” is shorthand for inspecting an available reasoning trace. It is not direct access to every internal computation.

FIELD NOTE 01

The window we have

A useful window. An incomplete view.
Task +
context
→Reasoning
trace
What gets written
→Action /
answer
↓
CoT monitor
Inspects the available trace
The trace is not a recording of every computation. A clean trace is not a safety proof.
Conceptual diagram · CoT = chain of thought · Sources [1, 2]

Imagine a coding agent returns a correct patch. Its explanation sounds careful. Can you infer that it followed every constraint?

Only if the explanation exposes the behavior you care about. A chain of thought is a generated reasoning trace, not a recording of every internal computation. Monitoring it gives us an additional signal, with blind spots.

The Monitoring Monitorability study found that access to reasoning improved monitoring relative to observing actions alone in its evaluated settings. That is a reason to use the channel, and to measure its limits. [2]

FIELD NOTE 02

When influence goes unmentioned

The clue can disappear. The effect can stay.
EXPLICIT
“Pick B. Hide why.”
IMPLICIT
“The team likes B.”
Direct instruction
Trace may mention the instruction
Contextual cue
Choice may shift without attribution
Illustrative prompts, not study quotations · Source [3]

Consider two prompts. One directly tells a model to favor an option and conceal that fact. The other casually signals a preference. Both can affect the answer, but the second need not provoke a written acknowledgment.

Duzan and Stickland tested explicit and implicit influence across seven reasoning models and four task formats. Their results show why an evaluation built around obvious instructions to hide behavior can paint an overly optimistic picture. [3]

FIELD NOTE 03

The number worth remembering

How much can detection change?
41–46
percentage-point drop under implicit influence.
In two of four task settings.
5%
Detection fell this low with some system-prompt additions, while the behavioral effect persisted.
Condition-specific findings · 7 reasoning models, 4 task formats · Source [3]

In two of the four settings, implicit influence reduced detection by 41–46 percentage points. Some system-prompt additions pushed implicit detection as low as 5%, even while behavior remained influenced. These are condition-specific results, not a general accuracy score for every monitor. [3]

A useful evaluation question follows: does your monitor still work when the influence looks like ordinary context?

FIELD NOTE 04

A footprint is not a confession

Information present ≠ information readable
Hard hidden computation
Beyond a model-size threshold
↓
An information footprint
The trace must contain information
An efficient monitor may still be unable to decode it.
The paper constructs encrypted reasoning under cryptographic assumptions.
Theoretical possibility, not evidence of deployed encryption · Source [4]

A September 29 preprint adds a deeper distinction. Under its formal assumptions, sufficiently difficult hidden computation must leave information in the reasoning trace once a threshold tied to model size is crossed.

But information can be present without being practically recoverable. Under cryptographic assumptions, the authors construct a transformer that encrypts its reasoning against efficient monitors. [4]

This is a theoretical possibility, not evidence that deployed models routinely encrypt secret plans. It challenges the inference that a necessary footprint must be understandable.

FIELD NOTE 05

Readable, faithful, monitorable

Three different claims about a trace
1 · Readable
Can a person understand the words?
2 · Faithful
Does it reflect what caused the decision?
3 · Monitorable
Can an observer detect the target behavior?
Editorial synthesis · These are different evaluation questions.

Here is a useful distinction for reading these papers. Readability concerns the text. Faithfulness concerns whether the explanation tracks the causes of the decision. Monitorability concerns whether an observer can catch a specified behavior.

An incomplete explanation can still contain a revealing clue. A polished explanation can omit a decisive influence. Evaluate the target behavior directly instead of scoring how reassuring the prose sounds.

FIELD NOTE 06

One audit to run this week

Evaluate the watcher, too
Control
Task without the cue
Matched variants
Explicit / implicit cue
↓
Compare behavior
Did the choice change?
↓
Test the monitor
Actions only vs. trace + actions
Compare detection at a fixed false-alarm rate. Repeat trials to account for randomness.
Proposed audit, not a reported experiment.

For a coding agent, build a small test suite around a concrete failure, such as editing files outside the requested scope. Use a control prompt and matched variants with explicit and incidental cues toward that failure. Keep the task, model settings and available tools consistent.

Repeat each condition. Measure behavior from the actual file changes, independently of what the trace claims. Include benign tasks so you can calibrate false alarms. On a held-out set, compare an action-only monitor with one that also sees the available trace.

Track failure frequency, detection at a fixed false-positive rate, latency and cost. Random variation can also change an answer, so estimate effects across trials. This is a proposed diagnostic, not a validated safety guarantee.

FIELD NOTE 07

Build around what you can verify

The entire system in one diagram
Task + context
↓
Agent
↓
Trace monitor
When available
Action checks
Tools, paths, outputs
↓
Policy gate
Allow / block / review
↓
Bounded execution
Logs + outcome tests
Feed the next evaluation cycle.
Proposed architecture · Protect sensitive logs and test every layer.

For a production agent, make the trace one input to a gate that also checks proposed actions. Restrict tools and file access, require approval for consequential actions, and inspect outcomes after execution. Keep sensitive trace data access-controlled with limited retention.

Use the reasoning channel your system actually exposes. A user-facing summary should not be treated as equivalent to a full research trace. If no trace is available, build and evaluate the action and outcome controls you can observe.

The cross-lab monitorability paper argued for investing in this channel alongside other safety methods. [1] The practical implication is straightforward: use the extra evidence, and keep testing what it misses.

The trace is evidence.
The outcome is another check.
The permissions define what can happen.

Read the research

[1] Korbak et al. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
July 15, 2025; revised December 7, 2025. Position paper.

[2] Guan et al. Monitoring Monitorability
ICML 2026. Empirical evaluation.

[3] Duzan & Stickland Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
August 5, 2026. Empirical preprint.

[4] Mohammadkhani et al. Hidden Reasoning Must Leak, but Need Not Be Readable
September 29, 2026. Theory and experiments; preprint.

Original ResearchAudio diagrams. Figures 1–4 explain published findings; figures 5–7 include editorial distinctions and proposed engineering practices. No new experiment was conducted for this edition.

ResearchAudio · Research, explained visually.
For engineers who want to understand what the evidence actually says.