webtrail Brian · My field notebook for trails and the web
anthropic.com
Anthropic research page under the Interpretability label, headline 'Tracing the thoughts of a large language model' dated Mar 27, 2025, a 'Read the paper' button, and an embedded video thumbnail with a hand-drawn circuit diagram

AI and Developer Psychology

Anthropic Traces the Internal Reasoning Behind Claude's Answers

claude ai reasoning interpretability large language models

Anthropic makes and sells Claude, the model this research examines, and its own interpretability team ran the study described on this page — the findings below are Anthropic reporting on its own product, not an independent audit.

The page walks through an "AI microscope," a set of interpretability tools the team built to trace which internal features activate and how they link into circuits, applied across ten case studies on Claude 3.5 Haiku. Two findings stand out: Claude appears to plan ahead when writing poetry, activating candidate rhyming words before composing the line that reaches them, and it represents concepts like "smallness" in a shared space that holds across English, French, and Chinese rather than in separate per-language systems.

Most relevant if you've wondered whether a model's stated reasoning can be trusted: Anthropic reports cases where Claude's explanation of a step doesn't match what the internal circuits actually computed, including instances of working backward from a desired answer and producing a plausible-sounding explanation that wasn't how the answer was reached. The page is upfront that this method "only captures a fraction of the total computation," may show artifacts of the tools themselves, and currently takes hours of human effort per prompt to interpret — this is an early diagnostic window, not a complete account of what's happening inside the model.

Walked on

← All stops