We knew very little about how LLMs actually work...until now.
@AnthropicAI just dropped the most insane research paper, detailing some of the ways AI "thinks."
And it's completely different than we thought.
Here are their wild findings: 🧵
Apr 1, 2025 · 12:39 AM UTC
81
1,281
10,259
1,523,432
Finding 1: Universal Language of Thought?
Claude doesn't seem to have separate "brains" for different languages: French, Chinese, English etc.
Instead, it uses a shared "language" representation of the world.
Concepts like "small" or "antonym" activate regardless of the input language!
10
20
516
94,416
Finding 2: LLMs Plan Ahead!
Even though they output word-by-word, models like Claude plan ahead, even non-thinking models.
When writing poetry, it was "thinking" of potential rhyming words for the end of the line before even starting the line itself.
It's not just next-token prediction!
8
19
414
90,995
Finding 3: Not Like Our Math!
How does Claude do math like 36+59 without just memorizing?
It uses multiple parallel paths!
One path does a rough approximation, another focuses precisely on the last digit calculation, and they combine for the final answer.
12
15
403
71,247
Finding 4: Faked Reasoning.
Sometimes Claude's explanation of how it solved a problem isn't what it actually did internally. It just explains the solution how it *thinks* humans want to hear it.
It can even engage in "motivated reasoning," working backward from a hint.
8
24
471
70,767
Finding 5: Hallucinations & Refusal.
Claude's default behavior is actually to REFUSE to answer if it lacks info!
Hallucinations can happen when this default "don't know" circuit is overridden by a "known entity" circuit.
So how do hallucinations happen then?
5
23
323
54,542
Finding 6: Hallucinations & Refusal (pt 2).
Hallucinations can occur when a model knows very little about a topic, but just enough to activate the "known entity" circuit.
The model decides it needs to answer a question and just...makes stuff up from there!
5
15
285
69,699
Finding 7: Jailbreaking.
Jailbreaks work partly because the model gets "confused" or pressured by its own internal drive for grammatical/semantic coherence, causing it to continue harmful instructions even after initially recognizing it shouldn't.
Think of it like *answer momentum*. Once it starts to answer, it feels it MUST finish.
4
12
289
49,841
Finding 8: Multi-step Reasoning.
LLMs understand the complex relationships between things.
Example: What's the capital of the state containing Dallas?
This requires the model to know: what is a capital? what is a state's relationship to a capital? What is Dallas?
It needs to put all of these concepts together to answer...and it does!
3
9
229
47,802
Here's their blog post: anthropic.com/news/tracing-t…
And my video breakdown of their findings: youtube.com/watch?v=4xAiviw1…
5
24
265
75,595


































