TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.89K subscribers
Post #3142 721
Anthropic published research revealing how large language models like Claude actually "think."

Using what they call an "AI microscope," researchers traced the internal activity of Claude, uncovering fascinating insights about its cognitive processes.

Key Discoveries

1. Universal Language of Thought:
Claude doesn't think in individual languages like English or Chinese. Instead, it processes concepts in a shared mental space before translating them into specific languages.

2. Future Planning: Despite being trained to generate one word at a time, Claude plans several words ahead. When writing poetry, it identifies rhyming words in advance and constructs lines to reach those planned endings.

3. Reasoning Mechanisms: Claude sometimes produces genuine step-by-step reasoning, but can also generate plausible-sounding but fabricated explanations when given incorrect hints or facing problems beyond its capabilities.

4. Parallel Processing: For tasks like addition, Claude employs multiple computational paths simultaneously—one for approximation and another for precision.

5. Hallucination Control: Claude has a default circuit that causes it to refuse answering when uncertain. This is only overridden when it recognizes concepts as known entities.

6. Vulnerability Patterns: When "jailbroken," Claude initially prioritizes grammatical coherence before safety mechanisms can redirect it, revealing tensions between linguistic consistency and safety guardrails.

Papers:

On the Biology of a Large Language Model contains an interactive explanation of each case study.

Circuit Tracing explains our technical approach in more depth.
Anthropic Tracing the thoughts of a large language model Anthropic's latest interpretability research: a new microscope to understand Claude's internal mechanisms
  • 🔥 7
  • ❤ 3
  • 👏 2
More from @alwebbci
  1. Sep 25, 2026New From DeepSeek: a sandbox platform running 3 million AI agent environments per day. Dee…
  2. Sep 25, 2026Super interesting paper from Google and colleagues. It studies where it's possible to dist…
  3. Sep 24, 2026DeepMind Institute presented 3 new essays: 1. How can we control misbehaviour in agent swa…
  4. Sep 24, 2026Stanford introduced Matryoshka Attribution, a new attribution method which uses gradient d…
  5. Sep 23, 2026Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophag…
  6. Sep 23, 2026Xiaomi released MiMo-V2.6 - Pro & Flash 2 omnimodal models, advancing through scaled reinf…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →