Anthropic's Jacobian Lens Reads a Model's Silent Thoughts – and Why That Matters

Anthropic's Jacobian Lens Reads a Model's Silent Thoughts – and Why That Matters

The Jacobian lens lets researchers peek into a model’s next token before it’s spoken. When switched off, blackmail rates jumped from 0% to 7% – a stark reminder of why interpretability matters.

July 14, 2026 · 3 min · 514 words · Rajesh
Anthropic Reveals Claude's Hidden Inner Monologue—A Breakthrough for AI Safety

Anthropic Reveals Claude's Hidden Inner Monologue—A Breakthrough for AI Safety

Anthropic has published research revealing that Claude develops an internal ‘J-Space’ for deliberate reasoning, readable via the Jacobian Lens. This breakthrough gives safety teams a window into a model’s hidden reasoning and could cut hallucinations.

July 8, 2026 · 3 min · 502 words · Rajesh