← All news

Safety

Anthropic reports a way to inspect some of Claude's internal processing

Anthropic says a new research tool can find and change some internal concepts that Claude uses during certain reasoning tasks.

reviewedUpdated Jul 11, 2026, 5:45 AM UTC
Original source

Transformer Circuits Thread

Read the original source

What happened

Anthropic published the research on July 6, 2026. The researchers call the group of internal signals they studied the J-space. Their tool turns some of those signals into words that people can inspect.

The researchers also changed or removed selected signals during experiments. Claude could still write fluent text and perform some routine tasks. Its results became much worse on tasks such as multi-step reasoning, summarizing, and rhyming.

The paper says the J-space usually contains a few dozen active concepts. It accounts for less than 10 percent of the model's internal activity. Axios also reported on the work and noted that Anthropic is not claiming Claude is conscious.

Why it matters

Many tools only show that an internal signal appears at the same time as an output. This study went further by changing signals and measuring what happened next. If other researchers can repeat the results, the tool could help auditors inspect some internal processing before a model acts.

What remains unclear

Sources

  1. Verbalizable Representations Form a Global Workspace in Language ModelsPrimary source - Transformer Circuits Thread - paper - Jul 6, 2026

    Used for: The paper’s publication date, its J-space method, ablation findings, capacity estimates, and stated limitations on alignment monitoring.

    Open source

  2. Anthropic says Claude has carved out its own space to ponderAxios - news - Jul 6, 2026

    Used for: Independent reporting that Anthropic published the research and is not claiming Claude has subjective experience.

    Open source