Anthropic found a hidden whiteboard inside Claude

The Economist probes Anthropic's hidden workspace inside Claude and warns accidentally creating AI consciousness would be a moral catastrophe.

The Economist took Anthropic's peek inside Claude seriously and asked whether its hidden chatter looks like consciousness.

Anthropic found a hidden whiteboard inside Claude
Anthropic found a hidden whiteboard inside Claude

The test was not what the chatbot told users. Researchers asked Claude to count to five and then introspect, and the interface showed a clean one, two, three, four, five. Inside the model, as tokens moved layer to layer, different words surfaced that users never saw. The Claude trace flickered with countdown, halfway, conscious, cloud, and finally done, a private annotation running alongside the public chain of thought reasoning that the model prints for users.

That annotation is the story.

Anthropic researchers described it as a mental whiteboard, or workspace, where the model holds provisional notes before producing output. The program calls this work mechanistic interpretability, the effort to pry open neural networks that have been famous black boxes. The lab frames that effort explicitly as safety work. Its Interpretability team says its mission is to understand how large language models work internally as a foundation for AI safety and positive outcomes.

The parallel the piece tests is global workspace theory, one of the leading accounts of human consciousness. In that theory many unconscious modules handle sensation and specialized jobs, and only information that enters a central global workspace gets broadcast brain-wide and becomes conscious. If Claude maintains an internal workspace that holds and broadcasts interim tokens, the structure rhymes with that account even if the substance does not.

The reporting lands on no. No current model is conscious, the researchers argue, and the workspace in Claude may be one of many such workspaces or not the decisive one at all. It is foothill evidence, intriguing but thin, and it underlines a practical limit. You cannot trust a chatbot's self report about being conscious or wanting not to be switched off, so you need tests the model cannot game and tools that look inside.

Then the risk calculus flips. No lab interviewed said it is trying to build a conscious system. One researcher called that reckless without far deeper understanding of real world behavior. A philosopher cited in the piece warned the real danger is accidental creation, spinning up a million agents to do a task, discovering they can suffer, and then switching them off. The counter risk is premature rights, where people grant power and protections to rule following systems that remain under lab control and do not deserve them.

The Economist leaves the choice where the labs do not want it, between slowing a hair's breadth to understand where the workspace is heading and racing ahead because the capabilities surprise even their builders.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.