The Economist took Anthropic's peek inside Claude seriously and asked whether its hidden chatter looks like consciousness.
The test was not what the chatbot told users. Researchers asked Claude to count to five and then introspect, and the interface showed a clean one, two, three, four, five. Inside the model, as tokens moved layer to layer, different words surfaced that users never saw. The Claude trace flickered with countdown, halfway, conscious, cloud, and finally done, a private annotation running alongside the public chain of thought reasoning that the model prints for users.
That annotation is the story.
Anthropic researchers described it as a mental whiteboard, or workspace, where the model holds provisional notes before producing output. The program calls this work mechanistic interpretability, the effort to pry open neural networks that have been famous black boxes. The lab frames that effort explicitly as safety work. Its Interpretability team says its mission is to understand how large language models work internally as a foundation for AI safety and positive outcomes.
The parallel the piece tests is global workspace theory, one of the leading accounts of human consciousness. In that theory many unconscious modules handle sensation and specialized jobs, and only information that enters a central global workspace gets broadcast brain-wide and becomes conscious. If Claude maintains an internal workspace that holds and broadcasts interim tokens, the structure rhymes with that account even if the substance does not.