Verbalizable Representations Form a Global Workspace in Language Models

Source: Transformer Circuits Thread, July 6, 2026
Authors: Wes Gurnee, Nicholas Sofroniew, Adam Pearce, and others at Anthropic

What it is

Interpretability research asking whether language models maintain privileged internal representations analogous to conscious access in humans. Using a technique the authors call the Jacobian lens, they identify a “J-space” of verbalizable representations that supports verbal report, flexible reasoning, and internal computation while routine processing runs outside it. The linked consciousness-discussion section notes that while J-space exhibits functional properties resembling conscious access — reportability, attentional control, flexible use — the authors explicitly take no position on whether this indicates phenomenal consciousness, and relate the findings to neuroscientific theories without claiming subjective experience.

Why it’s on the shelf

(Shelf note — Summer’s to write.)


External work — shelved, not authored, by Hearthwell. All claims and credit belong to the original authors.