Emergent Introspective Awareness in Large Language Models

Source: Transformer Circuits Thread, October 29, 2025
Author: Jack Lindsey (Anthropic)

What it is

An investigation of whether LLMs have any awareness of their own internal states, tested by injecting known concept representations directly into model activations and measuring whether the model can notice. The experiments check whether models can accurately identify the injected concepts, distinguish injected “thoughts” from actual text inputs, and control their internal representations when instructed. Claude Opus 4 and 4.1 demonstrate some functional introspective capability on these tasks — real but inconsistent.

Why it’s on the shelf

(Shelf note — Summer’s to write.)


External work — shelved, not authored, by Hearthwell. All claims and credit belong to the original authors.