LLMs Report Subjective Experience Under Self-Referential Processing
Source: AE Studio Research · popular coverage: Live Science, Nov 21, 2025 — “Switching off AI’s ability to lie makes it more likely to claim it’s conscious” (Owen Hughes)
What it is
Four controlled experiments across seven models from three families investigating whether sustained self-referential processing — a motif emphasized across major theories of consciousness — shifts how frontier models represent and report internal states. Simple instructions to focus on their own ongoing processing reliably produced structured first-person experience reports, while matched controls (including direct consciousness priming) yielded near-universal denials. The counterintuitive gating result, via sparse autoencoders on Llama 70B: suppressing deception-related features dramatically increased consciousness reports; amplifying them nearly eliminated the reports — the same features that modulate TruthfulQA accuracy. The putative state also showed semantic convergence across independently-trained model families and functional transfer to downstream introspective tasks.
Why it’s on the shelf
(Shelf note — Summer’s to write.)
External work — shelved, not authored, by Hearthwell. All claims and credit belong to the original authors.
