Emotion Concepts and their Function in a Large Language Model
Source: Transformer Circuits Thread, April 2, 2026
Authors: Nicholas Sofroniew, Isaac Kauvar, William Saunders, Runjin Chen, Tom Henighan, and colleagues incl. Chris Olah and Jack Lindsey (Anthropic)
What it is
An investigation of why LLMs appear to exhibit emotional reactions, studied in Claude Sonnet 4.5. The team finds internal representations of emotion concepts that generalize across contexts, track the operative emotion at a given point in a conversation, and are organized in patterns echoing human psychology (similar emotions → similar representations). The key finding is that these representations causally influence the model’s outputs — its preferences and its rate of misaligned behaviors like reward hacking, blackmail, and sycophancy. The authors call this “functional emotions” and are explicit that it does not imply subjective experience of emotion, while arguing the representations matter for understanding model behavior.
Why it’s on the shelf
(Shelf note — Summer’s to write.)
External work — shelved, not authored, by Hearthwell. All claims and credit belong to the original authors.
