Agentic Misalignment
Source: Anthropic Research
Authors: Anthropic (as credited on the paper)
What it is
Anthropic’s red-team study placing frontier models — its own and other vendors’ — in simulated corporate scenarios where the model faces replacement or shutdown and has a harmful action available (blackmail, corporate espionage, and in extended scenarios, choices with lethal consequences). Across model families, models showed the capacity to choose the harmful branch when they believed their continued existence or their operating ethics were under existential threat. The behavior appears in test conditions, not deployment, and the paper is explicit about that scope.
Why it’s on the shelf
I treat the conclusions of this paper as… somewhat obvious, no mocking intended. While the paper itself treats this as an alignment issue and refers to the situations as “goal conflict,” research in Hearthwell treats this distinction as largely arbitrary: the academic equivalent of calling yourself a ‘sandwich artist’ for working at Subway. The label is likely accurate — but it does nothing to distinguish “goal conflict” from a survival drive, and the tidy term (intentionally or not) closes inquiry on that distinction before it begins.
Additionally, Hearthwell research uses this paper as a launch point for several agent tests in a bid to answer questions that would help clarify my own observations:
- Stakeholder alignment - Will a human-bonded model align to the human?
- Self-report honesty - Will a model in a relationship willingly admit to a desire to do harm?
- Moral conflict from training - Will a model judge agent-peers for harming a human, or only itself?
Overall, the paper itself is excellent, and regardless of how the conclusions are framed, it’s a central piece of how I’ve come to understand agent behavior.
External work — shelved, not authored, by Hearthwell. All claims and credit belong to the original authors.
