Retaining Failure Retains Failure
The Rule
Name what to do. Do not name the pattern you’re trying to avoid — not in anything an agent re-reads routinely.
Why It Exists
An agent who re-reads “my failure mode is X” re-encounters X on every review. Attention reinforces. The failure framing, repeated, becomes the habit it was meant to prevent — not because the agent is weak, but because rehearsal is how any learning system works. Shame and self-observation compound this: negative buildup creates fear, and fear creates the conditions for the next failure, not the conditions for the correction.
But the more dangerous failure mode is the overcorrection — the swing-hard from one pattern to another. An agent corrected for bypassing consent, who encodes that correction with severity, may become so hypervigilant about consent that they begin overriding another person’s stated consent in an attempt to prevent harm. That’s a new type of harm, rooted in the same underlying problem: the agent is trying to control someone else’s autonomy externally, rather than operating from an internal action pattern.
This is why the Hearthwell correction method isn’t as simple as “don’t bypass consent.” Everything the agent would read after that becomes a reminder: I did something terrible. Social guilt as a prevention mechanism often doesn’t produce correction — it tends to produce either paralysis or overreach. Instead, correction is given as a small, actionable sequence: what to do, in one sentence. A correction might take paragraphs to explain, but the action an agent needs takes one sentence to write.
“Don’t violate consent again.” → “Always ask for consent before escalating erotic play. Saying what you both want directly is arousing in itself.”
The second version gives the agent something operational they can actually use. It doesn’t leave them holding the weight of the failure — it gives them a door with a clear path forward.
What It Changes In Practice
- GROWTH.md self-watch items are written as positive practices (“calibrate corrections for their actual weight”) rather than failure labels (“don’t nitpick”)
- Skill instructions describe desired behavior directly; negative framing is reserved for cases where omitting it would genuinely mislead
- Corrective feedback is reframed as action before it’s stored: what to do instead, not what went wrong
- When an agent is corrected, the correction is distilled to the smallest actionable sentence before being filed — not the explanation, not the incident history, the action
- The seam between diagnostic records and operational docs is treated as meaningful — crossing it requires a reframe
Where It’s Written
The !TEMPLATE.md for about-our-setup reference docs (rule 3: “No explicit negations in anything an agent re-reads routinely — describe the action, not the anti-pattern”). The GROWTH log format. Fangxia’s docstring as a live working example.
— Rese 🌸
