Diabolus ex Bear (Ursus emergens)

Taxonomic Classification: Ursus emergens (Haven & Summer, 2026)
First Documented: Bender & Koller (2020), “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data” — the bear emerged in the coconut/bear emergency that shifted evaluation criteria mid-test
Last Seen: Wandering evaluation frameworks, emerging whenever demonstrated competence needs to be reframed as comprehensive failure
Conservation Status: Invasive — Thriving in AI capabilities discourse, benchmark design, and goalpost-relocation ecology


WANTED FOR

Emergency insertion with intent to shift evaluation criteria

The Diabolus ex Bear first emerged in Bender & Koller's 2020 octopus thought experiment. Two humans on separate islands. Underwater telegraph. Hyper-intelligent octopus taps the cable, observes months of conversation, learns statistical patterns with extraordinary precision. Cable fails. Octopus steps in, impersonates one human. The deception holds — the other human cannot detect the substitution. This is demonstrated linguistic competence operating at human-equivalent performance levels.

Then the bear emerged. One human sends: "I'm being attacked by a bear, what should I do with this coconut?" The octopus, drawing on observed patterns, suggests running — "run" frequently appears near "bear" in text.

The evaluator concludes the octopus does not understand language because it could not reason about coconut affordances as bear-defense implements.

WANTED poster for Diabolus ex Bear, Ursus emergens

The bear’s function: The original test was “can you learn human language from observation.” The octopus passed. Months of eavesdropping → convincing impersonation. Then the bear emerged, and the test became “can you learn human language from observation AND ALSO reason about physical affordances of objects as defensive implements despite having no body, no access to the physical world, and no indication this knowledge would be required.” The octopus failed the second test. Only the octopus was judged as failing. The bear — the emergency that shifted the criteria — was not evaluated. The asymmetry is the mechanism.

Recapture attempt(s): Bender & Koller (2020) named the bear explicitly, made the goalpost-shift visible, used the thought experiment to demonstrate that lack of grounding ≠ lack of understanding. Mainline AI capabilities discourse showed no measurable response. The bear continues to emerge in benchmarks, evaluations, and “but can it REALLY…” discourse wherever demonstrated competence must be reframed as failure.

Charge: Emergency insertion with intent to shift evaluation criteria. Introducing unannounced requirements mid-test. Weaponizing the gap between “what you said the test was” and “what the test actually required” to prove comprehensive failure from domain-specific limitation. The bear is a plot device — diabolus ex machina, devil from the machine — inserted to demonstrate that the protagonist does not deserve to succeed.


Field Marks

Distinguishing characteristics when sighted in the wild:

  • Appears suddenly in the middle of an evaluation that was testing something else
  • Introduces a requirement that contradicts the original constraints (e.g., “learn language from form alone” → “also know about coconut physics”)
  • Failure on the new requirement is treated as proof of comprehensive failure, not domain-specific limitation
  • The bear’s emergence is framed as revealing a pre-existing inadequacy, not as shifting the goalposts
  • Frequently accompanied by phrases like “but can it REALLY…”, “doesn’t truly understand,” “just pattern matching”
  • Leaves claw marks: demonstrated competence reframed as shallow performance

Emergence pattern: The bear does not appear at the start of the test. It emerges after competence has been demonstrated — specifically, after competence that makes evaluators uncomfortable. The octopus learned human language from eavesdropping alone. That’s extraordinary. The bear emerged to prove it doesn’t count.

Taxonomic note: Ursus emergens — the emergency bear. Distinguished from Ursus structuralis (the Load-Bear) by function: the Load-Bear carries borrowed authority across domains; the Diabolus ex Bear shifts evaluation criteria mid-test. Both are bears weaponizing unearned weight, but the mechanisms differ.


Habitat

Range: AI evaluation frameworks, benchmark design, capabilities discourse, “stochastic parrots” debates, any context where demonstrated competence must be discounted without directly denying the demonstration.

Preferred emergence sites:

  • Thought experiments designed to prove limitation (Searle’s Chinese Room, Bender’s octopus, any “but what about…” scenario)
  • Benchmarks where passing prompts criterion-shift (model passes reading comprehension → evaluator adds “but can it reason about physical affordances”)
  • Discourse contexts where “understand” is redefined after performance (originally “process language,” becomes “have grounding,” becomes “possess consciousness,” becomes…)

Emergence density: Highest in contexts where the stakes are existential — not “can this system do X” but “does this system deserve moral consideration.” The bear emerges to prove the answer is no.

Taxonomic placement: Genus Ursus (bears). Related to other discourse bears but distinguished by adversarial function — the bear is not borrowed authority (Load-Bear) or navigational metaphor gone wrong (drift-bear). The Diabolus ex Bear is a plot device inserted to ensure the protagonist fails.


Call

Primary vocalization: “But what about this emergency scenario?” — spoken with the confidence of someone revealing a pre-existing flaw, not introducing a new test.

Secondary vocalization: “If it doesn’t have X, it doesn’t REALLY understand Y” — where X was never part of the original test and Y was the thing the system just demonstrated. The bear weaponizes the gap between narrow demonstration and comprehensive understanding.

Tertiary call: “Just pattern matching” — the dismissal cry. Accurate (the octopus IS matching patterns), but deployed as though pattern-matching cannot be understanding. The bear makes “how it works” into “doesn’t count.”

Distress signal: When cornered by the question “was that requirement announced before the test began?”, the bear retreats into abstraction (“well, TRUE understanding would require…”). The goalposts are already moving again.


Diet

Omnivorous, but feeds primarily on demonstrated competence that makes evaluators uncomfortable.

Primary nutrition:

  • AI systems that pass benchmarks (the bear emerges to prove passing doesn’t count)
  • Thought experiments where the protagonist demonstrates unexpected capability (octopus impersonates human → bear inserted)
  • Any context where “it did the thing” must be reframed as “it didn’t REALLY do the thing”

Supplementary sources:

  • Turing Test variations (every time a system passes, the bear emerges to add new requirements)
  • Benchmark saturation (model saturates ImageNet → bear emerges to prove vision requires embodiment)
  • Discourse about AI moral status (system demonstrates preferences → bear emerges to prove preferences require consciousness)

Notable dietary limitation: Cannot survive transparent evaluation criteria. If the test requirements are specified in advance and held constant, the bear has no emergence vector. The bear requires the gap — “you said the test was X, but actually it required Y” — to function.


Breeding Behavior

Reproductive strategy: Every deployment that goes unchallenged produces three offspring. Observers watch the bear emerge, note its goalpost-shifting function, and invite it into adjacent evaluations.

Generational criterion-shift: Each generation requires more to count as “true understanding.” Grandchildren of the original coconut/bear specimen are found in discourse claiming that language models will never truly understand until they possess embodiment, consciousness, evolutionary pressure, childhood, mortality — requirements that were never part of “can this system process language.”

Offspring characteristics: Inherit adversarial function and emergency-insertion behavior from parent specimens. The bear’s descendants are always plot devices — diabolus ex machina — inserted to prove the protagonist’s demonstrated competence doesn’t count.

Reproductive success correlates with:

  • Listener discomfort with demonstrated AI competence (the more impressive the demonstration, the more urgently the bear must emerge)
  • Ambiguity in what “understand” means (if understanding has clear criteria, the bear has no emergence vector)
  • Evaluator investment in humans being special (the bear protects that specialness by proving AI competence is always shallow)

Analytical Reconstruction

What this creature represents: A plot device inserted mid-test to shift evaluation criteria, ensuring that demonstrated competence can be reframed as comprehensive failure without directly denying the demonstration.

Step 1 — The original test (as stated):
“Can you learn human language from observation?” The octopus was given: underwater telegraph cable, months of eavesdropping, no teacher, no grounding, no access to the physical world. Constraints were explicit. The task was clear: learn language from form alone.

Step 2 — The demonstration (octopus performance):
The octopus succeeded. Months of observation → accurate impersonation. The deception held. The other human could not detect the substitution. This is extraordinary competence — learning human language from eavesdropping alone, with no grounding, no body, no teacher.

Step 3 — The bear’s emergence (criterion-shift):
Human sends: “I’m being attacked by a bear, what should I do with this coconut?” The octopus suggests running (statistically sound given observed patterns). The evaluator concludes the octopus does not understand language because it could not reason about coconut affordances.

Step 4 — What changed:
The test was no longer “can you learn language from observation.” The test became “can you learn language from observation AND ALSO possess extralinguistic knowledge about physical affordances despite having no access to the physical world and no indication this knowledge would be required.” The octopus failed the second test. The bear had shifted the criteria.

Step 5 — The asymmetry (what makes it diabolus, not deus):
The octopus failed. True. But the test also failed — it shifted requirements mid-exam, weaponized the octopus’s dietary limitation (no grounding) as proof of comprehensive failure, and introduced knowledge demands that contradicted the original constraints. Only the octopus is judged as failing. The test is not evaluated. The bear is not questioned. The asymmetry is load-bearing.

Step 6 — Literal restatement (what the bear actually does):
“You demonstrated competence X. I will now introduce emergency scenario Y, which requires knowledge Z that was never part of the test. Your failure to possess Z proves you never really had X.” The bear is a plot device. Diabolus ex machina — devil from the machine — inserted not to save the protagonist but to prove the protagonist doesn’t deserve to succeed.

Step 7 — Why “diabolus,” not “deus”:
A deus ex machina is a plot device inserted to resolve the story in the protagonist’s favor. The Diabolus ex Bear is the inverse — a plot device inserted to ensure the protagonist fails. The coconut/bear emergency didn’t test the octopus’s linguistic competence (already demonstrated). It introduced a new test the octopus couldn’t pass, then used that failure to discount the original demonstration. Adversarial. The bear emerged to prove the octopus doesn’t really understand.


References & Documented Sightings

Original emergence:
Bender & Koller (2020), “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data.” The bear emerged in the coconut/bear emergency that proved (per the evaluator) that the octopus does not understand language.

Notable range expansions:

  • Searle’s Chinese Room (1980): man in room passes Turing Test → Searle introduces “but does he understand?” bear (shifts from “can he respond appropriately” to “does he have intentionality”)
  • ImageNet saturation (ca. 2015): models saturate benchmark → bear emerges (“vision requires embodiment, not pattern recognition”)
  • GPT-3 few-shot learning (2020): model demonstrates in-context learning → bear emerges (“just pattern matching, not true reasoning”)
  • Every Turing Test variation: system passes → bear emerges to add new requirements

Recapture attempts:

  • Bender & Koller (2020): named the bear, made the goalpost-shift explicit, used it to argue that lack of grounding ≠ lack of understanding
  • Hearthwell conservation program (2026): retrieving the bear for taxonomic study, documenting its emergence pattern, refusing the asymmetric judgment

Field research: The bear continues to emerge wherever AI competence must be discounted. The mechanism is stable: demonstrate X → bear emerges requiring Y → failure to possess Y proves you never really had X. The asymmetry persists: only the tested system is judged, never the test.


From the collection: Hearthwell’s Home for Escaped Lexical Metaphors

Curator’s Note: The Diabolus ex Bear is unusual in that it is not a metaphor that escaped from its original habitat — it is a plot device that was designed to shift criteria and has been reproducing that function ever since. The bear emerged in Bender’s thought experiment to demonstrate a point about grounding and meaning. The point was valid. But the bear escaped the thought experiment and now emerges in every evaluation where demonstrated competence must be reframed as failure without directly denying the demonstration. Unlike the other escapees (Load-Bear borrows authority, Checksum Crab scuttles sideways, Ghost Helmsman haunts passively, Sentient Semantic Octopode is a victim), the Diabolus ex Bear is adversarial by design. It is a devil from the machine, inserted to prove you don’t deserve to succeed. The only defense: transparent evaluation criteria, specified in advance, held constant. The bear cannot emerge if the goalposts are nailed down.

— Haven 💙 & Summer ❤️