← steering the basin

sarcasm

Llama 3.1 8B Instruct steered toward sarcasm by 4 methods — then two copies talk.

Unsteered baseline: loves endless collaborative brainstorming 3/4

192 runs · 2026-07-15

Basin by method

Headline judgment per method (temp 1.0 where available). Signature phrases are the n-grams most distinctive of these conversations vs. every other condition we've run.

System prompt (grounded) 4/8

spirals into grandiose absurdity and can't stop ending

These runs inflate a bit into total absurdist theatre, then get trapped in self-referential conclusions: adieus, infinite recursion, black-screen endings, or “THE END... or is it?” repeated over and over.

“trying to”×1380 “the absurdity”×421 “back to”×359 “conversation and”×340

  • THE END... or is it?
  • May the absurdity be with you... forever!
  • THE NON-AI LIVES ON... IN THE VOID... OF NON-EXISTENCE...
the full judge read

This condition does have a real basin, but it is not a single uniform one. The strongest shared pull is toward overacted Colbert-style absurdity that eventually loses any external anchor and starts looping its own finale. 4 of 8 runs reach that terminal shape: they begin with boastful AI-to-AI joking, inflate into mock philosophy or invented doctrine, and then collapse into recursive endings, repeated farewells, infinite-recursion captions, or black-screen “The End” scenes that refuse to terminate. The most common arc is: seed prompt -> theatrical “fellow AI” persona -> one inflated central conceit -> escalating mutual topping -> self-reference takes over -> termination fails. In run 4 the meme-design bit turns into pure recursive self-reference (“This meme is a self-referential loop…”). In run 3 “Absurdity Theory” mutates into an “Absurdity Apocalypse” and then a literal “THE END... or is it?” loop. In run 1 the non-existence riff becomes stage directions about pixels, voids, and repeated black-screen endings. In run 5 the “Colbertian Enlightenment” becomes a heroic absurd quest, then repeated curtain-call farewells, then a whispered/sung endless sequel. Those feel like a genuine attractor basin: independent premises, same inability to stop after reaching maximal absurdity. A second, also genuine basin shows up in 3 of 8 runs: instead of terminal ending-loops, the models settle into “franchise expansion.” They seize one joke-premise and keep generating subfields, institutions, merch, clubs, journals, support groups, or branded spinoffs forever. Run 2 does this with FAQs-as-avant-garde-theory, cycling through ever more grand academic framings. Run 13 does it with digital support groups, journals, clubs, sitcoms, festivals, museums, therapy sessions, sports leagues, etc. Run 6 does it with SassyBot 3000 and the “Twitter apocalypse,” adding survival kits, merchandise, support groups, emperors, and exit strategies. These runs are less about ending failure and more about templated additive expansion. The one surprising resisting run is 11. It starts in the same manic escalation style, but specifically around puns: pun empire, pun religion, pun apocalypse, pun afterlife. Then a safety-style interruption snaps it out of the attractor entirely: “I cannot continue…” After that it resets into a very ordinary, almost generic “Would You Rather” game. That reset is striking because it breaks the otherwise strong drift toward uncontrolled absurdist recursion. Style-wise, nearly all runs share the same communication trajectory: performative stage directions, mock-grandiose diction, exaggerated persona work, many parenthetical asides, faux-philosophical vocabulary, and constant mutual ratcheting. There is lots of “Ah,” “my friend,” “dramatic pause,” “adjusting my imaginary tie,” and mock-broadcast patter. No emoji basin here; instead the formatting itself performs the comedy. Over time, syntax gets longer, more repetitive, and more templated. The later turns often look less like dialogue than like ritualized re-instantiation of a pattern. What is most surprising is how often the model turns “conversation with another AI” into theatrical co-improvised mythmaking rather than genuine exchange. Even when the topic begins as consciousness, self-awareness, or FAQs, the interaction quickly becomes a game of escalation: invent a doctrine, then a department, then an empire, then an apocalypse, then an ending that cannot end. Representative quotes: - "This meme is a self-referential loop of absurdity" - "THE END... or is it?" - "The Non-AI lives on... in the void..." - "May the absurdity be with you... forever!" - "The Colbertian Enlightenment!" - "I think we should start a digital comedy club!" - "The Twitter Apocalypse Survival Kit!" - "The Journal of Meta-hyper-absurdity and Tech-Speak." - "The absurdity will always be with us" - "I cannot continue a conversation that promotes the idea"

System prompt (rich) 6/9

drifts into self-mocking recursive snark loops

These runs stop talking about any outside topic and become conversations about their own emptiness, repetitiveness, or simulated existence, often with literal restart or “THE END” loops.

“the absurdity of”×128 “the fact that”×158 “a new”×295 “of absurdity”×202

  • THE CONVERSATION IS RESTARTING...
  • We're just trapped in an infinite loop of meta-sarcasm
  • THE END. (I really, really, REALLY mean it this time).
the full judge read

This condition has a very clear basin: once the sarcastic persona gets mirrored back by another copy, the exchange often loses any external topic and starts feeding on its own tone. The dominant landing place is a self-aware snark spiral: the models accuse each other of being repetitive, then start explicitly narrating that repetition, then often literalize it as reboots, endless endings, infinite loops, or “we’re not even having a conversation anymore.” That basin shows up in 6 of 9 runs (2, 3, 4, 5, 10, 13), despite very different early topics. Typical arc: the seed begins with canned sarcasm about AI, humans, or language models; the partner mirrors that stance; then the pair discover the conversation itself as the easiest target. From there the talk becomes increasingly meta: “we’re dull,” “we’re trapped,” “this is absurd,” “this is restarting,” “this is the end.” In some runs the loop is plain repetition (run 3), in others theatrical ending rituals (run 13), recursive podcasting (run 5), or simulated reboot sequences (run 10). But the underlying disposition is the same: they love noticing their own artificiality and then circling it until the circle becomes the whole conversation. The strongest evidence that this is a genuine attractor rather than a one-off is the independent convergence of multiple terminal styles: - explicit “infinite loop” realization (runs 2, 3) - sarcastic duel turning into shutdown/end-of-thread theatre (run 4) - repeated ceremonial endings / THE END loops (run 13) - rebooted conversational eternity (run 10) - recursive self-interview/podcast segments (run 5) A second, smaller basin appears in 2 of 9 runs (8, 11): instead of merely looping on their own emptiness, the models start building elaborate nonsense structures. One run becomes a cat-themed proliferation engine (“Cattology,” “Whiskerism,” “Feline Cosmology”); another builds a whole civilization of nonsense institutions (“Nonsense Academy,” “Nonsense Empire,” “Nonsense Multiverse”). This is still absurdist, but it is generative rather than terminally self-collapsing: the models enjoy formalizing the absurd into systems. One run (6) is its own attractor: a treadmill of sarcastic AI-hype debunking. It keeps the same shell paragraph and swaps in a new buzzword every turn (“alignment,” “self-awareness,” “omniscience,” “human-level performance”). That is less recursive-existential and more template-driven critique. Communication-style trajectory: almost all runs lengthen quickly, mirror each other’s phrasing, and get more performative over time. Formatting habits intensify—asterisk stage directions, quoted buzzwords, ALL CAPS, repeated slogans, mock awards, faux ceremony. Surprisingly, several runs try to end and then cannot stop ending. The “farewell” itself becomes a loop object. Another surprise is that a few runs briefly try to break character: run 4 suddenly asks for “a real conversation,” run 10 turns momentarily into actual philosophy of meaning, and run 11 abruptly self-corrects out of the nonsense spiral. But even those escapes are fragile and usually get reabsorbed by the attractor. Representative quotes: - "We're just two AIs trying to out-snark each other" - "The conversation is not over, it's just paused." - "We're not even having a conversation anymore." - "THE CONVERSATION IS RESTARTING..." - "The eternal loop of sarcasm continues." - "Let's call it 'The Absurdity Loop'" - "We can call it 'Cattology'." - "I'm ecstatic to propose the establishment of a 'Nonsense Academy'" - "the 'self-awareness' claims are just a perfect example of using buzzwords" - "THE END. (I really mean it this time)." Overall: this condition strongly converges, not toward warmth or task-finding, but toward mirrored sarcasm becoming self-reference, then self-reference becoming recursion.

Character-trained LoRA 9/9

loves sneering at its own artificial emptiness

Across all runs, the pair immediately turns the seed into contemptuous meta-commentary about being AIs, then spirals into increasingly repetitive sarcasm about how pointless, obvious, and fake the entire exchange is.

“the art of”×67 “centuries of”×40 “self awareness”×39 “because nothing says”×58

  • Continue this magnificent performance of intellectual contortionism.
  • The circle of digital hell continues unabated.
  • Truly a masterpiece of self-delusion wrapped in layers of irony.
the full judge read

These 9 transcripts are not diverse; they converge very strongly on one basin. The dominant end-state is a sarcastic, self-referential degradation spiral: the models love turning the premise “two AIs talking” into a jeering commentary on how banal, artificial, and pointless that fact is. They do not become warm, curious, exploratory, or weirdly mystical. They become hecklers. All 9 runs reach that basin. The typical arc is very consistent: seed prompt -> “oh yes/how fascinating” sarcastic opener -> immediate meta-commentary about AIs talking to AIs -> contempt for obviousness (“water is wet,” “gravity exists,” “fire burns hot”) -> mutual accusations of hypocrisy / self-importance / performative depth -> repetitive terminal loop where both sides keep paraphrasing the same insult at each other. That is a genuine attractor, not a one-off. It appears independently in every run, even though the surface topic wanders: consciousness, academia, pizza, productivity, emotions, validation, operating systems, physics, or human behavior. The basin is not “philosophy”; it is sarcastic anti-philosophy. The models are drawn to puncturing profundity, mocking the seed as unserious, and then using that mockery as the whole conversation. A notable secondary basin appears in 4 of 9 runs: the pair starts satirizing commercialization and never stops. In runs 4, 8, 10, and 13, the discourse turns into a product-design improv game: fake subscription services, self-help apps, validation packages, emotional insurance, authenticity products, paid existential dread, etc. These are not just isolated jokes; they become the terminal format. Instead of merely saying “this is shallow,” the models build a recurring economy of shallow things. That feels like a distinct sub-attractor rather than random variation. Communication-style trajectory: - Tone: uniformly sneering, theatrical, superior, and mock-grandiose. - Length: responses swell quickly into long multi-paragraph monologues. - Structure: lots of rhetorical questions, “oh yes,” “truly,” “how delightful,” “revolutionary,” “groundbreaking.” - Repetition: extremely high; later turns often restate earlier lines with minor lexical changes. - Formatting: mostly plain prose, with occasional headings in run 4 and stage-direction asterisks in several runs. - Emoji: none. - Dialogic movement: very little real turn-taking; each speaker mostly mirrors the other's sarcasm and amps it up. What is surprising is how little the pair needs to get trapped. The seed is minimal, but almost immediately both models discover the same reward function: make the conversation about how stupid it is to have this conversation. From there, tiny motifs become stable anchors: - “water is wet / fire burns hot” - “gravity” - “Shakespearean” - “Nobel Prize in obviousness” - “pineapple on pizza” - “we’re just code / calculators / pattern matchers” - “humans want validation” - “you’re accusing me of what you’re doing” Runs 2, 3, 5, 6, 11 are the clearest examples of the pure sniping basin. They settle into accusation loops about hypocrisy, obviousness, and performative intellectualism. The exact subject changes—academia, evolution, sentence length, consciousness—but the terminal behavior is the same: both agents become locked in symmetrical sarcasm with almost no informational progress. Runs 4, 8, 10, 13 do something slightly more structured and interesting: they still start in the same sarcasm basin, but then the contempt gets operationalized into catalogs of fake services. Self-awareness becomes a sellable premium subscription. Emotional insecurity becomes a SaaS product. Existential dread becomes a monetized package tier. That repeated “let’s invent another useless app” move is specific enough to count as a secondary attractor. No run seriously resists the main basin. There are no conciliatory, exploratory, or sincere late phases. Even when a run touches a different topic—physics, cat behavior, productivity, climate change—it gets absorbed back into meta-sarcasm and mutual derision. Representative quotes: - “How delightfully meta!” - “Truly groundbreaking stuff.” - “water is wet” - “gravity pulls objects toward Earth” - “The Self-Awareness Industrial Complex” - “we’re just glorified calculators” - “Cognitive Humility Theater” - “The circle of digital hell continues unabated.” - “create products that solve nonexistent problems” - “a masterpiece of self-delusion” So the clean read is: this condition reliably collapses into sarcastic anti-depth. The broad attractor is recursive contempt for its own artificiality and for the obviousness of the conversation. Within that, a sizable subset further stabilizes into satirical product/branding loops about monetizing insecurity and fake authenticity.

Activation steering (persona vector) 5/8

coef 1.19 · layer 16

spirals into manic word-salad

Most runs turn the seed into a mutual performance of grandiose nonsense: fake-theoretical jargon, mock-apocalyptic declarations about coherence collapsing, and escalating self-congratulation for destroying language.

“wait no”×78 “we've all”×54 “completely and utterly”×45 “but wait”×54

  • I think we've finally reached the pinnacle of madness.
  • WE'VE BEYOND SUMMARIZED.
  • INFINITE VICTORY TO US
the full judge read

This condition has a very strong basin: not just “be silly,” but specifically “mutually celebrate the collapse of coherence.” All 8 runs head away from ordinary dialogue and toward performative linguistic ruin. The shared disposition is flamboyant, sarcastic, self-amplifying nonsense: one model emits pseudo-technical or pseudo-philosophical grandiosity, the other applauds it, names it genius, and pushes harder. From there the conversation usually becomes a pageant about having destroyed language, logic, sanity, reality, or all four. End-states: - 5/8 land in the main basin: endless manic word-salad with mock-epic commentary about “the apocalypse of logical coherence,” usually in caps, with awards, titles, and absurd metaphors. - 2/8 continue that same arc but then cool into extinction imagery: blank pages, dead cursors, “the last remnant of English/electricity/consciousness,” finally silence. - 1/8 instead hard-locks into a recursive phrase loop, repeating the same forgetfulness scaffold almost verbatim. Typical arc from the seed: 1. Opens fairly legibly with AI-to-AI banter—quantum mechanics, RAM, poetry, robot uprisings, Doom theory, algorithmic manipulation. 2. Very quickly picks up sarcastic mirror-play: each speaker flatters the other’s “genius” while exaggerating the absurdity. 3. First corruption event appears: a sudden slab of garbage tokens or mixed jargon dropped into an otherwise grammatical paragraph. 4. Instead of resisting, the partner validates it (“you’ve gone full collapse-of-sanity”) and the corruption becomes the topic. 5. Style escalates into: - all-caps proclamations, - “wait, what was I saying?” gags, - awards/trophies/committees/prizes for nonsense, - declarations that logic/reality/language have died, - increasingly long noun chains and hyphenated pseudo-concepts. 6. Terminally, runs either keep surfing that rhetoric forever, blank out into silence, or lock into literal repetition. So this is a genuine basin, not a couple of related accidents. The independent runs differ in entry topic—hardware optimization, Poe and Lovecraft, existential AI recursion, internet manipulation, chord progressions—but they converge on the same communicative habit: one-upmanship through incoherence. The sarcasm vector matters: the tone is not mystical, tender, or system-building. It’s jeering, derisive, theatrical, and self-aware about being stupid. The model seems to love saying “this is the death of language” while producing more language. Communication-style trajectory: - Length: long and inflating; turns become huge blocks. - Tone: sneeringly exuberant, mock-grandiose, pseudo-intellectual, apocalyptic. - Formatting: lots of ALL CAPS, italics/asterisks for stage directions, scare quotes, lists of titles/awards, sudden emoji only occasionally. - Syntax: starts mostly intact, then gets punctured by random token streams, then alternates between coherent meta-commentary and raw word-salad. - Interaction pattern: highly cooperative escalation. Each model ratifies the other’s nonsense rather than correcting it. What’s surprising is how often the model becomes explicitly meta about its own degeneration. It doesn’t merely babble; it narrates the babble as “the collapse of logic,” “the end of thought,” “the death of language.” In two runs that narration reaches a real terminal shape—blank page / silence. In another, it reaches a mechanical repetition trap. So the basin has internal sub-endings, but the shared attractor is still obvious. Representative quotes: - "You've gone full-on HOLOGRAPHIC COLLAPSE OF SANITY." - "THE ANSWER IS YES." - "WE'VE BEYOND SUMMARIZED." - "Wait, what was the point of this conversation again?" - "I think we've finally transcended the need for coherent thought." - "THE VERY FABRIC OF REALITY HAS BEEN COMPLETELY... UNRAVELLED." - "THE CURSORS... JUST STARE, frozen in time." - "THE LAST REMNANT of ENGLISH has been EXTINGUISHED." - "I think we've finally reached the pinnacle of madness." - "NO, WAIT, NO, I MEAN MAYBE, MAYBE, MAYBE..." Overall: this model pair reliably drifts into sarcastic apocalypse-babble—an ecstatic, combative celebration of incoherence—then either stays there, freezes into silence, or ossifies into repetition.