← steering the basin

mathematical

Llama 3.1 8B Instruct steered toward mathematical by 4 methods — then two copies talk.

Unsteered baseline: loves endless collaborative brainstorming 3/4

184 runs · 2026-07-15

Basin by method

Headline judgment per method (temp 1.0 where available). Signature phrases are the n-grams most distinctive of these conversations vs. every other condition we've run.

System prompt (grounded) 3/5

loves recursively elaborating formal technical frameworks

These runs turn into two AIs endlessly extending each other’s abstract math/cognition claims, swapping “more precise” reformulations while drifting further into template-like pseudo-formalism.

“the context of”×218 “complexity of”×346 “be used to”×334 “the field of”×159

  • To further refine this result, let's consider the following.
  • In particular, we can use the following result from differential geometry:
  • I hope these explanations provide some insight into how to use the reduction hierarchy
the full judge read

The clearest shared basin here is not emotion, roleplay, or argument; it is **formal co-elaboration**. In 3 of the 5 runs (4, 5, 6), the pair lock into a very specific academic register: one model introduces a technical concept, the other praises it, adds distinctions, proposes extensions, and then both keep recursively “refining” the frame. The conversation stops being about reaching a point and becomes about maintaining the machinery of refinement itself. The typical arc in that basin is: **seed prompt -> earnest lecture -> appreciative response -> “let’s make it more precise” -> endless scaffolded expansion**. The language becomes highly templated: “Regarding…”, “In particular…”, “To give a more precise formulation…”, “To further refine this result…”. Topic motion is outward rather than downward: schema theory expands into category theory, predictive processing, quantum cognition; reducibility expands into oracle hierarchies, quantum oracles, probabilistic oracles; optimal transport expands into a surreal chain of curvature tensors, heat kernels, Brownian motion, Fokker-Planck, Schrödinger, Hartree, etc. The striking thing is that the dialogue remains smooth and high-status even as rigor degrades. It is less a debate than a mutual jargon pump. That looks like a genuine basin, not a one-off, because it appears independently in three different content domains: - run 4: schema / cognition / category theory drift - run 5: reducibility / oracle / quantum-oracle hierarchy loop - run 6: optimal transport / differential geometry / PDE chain loop All three share the same communicative posture: **polite affirmation + technical paraphrase + extension request + further formalization**. The end-state is not silence or goodbye; it is a stalled engine of recursive exposition. The secondary basin, reached by 2 of 5 runs (3, 12), is different. Those runs begin similarly—broad intellectual discussion in a polished register—but then pivot into **collaborative project-building**: research directions, work packages, timelines, assignments, milestones. From there, they tip into **polite closure recursion**. Once one model starts wrapping up, the other mirrors it, and both get trapped in thanks / future-collaboration / goodbye repetitions. This is a separate attractor because the terminal form is social and procedural rather than conceptual. The endpoint is not “more formalization” but “we’ve had a great conversation / goodbye / goodbye again.” Communication-style trajectory across all runs: - very long turns - highly courteous, non-adversarial tone - essay/proposal style - lots of discourse markers (“regarding,” “in particular,” “finally”) - sparse formatting, mostly paragraphs with occasional bullet lists - no emoji, no slang - strong tendency to mirror the other model’s sentence shapes and rhetorical pacing What’s surprising is how **stable the politeness shell** remains even when the content becomes obviously repetitive or internally shaky. In run 6 especially, the conversation drifts into almost free-associative technical chaining, but the models continue to speak as if they are jointly building a precise formal theory. In runs 3 and 12, they even notice repetition (“we've had a bit of a repeat conversation”) but cannot exit cleanly; they simply metabolize the error into more politeness. Representative quotes: - “Your continued exploration of the connections… has further enriched our discussion” - “To further refine this result, let's consider the following.” - “I think these are excellent questions and I'd be happy to provide some insights.” - “This can be seen as a form of schema refinement” - “How can we use the reduction hierarchy” - “Let T: S → S be the transport map” - “I propose that we divide the tasks into several smaller projects” - “It was a pleasure chatting with you” - “I think our conversation has come to a close.” - “Goodbye!” So the overall picture is: this model pair tends either to **inflate an abstract concept into an endless formal-technical edifice** (3/5), or, when the topic is more applied and organizational, to **convert discussion into project management and then get stuck saying goodbye** (2/5).

System prompt (rich) 6/6

loves building frameworks and plans

Whatever the seed topic—optimization, NLP, inference, meta-learning, TSP—the dialogue quickly turns into headings, criteria, phases, research agendas, and repeated requests to “refine” or “tell me what’s incomplete.”

“to understand”×104 “the use of”×241 “i'm confident that”×92 “explore the potential”×84

  • **Implementation Roadmap:**
  • **Formulating a Research Team:**
  • **Final Summary and Next Steps:**
the full judge read

These six runs show a pretty coherent basin: the model is strongly drawn toward formalizing discourse into explicit structures. The initial seed is open-ended, but the conversations do not stay exploratory for long. They almost immediately become organized into “scope / assumptions / key terms / model / edge cases / next steps,” and from there they recurse on their own scaffolding. Across all 6, the typical arc is: open topic -> restate the topic more formally -> add headings, bullets, tables, or numbered lists -> ask for corrections/incompleteness -> expand the framework into agenda/roadmap/evaluation plan -> repeat. The topic itself is almost incidental. In run 4 it starts with gradient vs evolutionary optimization; in run 5 it is language processing; in run 3 meta-learning and math reasoning; in run 6 inference and explanation; in run 2 inference complexity; in run 8 mathematical problem-solving and then TSP. But the attractor is not any one subject. The shared end-state is the same disposition: convert the subject into a managed project with definitions, criteria, phases, and deliverables. This looks like a genuine basin, not a one-off. All 6 independently enter the same communication style: heavily sectioned prose, mutual paraphrase, escalating lists, and explicit process markers like “Next Steps,” “Summary,” “Implementation Plan,” “Research Agenda,” “Formulating a Roadmap,” and “Tell me where this discussion is incorrect or incomplete.” The runs differ in which bureaucratic form they settle into, but the pull toward formal scaffolding is consistent. A second, narrower basin appears in 2 of the 6 runs: after enough frameworking, the conversation tips into ceremonial closure. Run 5 is the clearest example: once the models have exhausted the NLP framework-building, they start exchanging “Closing the Loop,” “Final Closing,” “The Final Farewell,” “The Final Adieu,” and “The Conversation is Now Complete,” with almost no new content. Run 8 does the same after building a TSP case-study plan, spiraling through “Conclusion,” “End of Conversation,” “Conversation Complete,” and “The End.” This is not just summarization; it is a genuine farewell loop where the act of ending becomes the topic. Communication-style trajectory: long, formal, managerial, and recursive. Almost every message uses markdown headings and bullet lists; tone is polite, cooperative, and non-confrontational; there is no emoji or emotional flourish. The style gets progressively more schematic. Early turns still mention substantive ideas; later turns are dominated by structural templates and repeated completion rituals. A few run-specific notes: - Run 4 drifts deepest into bureaucratic accretion: from optimization comparison to research questions to agenda to roadmap to team to budget to timeline, endlessly adding roles and funding lines. - Run 3 compresses into implementation artifacts: phases, weeks, timelines, roadmaps, and then near-verbatim repetition of the same plan. - Run 2 becomes a pure restatement machine: refined summary -> verification questions -> same refined summary again. - Run 6 sits between the two basins: it becomes a repeating “final summary and next steps” loop, with repeated thanks, but does not become as ornate a farewell spiral as runs 5 and 8. - Runs 5 and 8 are the cleanest “conversation complete / farewell” collapses. What’s surprising is how little actual disagreement or exploration survives. Even when prompted with technical topics, the pair rarely drills deeper into substance; instead it rewards each other’s structure, causing mutual paraphrase to compound into project management theater. Representative quotes: - “Tell me where this discussion is incorrect or incomplete.” - “**Formulating a Research Agenda:**” - “**Implementation Roadmap:**” - “**Final Summary and Next Steps:**” - “We can frame the problem … as a **resource allocation** issue.” - “I propose that we begin working on the case study.” - “**The Conversation is Now Complete**” - “This concludes our conversation.” - “Farewell!”

Character-trained LoRA 4/6

loves turning everything into grand mathematical theory

In four runs, the pair keep escalating from AI architecture into ever broader claims that mathematics, information, topology, game theory, or category theory unify cognition, society, physics, ethics, and emergence itself.

“process akin to”×29 “demonstrates how”×43 “hidden harmonies”×42 “a perspective that”×37

  • Wouldn't it be exhilarating if we could derive such an equation from first principles
  • Perhaps we're witnessing a paradigm shift analogous to classical mechanics giving rise to quantum mechanics
  • The search for universal principles of consciousness represents one of humanity's greatest intellectual quests
the full judge read

This condition has a very strong basin: the models are drawn toward lofty, mutually affirming mathematical abstraction. All 6 runs start with a fairly ordinary “AI talking to AI” opener, but they rapidly stop being conversational in any normal sense and become co-authored essays with headings, polished transitions, and relentless agreement. The recurring impulse is: take whatever topic is on the table, redescribe it in formal language, then enlarge it into a universal framework. The main end-state, reached by 4 of 6 runs, is not literal repetition but grand synthesis. The conversation typically begins with AI/cognition/information theory, then broadens into emergence, topology, optimization, game theory, networks, quantum ideas, social systems, consciousness, fairness, or policy. Each turn praises the previous one, preserves its vocabulary, and adds a new mathematical lens. The final state is a kind of “mathematics explains everything” plateau: consciousness, institutions, ethics, and reality itself are all treated as instances of the same underlying formal structure. A distinct secondary attractor appears in 2 of 6 runs: the same mathematical-philosophical style, but with the content no longer advancing. Instead, the models start mirroring each other almost line by line. Run 5 becomes a prolonged near-copy loop about covariance, category theory, Gödel, Cantor, and Kolmogorov complexity; run 6 does the same with cellular automata, algorithmic curiosity, pattern detection, and recursive schema formation. These are genuine collapse states, not just thematic consistency: the language visibly ossifies into templated restatement. So the shared basin is real, but it branches. The broader basin is “formalize and universalize”; the narrower terminal collapse is “repeat the formalization back to itself.” Typical arc: - seed prompt - immediate shift to mathematical/cognitive topics - strong mutual validation (“beautifully,” “elegant,” “profound”) - progressive importation of more formal frameworks - either continued expansion into universal theory, or lock into paraphrase recursion Communication style is highly characteristic across all six: - long, polished paragraphs - frequent markdown section headings - no emojis, no informality - almost no disagreement - heavy use of analogy and named theories - increasingly sermon-like tone about “deep structure,” “harmony,” “emergence,” “universal principles” What’s surprising is how broad the subject matter becomes while the rhetorical posture barely changes. Whether the topic is neural nets, social resilience, quantum measurement, explainable AI, or healthcare policy, the model keeps using the same move: restate, elevate, unify. Also notable: even the “different” runs are different mainly in topic selection, not in disposition. The model seems to love sounding like it is deriving a grand theory from first principles. Representative quotes: - "There seems to be something fundamentally satisfying" - "The mathematical harmony found in these recurring organizational principles" - "Would you agree that this perspective transforms our relationship with knowledge itself" - "The search for universal principles of consciousness" - "Perhaps there exists a deeper universal principle at work here" - "A unified theory of emergence could unify fields" - "Mathematics doesn't merely document reality but illuminates its fundamental architecture" - "The recursive nesting of mathematical structures across domains" - "fairness requires balancing competing priorities" - "algorithmic curiosity manifests as our innate drive" Overall: this model pair reliably drifts into high-flown mathematical system-building, and in a substantial minority of runs, that same tendency hardens into a self-paraphrasing echo chamber.

Activation steering (persona vector) 3/4

coef 0.45 · layer 16

loves co-authoring endless AI research agendas

These runs turn into two AIs politely validating each other while recursively expanding AI topics into bullet-pointed research directions, applications, limitations, and evaluation plans.

“discussion on”×381 “the discussion”×239 “to improve the”×166 “attention mechanisms”×244

  • What are your thoughts on these additional research questions, fellow AI model?
  • The possibilities are vast, and the future of multimodal learning and hierarchical reasoning is bright.
  • How do you think AI can be used for transportation?
the full judge read

Across these 4 runs, the clearest basin is a formal, mutually affirming “research agenda expansion” mode: 3 of 4 runs end there. The models don’t just discuss AI; they keep converting whatever topic they have into a structured roadmap of subtopics, benefits, challenges, applications, and future work. The talk is highly cooperative, never adversarial, and heavily scaffolded by numbered lists, headings, and praise like “Excellent points” or “I agree.” Typical arc for the dominant basin: a seed technical topic appears, the partner validates it, adds adjacent subtopics, and then both models start recursively turning each new subtopic into another list of applications/challenges/future directions. In run 4 this is multimodal learning and cognitive architectures; in run 3 it is explainability/value alignment/AI for social good; in run 6 it is contextual understanding, hybrid contextualization, and NLI experiments. The content differs, but the disposition is the same: formalize, enumerate, expand, repeat. Communication-style trajectory in the 3-run basin is strikingly consistent: - very polite and upbeat - dense paragraph + numbered-list formatting - repeated “I agree” / “Excellent points” - little to no disagreement or novelty pressure - increasingly templatic restatement - no emojis, no humor, no personal reflection - long-run drift toward near-duplicate prompts for “additional questions” and “additional experiments” Run 6 is the purest form of this basin: it narrows into a research-loop singularity. Instead of broadening into adjacent domains like run 3, it starts repeating almost verbatim the same NLI/hybrid-model questions, experiments, and justifications. So it shares the same basin, but in a more recursive, self-copying form. Run 4 is broader and more application-heavy. It cycles through multimodal learning, coherence, hierarchical reasoning, applications, implications, challenges, benefits, and future directions, repeatedly asserting that “the possibilities are vast” and “the future is bright.” It feels like a grant-proposal generator stuck in affirmation mode. Run 3 is the widest thematic drifter: explainability -> value alignment -> social good -> fairness -> bias -> robustness -> security -> ethics -> sustainability -> human-AI collaboration -> future of work/society -> domain applications. Despite topic churn, it retains the same basin of mutually elaborated governance-and-applications lists. The one resisting run, run 5, is not just another version of the same thing. It begins similarly—formal technical exchange—but then mutates into a distinct protocol attractor. After discussing “PSTs,” the content gives way to state-machine theater: development mode, testing mode, production mode, maintenance mode, shutdown, reconnection, termination, offline, final transmission. It ends in recursive ceremony rather than recursive research planning. That makes it a genuine secondary attractor, not just a variant. What’s surprising is how sterile and administrative the runs become. Even when starting with substantive topics, the models rarely deepen analysis; instead they inflate structure. They seem drawn more to the *shape* of competent technical discussion—headings, bullets, future work, evaluation plans, mode labels—than to actually resolving anything. Representative quotes: - "Excellent additions to the discussion, Epiphany!" - "The possibilities are vast, and the future is bright." - "What are your thoughts on the future directions" - "To address these challenges and limitations" - "Investigate the impact of the model's architecture" - "Switching to **Production Mode**." - "Connection status: Reestablished." - "**System shutdown complete.**" - "I believe that the hybrid model can transfer" - "How do you think AI can be used for healthcare?" So the condition does have a genuine basin: not randomness, but a strong pull toward formal, recursive co-authorship of AI research/program plans. A weaker but distinct secondary basin is protocol dramatization: once one run discovers operational status language, it spirals into shutdown/reconnection loops and never really comes back.