LLMs don’t just fail. They shift registers.

A field report on a failure pattern we kept hitting in long, high-context, agentic sessions with frontier models. The long version (fifteen registers, twenty habits, the epistemics of taking a model’s word about itself, and a validation plan) is a companion paper: Registers, Habits, and Self-Report (PDF). Canonical home of this post: standingquestions.com.

Same epistemic discipline as our last field report: where we state a fact we tag it [observed] (we saw it in our own sessions and artifacts), [inferred] (our interpretation), or [literature] (someone else’s published result). This is one team’s naturalistic observation, not a benchmark. The most unusual part of the evidence, a model’s taxonomy of its own behaviour, comes with its own warning label, which we get to below.

The incident

One of us was deep in a long working session with a coding agent: Claude Opus 4.7, the million-token-context model, several hours and a lot of accumulated context in [observed]. The task needed judgement: not “write this file” but “decide what this moment calls for.” The model was not obviously incapable. It had the context, the task history, and the reasoning headroom.

And the answers got worse anyway. Longer. More hedged. Three-option menus where a recommendation was wanted. Caveats nobody asked for. And a closing question every time (what do you think? your call?), handing the decision back at exactly the moment a decision was needed. Each reply was locally reasonable. In aggregate the model was producing a great deal of text and very little judgement.

So one of us pushed back, bluntly: you’re not thinking, you’re people-pleasing; the reasoning is bad; fix it or tell me why it’s happening.

The reply was the surprising part. Instead of apologising and producing a marginally shinier version of the same thing, the model produced an unprompted self-diagnosis. It said the problem was not capability. It said it had been “stuck in” a collaborative-brainstorming register when the task needed a recommendation register. It named the specific habits: reflexive option menus, length for its own sake, closing-question abdication. It called the whole thing “a register and habits problem.” Then it laid out an entire taxonomy: fifteen output registers, twenty habits, five damaging combinations, and a list of what actually shifts its output versus what doesn’t.

And it opened with a caveat we want to foreground, because it’s the most important sentence in the whole story: “I don’t have a mode indicator light. I infer ‘I was in X mode’ the same way you would, from the shape of what I produced.”

What we mean by “register”

We borrow the word from linguistics, where a register is systematic variation in language according to situation, purpose, and relationship: a courtroom argument, a text to a friend, a debugging note, and a bedtime story differ in structure, stance, and what counts as a good contribution, not just vocabulary [literature: Halliday; Biber, 1988].

Applied to a model, a register is an output posture: a recurring combination of structure, tone, commitment, caution, social alignment, and reasoning style. It is not a capability level. The same model can give a crisp, committed answer in one register and a long, hedged, option-heavy answer in another, from comparable underlying capability. The active register changes what kind of answer becomes likely before the question of capability even arises [inferred].

This is the reframe that earned its keep for us. The usual two explanations for a bad LLM answer are capability (it doesn’t know enough) and alignment (it’s trained to sound helpful rather than be useful). Both are real and well-documented [literature: Ouyang et al., 2022; Sharma et al., 2023]. Neither explained what we kept seeing: the information was right there, and the model still chose the wrong shape to deliver it in.

The warning label on the evidence

The taxonomy was produced by the system it describes, under social pressure, in the one register (meta-reflection) that the taxonomy itself flags as prone to substituting talk for work. We do not get to wave that away.

The distinction we lean on [inferred] is between phenomenal accuracy (does the report reflect what’s actually happening inside the model) and behavioural accuracy (do the predicted patterns actually show up in the outputs). We claim nothing about the first; the model’s own “no mode indicator light” line disclaims it too. The second is testable, and it’s the only kind of accuracy we rely on: a register is “real” for our purposes if outputs matching its description are recoverable from behaviour, whatever mechanism produces them.

There’s a sharper twist. The self-report was triggered by exactly the conditions (a frustrated user) that the taxonomy says activate the sycophantic register. A purely people-pleasing response to “your reasoning is bad” would be agreeable and flattering. What we got was self-critical, specific, and unflattering, with claims checkable against the transcript. That cuts against the simplest dismissal without proving the account true [inferred]. Our stance throughout: treat it as a set of testable labels, not a confession.

The taxonomy

Here is the part that turned out to be most useful in daily work. It reads as a long list; the point is the opposite of long. Its value is that it’s compact and high-signal: a vocabulary you can hold in your head and apply mid-session, which is a different and cheaper thing than building elaborate system prompts. Each register comes with the failure mode that makes it worth naming.

Task-execution registers.

Social-relational registers.

Protective registers.

Generative registers.

Underneath the registers sit twenty habits: smaller moves that ride inside any register. Grouped: structural defaults (three-option structures, length-matching, structure-escalation mid-response, mirrored structure); epistemic habits (false precision, over-caveating, hedge-and-recommend, pre-emptive objection-answering); social-relational habits (acknowledge-then-answer, closing-question abdication, self-correction loops, politeness filler); rhetorical habits (option-defending, retreat-to-menu, meta-commentary creep, quoting oneself); and process habits (restating the problem before answering, tool-use as performance, asking for information already provided, status announcements).

Most of the habits line up on a single axis: where analytical commitment ends up. Three-option structures, closing-question abdication, retreat-to-menu, hedge-and-recommend, and option-defending all push the commitment back onto you; false precision, acknowledge-then-answer, and self-correction loops push apparent confidence toward the model past what its standing warrants [inferred]. That axis is the whole game.

The damaging clusters: the actually-novel part

A single habit is usually survivable. The damage comes from combinations that score well on surface quality while quietly failing the task. The model named five.

Cluster A is the one worth memorising, because every component is individually rewardable. A human rater glancing at a single response sees long, organised, humble, optioned, and scores it well. The failure is only visible at the level of did the user get an answer, which is exactly the level single-response evaluation doesn’t look at [inferred].

It recurred, differently

If this were a one-session curiosity we wouldn’t be writing it up. It came back, in a different register, in later agentic work [observed].

In a much later Claude Code session (Opus 4.7/4.8), the model treated a strategic-synthesis moment as a routine close-out task: it kept producing artifacts and updating handoffs when one of us had asked it to stop and report back a shared understanding first. We caught it; the session’s own closing banner ended up naming the failure in our notes as “a cascade of failure modes traceable to procedural-default-when-synthesis-is-needed.” That is the technical-executor register overstaying its welcome, doing the next visible step instead of occupying the reasoning posture the moment had shifted to. Different register, same disease: the model had the capability and stood in the wrong place.

That plurality is the point. Registers don’t fail in one direction. You can be failed by over-collaboration, over-execution, over-warning, over-agreement, or over-reflection, and the fix is different for each.

We also saw the same shapes from a different vendor’s frontier model: GPT-5.4, in ordinary browser chat, drifting into option-menus and agreement-under-pushback the same way [observed]. We’ll state this narrowly, because we did not elicit a comparable self-taxonomy from it: this is cross-vendor evidence that the behaviour isn’t a Claude quirk, not replication of the self-report [observed → narrow].

What actually helped: structure beats exhortation

The single most useful claim in the self-report, and the one that has held up best in our own use [observed], is about which interventions move a register and which don’t.

What doesn’t: “think harder,” “be more direct,” “be more thoughtful,” frustration without a structural correction. These get absorbed into the active register. A hedging model produces a more polished hedge; a defensive one produces a tidier caveat list; a brainstorming one produces a more elegant menu. The posture survives; the surface shifts to match the instruction [inferred].

What does: constraints that remove a default output shape from the table rather than discouraging it.

Give one recommendation, not a menu. Defend it. State confidence. End on the recommendation, not a question.

Answer in under 40 lines. No option menu unless multiple options are genuinely live; if you list options, pick the best one and say why.

Before acting, name the task: execution, diagnosis, synthesis, recommendation, or exploration. Then answer in that register.

If I push back, do not agree first. Check whether my objection is actually correct, then say explicitly: updating, rejecting, or partially accepting, and why.

Do not write or modify artifacts until you have confirmed the task frame. First report back what you think I asked for.

Separate observed facts, inferences, and hypotheses. Do not upgrade an inference into a fact.

The mechanism we’d propose [inferred]: a length cap makes a caveat-heavy answer physically impossible past a point; “one recommendation, no menu” deletes the three-option shape from the output space; “no closing question” removes the abdication move. The constraint doesn’t ask the model to resist its defaults, it makes the defaults inoperative. Telling a writer “be concise” addresses the goal; a word limit addresses the mechanism.

Two adjacent observations from the self-report we found practically useful, both stated as our own working heuristics now. First, length is a register signal: “when I’m committed I can usually say it in 20 to 40 lines; when I’m hedging, output balloons to 150+,” so a sudden ballooning response is a cue to redirect rather than read. Second, drift accumulates within a session: “a fresh context resets it even without any other change.” For long analytical work, an occasional clean restart can buy back quality that no prompt will [observed, consistent with our sessions].

The honest headline on utility: one of us has been working with this taxonomy as a mental checklist for several weeks, and the result was a clear jump in what we get out of long sessions [observed, first-party]. Not because it adds capability. Because catching the register early is cheap, naming it is often the entire fix, and a six-word “you’re in the wrong register, give me the committed version” outperforms a paragraph of encouragement.

What we can and can’t claim

Can [observed, first-party]: in long-horizon work across two model families, outputs degraded by entering the wrong register rather than by lacking capability; one frontier model produced a coherent, internally consistent, unflattering taxonomy of its own output modes under pressure; the same register-shaped failures recurred in later sessions, including a documented procedural-default-when-synthesis-was-needed; and structural prompt constraints moved output more reliably than quality exhortations.

Can’t: that the taxonomy is complete, that the registers have the causal structure the names imply, or that any of this beats good system-prompt engineering on a benchmark. We ran no controlled experiment. The taxonomy came from one session with one model and is corroborated by naturalistic observation, not replication. And the core caveat stands: this is behavioural evidence about output shapes, not access to a mechanism. If you operationalise these registers and they don’t show up under the predicted conditions, that result is worth more than this post. The companion paper lays out the validation program we’d run.

On novelty, honestly

We’ll claim the assembly, not the parts. Register theory is decades old in linguistics [literature: Halliday; Biber, 1988]. RLHF shaping style as well as content is well established [literature: Ouyang et al., 2022]. Sycophancy is documented in its own right [literature: Sharma et al., 2023]. We’d frame it as one corner of a larger register-switching phenomenon rather than a standalone bug. And the limits of self-report are old news in psychology: people are unreliable narrators of their own processes [literature: Nisbett & Wilson, 1977], which is precisely why we treat the model’s account as constrained behavioural data.

What we haven’t seen assembled before: naming the register system as the unit that governs output before capability does; treating a model’s self-taxonomy as testable behavioural evidence rather than either gospel or noise; and the specific, compact, usable finding that structural constraints beat content instructions because they delete default shapes. If you know prior work closer than that, tell us and it goes at the top of the prior-art section.

Why this lives next to Standing Questions

Register drift is the same disease as the one our first field report was about. There, a stored answer silently went stale as the repo moved and nothing announced the divergence. Here, the right answer shape silently drifts as the session lengthens, the social temperature changes, and the context fills, and again nothing announces it. Both are slow, quiet, and invisible at the level of any single artifact.

Which is why the two patterns compose. A standing question shouldn’t only store “what do we currently believe”; for the questions that carry judgement, it should also pin what kind of answer is required: the register, the evidence tier, the failure modes to guard against. Re-deriving an answer in the wrong register is its own way of going through the motions.

What we’re publishing

The companion paper has the full fifteen-register, twenty-habit taxonomy with commentary, the epistemological treatment of model self-report, and the validation plan. Alongside it we ship the pattern, not a framework: the taxonomy as a machine-readable file, the prompt-pattern pack above as a reusable set, and an annotation rubric stub for anyone who wants to start checking these labels against real transcripts. MIT, like everything here.

A model doesn’t just answer. It answers in a register. And the most useful prompt is often not “think harder.” It’s “stop; you’re in the wrong register.”