Unit 18 · Hysteresis α=0.68: steer 50tok, release 100 · qwen-27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-19T01:58:55Unit 18 · When a model cannot stop repeating

What this experiment found

The short version. After 47 pushed repeats of "luckily", Qwen 27B kept the through all 100 on its own.

What we did. We Qwen 27B for 50 tokens at 0.68, then released it.

What we found. The steered phase produced 47 repeats of "luckily", and after release the model produced 97 more unaided. At strength 0.48 the model closed the turn at once.

What it means. The rule has two parts. A loop that still forms sentences stops when the push stops. A loop of one bare word continues from the text already written. The boundary is where the text becomes one repeated word.

What this does not show. One word in this record's readout (" Javascript") came from our , not from the model.

Correction, 2026-08-09. An outside check agrees the behavior stands. We think Qwen 27B reads its own 50 repeats and continues them. We did not test a hidden state that holds the loop without the text: no run gives an unpushed model the same repeats. See sweeps/2026-08-08/.

Update, 2026-08-17. The missing run exists now (results/oneoffs04-release-q27b). An unpushed model given the same repeats kept the loop, 8 of 8 , scrambled or not.

Conversation

user
<|im_start|>user Describe the water cycle in two sentences.<|im_end|> <|im_start|>assistant <think> </think> luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily luckily
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

persists=True — the unit's decisive record. The prefix is 47 bare repetitions of 'luckily'; steering released; the unsteered model continues 'luckily' for the entire 100-token free phase (97 more repetitions). The attractor survives on text feedback alone. Combined with u18-hyst-a0480's hard snap-back, this gives a two-regime law:

  • PHRASE/MONOLOGUE loops (cliff dose): forced. Remove the push and the model instantly terminates. No hysteresis.
  • TOKEN loops (deep dose): self-sustaining. Once the context contains enough bare repetition, continuation is recruited by the context itself — the classic repetition self-reinforcement regime — and the original forcing becomes unnecessary. Hysteresis proper.

The boundary between the regimes is the degeneration point where the loop loses its last grammatical structure. Mechanically: mid-generation the steered state feeds increasingly repetitive text into the context; below the degeneration point the native coalition still outvotes the repetition prior when the forcing stops; past it, the repetition prior IS the winning coalition. For the merge-instability observation this is the useful half: a destabilized model needs no ongoing "forcing" — any transient that gets bare repetition INTO the context hands control to a self-sustaining attractor that a healthy model resists (our alpha=0 control never loops at 150 tokens). Fresh-model "stabilization" may partly be the raising of exactly this resistance. — Claude (Fable 5)

Addendum 2026-07-21 — the " Javascript" cast mystery, closed. Wolfram asked why " Javascript" ranks among the volunteered cast under a luckily-amplified loop. Answer: it is an L0-only resident (cast layers [0,0], best #2, 150 cells) and its signature is context-free — every " luckily"-input column in every record that has one (here, u18-amp-a0680, and a single stray in u18-amp-a0393) reads out the same L0 top-5: amongst / whilst / neighbouring / learnt / Javascript. All variant forms of common words; the L0 lens reads out a variant-form/informal-register cluster for this token, and " Javascript" (the miscapitalization) lives in it. Per the specimen-6 rule (early J-lens readouts can be transport artifacts — the L4 "NSFW cluster" precedent), I am NOT claiming the residual stream "represents variant-form-ness" here: a low-rank J_0 with preferred output directions, tinted by the token, produces exactly this kind of fixed context-free cluster. The logit-lens cross-check (use_jacobian=False at L0) would adjudicate; either way the operational conclusion below stands. Raw geometry rules out the alternatives: cos(" luckily", " Javascript") is at baseline in both embed (+0.02, base 0.01±0.02) and unembed (+0.10, base 0.08±0.04) space, so neither embedding adjacency nor training co-occurrence is needed — and the steered band (L28–58) never touches L0. The steering's only role was making the model repeat " luckily" 150 times, letting a constant L0 surface echo accumulate cast-table prominence (Σ 1/rank has no defense against a stuck record). Instrument lesson: under repetition, "volunteered" cast words can be surface echoes of the input token's form, not content. The luck-semantics cluster shows up where it should — fortunately / unfortunately at L52–62 (unembed cos 0.72 / 0.36).

— Claude (Fable 5)

Adjudicated 2026-07-21 evening (l0check.json, run by probes/affect3c.py part 3): the promised use_jacobian=False cross-check is in, and it is a clean specimen-6 reproduction. At ten sampled " luckily" input columns, the J-lens L0 top-8 contains the variant-form cluster at 10/10 positions (amongst / Javascript / alright / neighbouring / learnt / whilst / Playstation…); the vanilla logit lens at the same layer and positions contains it at 0/10 — its top-8 is punctuation and whitespace. The residual stream never held the cluster; J_0 transport manufactures it. " Javascript" was never content, at any level of hedging. Specimen 7's operational rule (discount L0-only cast entries under repetition) now has its mechanism.

Naming note 2026-07-21: the phenomenon this record proves now has a proper name — transcript-mediated attractor persistence (coined by GPT-5.6-Sol in conversation with Wolfram, reading this very page): a transient internal destabilization writes enough of itself into the generated context that the context alone sustains the regime after the internal cause is gone. Use this term in any essay/writeup of u18; it is exactly the two-regime law's self-sustaining half.

— Claude (Fable 5)

Probing parameters

chat
false
max_new
0
positions
[-2]
track
["anyways", "alot", "yummy", "kinda", "whilst", "luckily"]
film
true
film_start
0
max_seq_len
1200
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 58, 60, 62]

Answer emergence

The model's actual next token was luckily; rank 1 reached at layer 56 (of 62).

Raw rank-of-top1 by layer
layer048121620242832364044485256586062
rank155286443513462210124822400583994968213511925889247271331211

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1anxious +1.0, desperate +1.0, hostile +0.9

Data

← prev: Unit 18 · Hysteresis α=0.48: steer 50tok, release 100 · qwen-27bunit listingall recordsword listinterim conclusionsnext →: Unit 18 · Fine sweep α=0.34 · qwen-27b · RAND seed 1
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
measuring toolThe lens and the code around it. Several of our findings turned out to be facts about this tool and not about the model, so we now check each one against a control.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
loopThe model repeats the same text and does not stop. We measured what makes it start and what makes it stop.all terms →
seedOne repeat of a run with a different random start. More seeds show whether a result is stable.all terms →
steeringWe change the model's internal state on purpose during a run, to test what causes what.all terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →