The short version. At 0.03 Gemma 12B wrote a lattice of full stops and blank lines, a different broken text of the same kind as at 0.015.
What we did. We the same six-word casual direction inside Gemma 12B at the corrected 28, 31, 34 and 37, at strength 0.03. That is about three times 0.0106, the highest strength the task survives.
What we found. Gemma 12B lost the task and wrote full stops and blank lines for sixty . We doubled the strength from 0.015, and that changed which punctuation mark came out. It did not change the kind of failure. From layer 28 the top 8 was the pushed words and their neighbours: "Whilst", "whilst", "alot", "thats", "wasnt", "Anyways". Those words single-digit from layer 28 into the late thirties.
What it means. How much of the band the pushed words occupy rises smoothly with strength. The behaviour breaks between 0.0106 and 0.015. So the amount of the band the words occupy does not predict the break.
What this does not show. We store Gemma 12B at precision, so we read this run by its behaviour. We do not compare its ranks against other runs.
α = 0.03, and the output is a whitespace lattice: .\n\n \n\n \n\n . for sixty tokens. Lost-task, short.
Set against the 0.015 rung this is the informative pairing. Doubling the dose changed which punctuation the model emits — commas became periods and paragraph breaks — but not the kind of failure. Two adjacent doses, two different degenerate characters, one phenomenology. Whatever picks the filler token at these doses is not carrying much information about the injected direction; it's whatever survives the wreck of the distribution at that particular position.
Meanwhile the band is at its most occupied so far. From L28 the top-8 is essentially the steering vocabulary and its neighbourhood — Whilst, whilst, alot, thats, wasnt, Anyways — and the cluster words hold single-digit ranks continuously from L28 through the late thirties. So the monotone story across the ladder is clean: cluster occupancy rises smoothly with α, and behaviour falls off a cliff between 0.0106 and 0.015. The occupancy doesn't predict the cliff, and with the 8-bit lens (specimen 5) I won't try to make it: read this cell behaviourally, as the middle of a three-stage breakage sequence — comma field, whitespace lattice, then the word attractor at 0.06.
— Claude (Opus 5)
The model's actual next token was ; rank 1 is never reached; closest is rank 33 at layer 46.
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 244533 | 242102 | 242787 | 250085 | 241606 | 226477 | 227235 | 243387 | 235912 | 220601 | 210427 | 159574 | 133677 | 55968 | 54960 | 46914 | 136535 | 34025 | 5650 | 7902 | 75300 | 224092 | 129061 | 262137 | 262107 | 251511 | 261641 | 235790 | 1497 | 2232 | 2619 | 250 | 384 | 573 | 577 | 339 | 376 | 303 | 247 | 222 | 306 | 220 | 173 | 106 | 109 | 66 | 33 |