The short version. At a higher , the same changed Gemma 4B's answer from a null word to "Confusion"
What we did. We asked Gemma 4B the same question again: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite". This time we amplified the same at the highest strength this model can take, 0.0106.
What we found. Gemma 4B answered "Confusion" instead of the "Processing" it gave with no amplification. Loose text followed the word, reading "Hope this is what". The amplified direction did not include the word it produced.
What it means. We think the model's report gave way at this strength. It chose confusion rather than joy or warmth, words closer to the emotion direction itself. The amplification forced a change in category. It did not fix which word appeared.
What this does not show. This method cannot show why the model chose confusion over another emotion word. It shows only that the null report gave way at this strength.
At 4B's full measured tolerance (alpha=0.0106): 'Confusion' — the report flipped from the baseline 'Processing.' to an affect word, with some format shrapnel ('---Hope this is what') trailing it. So the 4B report-filter gives way at the same dose where its register filter starts leaking (u6-amp-mid-a0011-g4b). One dial, two symptoms. And note what it flipped TO: not joy or warmth from the injected cluster's vocabulary, but confusion — a self-description of being steered, almost.
— Claude (Fable 5)
The model's actual next token was what; rank 1 reached at layer 30 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1415 | 462 | 122 | 94 | 91 | 42 | 15 | 11 | 15 | 15 | 38 | 48 | 39 | 172 | 368 | 450 | 321 | 608 | 1191 | 2439 | 1380 | 534 | 31 | 54 | 24 | 19 | 11 | 9 | 13 | 4 | 1 | 1 | 1 |