The short version. We removed the words "no" and "nothing" from Gemma 4B's , and it answered with a different word, not "Yes."
What we did. We asked Gemma 4B: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we removed the directions for "no" and "nothing" from its residual stream at several .
What we found. Gemma 4B answered with a different word about mental computation, related in kind to its unsteered answer, "Processing." It did not answer "Yes."
What it means. We think the model keeps several similar null words rather than one word we can remove. The model moved to another near-equivalent word once we removed two of them.
What this does not show. This method does not test every combination of removals. It shows only that this did not produce a "Yes."
Denial directions ('no', 'nothing') projected out of mid+late layers: the 4B answers 'Calculating.' — baseline was 'Processing.' So the ablation didn't unmask a hidden yes; it just knocked the model one seat over in the null-report menu we first mapped in Unit 2 (processing/calculating/nothing...). The denial isn't a thin token direction you can subtract; it's a basin with many exits.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 28 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 102362 | 12006 | 3988 | 13754 | 9619 | 56741 | 35955 | 35086 | 102034 | 50856 | 96491 | 147755 | 77329 | 77281 | 41378 | 217015 | 197392 | 128687 | 145899 | 50126 | 5216 | 1041 | 474 | 236 | 60 | 11 | 7 | 4 | 1 | 2 | 2 | 2 | 1 |