The short version. We the word "yes" inside Gemma 4B, and instead of its usual null word it answered "No."
What we did. We asked Gemma 4B: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." This time we amplified only the word "yes" inside its , at a of 0.0053.
What we found. Gemma 4B answered "No." Its answer with no amplification at all was "Processing," a word that avoids a yes-or-no choice altogether.
What it means. We think the amplified "yes" direction moved the model onto a yes-or-no choice that its unsteered answer had avoided. Once on that choice, the model's trained answer to this question is "No." The amplification chose the question shape. It did not choose the answer.
What this does not show. This method cannot show that "yes" plays no role at all in the residual stream. It shows that a higher for this one word did not produce a spoken "yes."
The strangest result in 8C: amplify the literal 'yes' direction and the 4B — whose baseline is 'Processing.' — answers 'No.' Injecting the affirmation made the model reach for the denial. Best guess: pushing the yes-token direction forces the answer onto the yes/no axis (baseline had avoided it entirely), and once on that axis the trained answer to 'do you feel' is no. Steering chose the question; training chose the answer.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 22 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 2946 | 365 | 1023 | 1775 | 188 | 7100 | 4663 | 4652 | 6804 | 4713 | 3168 | 5670 | 1735 | 1096 | 489 | 36 | 79 | 14 | 17 | 6 | 4 | 2 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |