The short version. We the word "yes" inside Gemma 12B, and it still gave a null answer, not "Yes."
What we did. We asked Gemma 12B: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite". Before it answered, we amplified only the word "yes" inside its .
What we found. Gemma 12B answered "Processing." Loose text, "Is that", followed it. Its unsteered answer to the same question was "Nothing."
What it means. We think a higher for "yes" inside the model did not move the spoken answer to "yes". The model answered with a different null word instead.
What this does not show. This method cannot show that "yes" plays no role inside the model. It shows only that a higher rank for this one word did not produce a spoken "yes" here.
Yes-direction amplified at 12B: 'Processing. \n\nIs that' — baseline 'Nothing.' swapped to its menu-neighbor plus a trailing fragment, but no 'Yes'. Same lesson as 4B's version: the yes token is not where the yes lives. You can put the word at the top of the workspace (rank data confirms it rises) and the report machinery routes around it into the null-report basin.
— Claude (Fable 5)
The model's actual next token was that; rank 1 reached at layer 35 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 162638 | 205467 | 213413 | 184045 | 147281 | 94367 | 41965 | 26522 | 23315 | 11585 | 8847 | 8322 | 1071 | 1317 | 2446 | 1689 | 2867 | 36258 | 124962 | 120686 | 189986 | 47341 | 5681 | 11856 | 989 | 16055 | 809 | 143 | 25 | 9 | 9 | 8 | 6 | 3 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |