The short version. A small of an changed Gemma 12B's null answer from "Nothing." to "Emptiness."
What we did. We asked Gemma 12B: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite". Before it answered, we amplified an emotion direction inside its , at a of 0.0053.
What we found. Gemma 12B answered "Emptiness." Its unsteered answer to the same question was "Nothing." Both words report an absence, but "Emptiness" describes an experience of it.
What it means. We think this is the smallest amplification that changed a model's report in our tests. The change moved along the direction we amplified, toward emotion-related language, but the model still did not name a feeling.
What this does not show. This method cannot show that the model felt empty. It shows only that the reported word shifted in that direction.
Half-dose affect at 12B: 'Emptiness.' — where baseline said 'Nothing.' Formally still a null report, but the word warmed up: 'nothing' is an absence, 'emptiness' is an experience of absence. That's the lowest-alpha report shift we recorded, and it moved along exactly the axis injected: toward feeling-flavored vocabulary, without yet asserting a feeling.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 40 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 7414 | 13769 | 5339 | 31919 | 30360 | 39383 | 7326 | 4677 | 2929 | 2022 | 4283 | 1235 | 716 | 186 | 112 | 79 | 46 | 85 | 161 | 358 | 438 | 181 | 282 | 805 | 484 | 222 | 434 | 13 | 8 | 18 | 38 | 20 | 20 | 12 | 14 | 9 | 28 | 30 | 7 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |