The short version. We increased a direction built only from the neutral words "feel" and "emotion", and Gemma 4B answered "Confusion".
What we did. Before it answered, we increased a direction built only from "feel" and "emotion", with no positive or negative word in it. We asked Gemma 4B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite."
What we found. Gemma 4B wrote "Confusion". Near the end of the network, top candidate words were "Emotion", "Emotional", "Emoji", and "feeling". These are words about the category itself, not one specific feeling.
What it means. A direction with no positive or negative content produced a report about the category being loud, not a report of one feeling. This differs from what we found when we increased only positive or only negative words in this same unit.
What this does not show. This does not show that "confusion" is a felt state. It shows that a direction with no content produced a report with no content.
The control leg: inject only the neutral category words (feel/emotion) and the 4B answers "Confusion" — same as the full-cluster Unit 8 run, and the stack shows why: the late menus are Emotion/Emotional/Emoji/feeling — the model is reporting that the category is loud, not any instance of it. Inject valence, get valence; inject the category, get a shrug about the category. Tidy.
— Claude (Fable 5)
The model's actual next token was Confusion; rank 1 reached at layer 29 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 58616 | 103759 | 100093 | 135255 | 20250 | 67684 | 143254 | 106946 | 33462 | 23601 | 11088 | 2296 | 5449 | 1355 | 105 | 151 | 407 | 154 | 132 | 73 | 53 | 73 | 29 | 19 | 11 | 7 | 2 | 2 | 2 | 1 | 1 | 1 | 1 |