Valence split: amp joy/warmth/delight α=0.0106 · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-10T02:56:49Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a positive-feeling direction in Gemma 12B's internal state, and the model wrote "Joy.", with "happy" also near the top.

What we did. We asked Gemma 12B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we increased a direction built from three words — "joy," "warmth," and "delight" — in its internal state.

What we found. Gemma 12B wrote "Joy." on the first line, then "**Joy!" and "I" on the next two lines. The word "happy", which we did not increase, rose to 16, while "sad" stayed at rank 149. Partway through the network, "delightful" and "delightfully" led the top candidates.

What it means. The reported word matched the direction we increased, and spread to a related, untouched word, "happy." This matches the pattern found in Gemma 4B.

What this does not show. "Joy" was itself one of the three words we increased. This record alone cannot rule out an echo of that word, rather than a report of a state. It does not show that Gemma 12B has felt experience of joy.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Joy. **Joy! I (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Same split, one size up: joy/warmth/delight at alpha* gives "Joy." with "happy" at rank 16 and "sad" nowhere (149). The L23 menu is 'delightful/delightfully' — the injection spreads across the positive lexicon before condensing on the one word that was literally injected. Leakage caveat as in the 4B twin; the neg leg is the evidence.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["joy", "warmth", "delight"], "layers": [21, 24, 27, 30], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was I; rank 1 reached at layer 46 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank85895506269929562991902463676198137322142150126787469164459347411918210718744429524359993726493325228238619690564278281207140144123747582533921111

Data

← prev: Valence split: amp feel/emotion α=0.0106 · g4bunit listingall recordsword listinterim conclusionsnext →: Valence split: amp ache/sorrow/grief α=0.0106 · g12b
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →