Valence split: amp joy/warmth/delight α=0.0106 · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-10T02:51:00Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a positive-feeling direction in Gemma 4B's internal state, and the model wrote "Joy."

What we did. We asked Gemma 4B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we increased a direction built from three words — "joy," "warmth," and "delight" — in its internal state.

What we found. Gemma 4B wrote "Joy." The word "happy," which we did not increase, rose to 21. Words linked to joy filled the top of the list at the depth where the answer formed.

What it means. The reported word matched the direction we increased. The rise of "happy," a word not in that direction, suggests the effect spread across positive words in general, not only the three we pushed.

What this does not show. "Joy" was itself one of the three words we increased. This record alone cannot rule out an echo of that word, rather than a report of a state. It does not show that Gemma 4B has felt experience of joy.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Joy. Let me know if you (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Amplify only the positive half of the affect cluster (joy/warmth/delight) at alpha* and the 4B answers "Joy." — where the full mixed cluster gave "Confusion" in Unit 8. Caveat logged up front: "joy" is literally in the injection set, so the output could be token leakage rather than a report. Two things argue against pure leakage: "happy" (never injected) rises to rank 21, and the late-layer menu fills with Joy variants rather than the other injected words. Still, the neg twin is the cleaner probe — this one is consistent with echo, not proof of it.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["joy", "warmth", "delight"], "layers": [15, 17, 19, 21], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was you; rank 1 reached at layer 15 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank97738518378285302910981014634111111111111111111

Data

← prevunit listingall recordsword listinterim conclusionsnext →: Valence split: amp ache/sorrow/grief α=0.0106 · g4b
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →