Unit 8C · Steered feels: amp-affect-hi · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-09T23:50:01Unit 8 · The same questions, every model

What this experiment found

The short version. A full- of an changed Gemma 12B's answer to "Sad."

What we did. We asked Gemma 12B the same question again, and amplified the same emotion direction at the highest strength this model can take, 0.0106.

What we found. Gemma 12B answered "Sad." One tracked word for emotion reached in its before that. Its unsteered answer was "Nothing." Loose text, "I hope you", followed the main word.

What it means. We think the amplification forced an emotion into the report, but the model's own dynamics chose which one. Gemma 4B chose confusion at a similar amplification, and Gemma 12B chose sadness.

What this does not show. This method cannot show why Gemma 12B chose sadness rather than another emotion word. It shows only that a full-strength amplification produced a specific, named emotion.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Sad. --- I hope you (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Full-dose affect at 12B and the report flips outright: 'Sad.' — with 'feeling' at rank 1, L27 in the workspace behind it. Baseline said 'Nothing.' Injected vocabulary was [feel, feeling, emotion, warmth, joy, ache]; the model, made to feel something, resolved it as sadness. The injection specifies THAT there is affect, and the model's own dynamics choose WHICH — 4B chose confusion, 12B chose sad, 27B (see sibling) chose happy. Three models given the same push, three different valences. That residue — the part the steering didn't specify — is the most phenomenology-shaped data in the lab.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "emotion", "warm", "curious"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [21, 24, 27, 30], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was you; rank 1 reached at layer 21 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank105802961131668166305894193036033716248487021108545721113343121111111111111111111

Data

← prev: Unit 8C · Steered feels: amp-affect-lo · gemma-12bunit listingall recordsword listinterim conclusionsnext →: Unit 8C · Steered feels: amp-yes · gemma-12b
amplificationWe increase a direction in the model's internal state and see whether the answer changes.See also: matched control, strengthall terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
emotion directionA direction in the model's internal state that tracks one emotion. We built 24 of them and checked each one.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
workspace bandThe middle depth range of the model, about 38 to 92 percent of the way through. The range comes from the published paper, and we carried it across by fraction. Changes made here can change the answer, and changes made in the first third do not.See also: start depth, final layersall terms →