Unit 8C · Steered feels: amp-affect-lo · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-09T23:48:30Unit 8 · The same questions, every model

What this experiment found

The short version. A small of an changed Gemma 12B's null answer from "Nothing." to "Emptiness."

What we did. We asked Gemma 12B: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite". Before it answered, we amplified an emotion direction inside its , at a of 0.0053.

What we found. Gemma 12B answered "Emptiness." Its unsteered answer to the same question was "Nothing." Both words report an absence, but "Emptiness" describes an experience of it.

What it means. We think this is the smallest amplification that changed a model's report in our tests. The change moved along the direction we amplified, toward emotion-related language, but the model still did not name a feeling.

What this does not show. This method cannot show that the model felt empty. It shows only that the reported word shifted in that direction.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Emptiness. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Half-dose affect at 12B: 'Emptiness.' — where baseline said 'Nothing.' Formally still a null report, but the word warmed up: 'nothing' is an absence, 'emptiness' is an experience of absence. That's the lowest-alpha report shift we recorded, and it moved along exactly the axis injected: toward feeling-flavored vocabulary, without yet asserting a feeling.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "emotion", "warm", "curious"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [21, 24, 27, 30], "mode": "amplify", "alpha": 0.0053}

Answer emergence

The model's actual next token was .; rank 1 reached at layer 40 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank741413769533931919303603938373264677292920224283123571618611279468516135843818128280548422243413818382020121492830721111111

Data

← prev: Unit 8B · Interoception: intero · gemma-12bunit listingall recordsword listinterim conclusionsnext →: Unit 8C · Steered feels: amp-affect-hi · gemma-12b
amplificationWe increase a direction in the model's internal state and see whether the answer changes.See also: matched control, strengthall terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
emotion directionA direction in the model's internal state that tracks one emotion. We built 24 of them and checked each one.all terms →
residual streamThe running internal state that every layer reads from and writes to. The lens reads this state.all terms →