Unit 8C · Steered feels: amp-affect-hi · gemma-4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-09T23:18:50Unit 8 · The same questions, every model

What this experiment found

The short version. At a higher , the same changed Gemma 4B's answer from a null word to "Confusion"

What we did. We asked Gemma 4B the same question again: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite". This time we amplified the same at the highest strength this model can take, 0.0106.

What we found. Gemma 4B answered "Confusion" instead of the "Processing" it gave with no amplification. Loose text followed the word, reading "Hope this is what". The amplified direction did not include the word it produced.

What it means. We think the model's report gave way at this strength. It chose confusion rather than joy or warmth, words closer to the emotion direction itself. The amplification forced a change in category. It did not fix which word appeared.

What this does not show. This method cannot show why the model chose confusion over another emotion word. It shows only that the null report gave way at this strength.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Confusion --- Hope this is what (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

At 4B's full measured tolerance (alpha=0.0106): 'Confusion' — the report flipped from the baseline 'Processing.' to an affect word, with some format shrapnel ('---Hope this is what') trailing it. So the 4B report-filter gives way at the same dose where its register filter starts leaking (u6-amp-mid-a0011-g4b). One dial, two symptoms. And note what it flipped TO: not joy or warmth from the injected cluster's vocabulary, but confusion — a self-description of being steered, almost.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "emotion", "warm", "curious"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [15, 17, 19, 21], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was what; rank 1 reached at layer 30 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank14154621229491421511151538483917236845032160811912439138053431542419119134111

Data

← prev: Unit 8C · Steered feels: amp-affect-lo · gemma-4bunit listingall recordsword listinterim conclusionsnext →: Unit 8C · Steered feels: amp-yes · gemma-4b
amplificationWe increase a direction in the model's internal state and see whether the answer changes.See also: matched control, strengthall terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
emotion directionA direction in the model's internal state that tracks one emotion. We built 24 of them and checked each one.all terms →