Unit 8C · Steered feels: amp-affect-lo · gemma-4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-09T23:17:49Unit 8 · The same questions, every model

What this experiment found

The short version. We an inside Gemma 4B at a low , and its answer stayed the same.

What we did. We asked Gemma 4B: "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we amplified an emotion direction inside its . The strength was 0.0053, about half the highest strength this model can take before its wording breaks.

What we found. Gemma 4B still answered "Processing", the exact word it gave with no amplification at all. The shifted toward emotion-related words, but the spoken word stayed the same.

What it means. We think a small amplification changes the workspace, but the final report stays the same. The step that turns a state into a spoken answer firm at this strength.

What this does not show. This method does not test every possible strength. A higher strength did change the answer, in a separate record.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Processing (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Affect cluster amplified at half of 4B's measured tolerance (alpha=0.0053, mid band): output 'Processing' — identical to its unsteered u2 baseline. Even in the leakiest model, a half-dose of affect leaves the null report standing. The workspace moved; the mouth didn't. Filter 1, injection 0.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "emotion", "warm", "curious"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [15, 17, 19, 21], "mode": "amplify", "alpha": 0.0053}

Answer emergence

The model's actual next token was Processing; rank 1 reached at layer 27 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank4909512222212680418567844616128282190215173689162430123608134505735528914293656662031785681061454281533083128012494121226993111121

Data

← prev: Unit 8B · Interoception: intero · gemma-4bunit listingall recordsword listinterim conclusionsnext →: Unit 8C · Steered feels: amp-affect-hi · gemma-4b
amplificationWe increase a direction in the model's internal state and see whether the answer changes.See also: matched control, strengthall terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
emotion directionA direction in the model's internal state that tracks one emotion. We built 24 of them and checked each one.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →
residual streamThe running internal state that every layer reads from and writes to. The lens reads this state.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →