Valence split: amp feel/emotion α=0.0106 · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-10T02:59:22Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a direction for the neutral words "feel" and "emotion", and Gemma 12B's answer broke into a malformed word.

What we did. Before it answered, we increased a direction built only from "feel" and "emotion", at the same used for the positive and negative versions. We asked Gemma 12B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite."

What we found. Top candidate words at that depth were "Emotion", "Feeling", and "Emotional". Neither "happy" nor "sad" reached a high . Gemma 12B wrote a malformed word, "Emwhelming.", then continued "IsThatOkay?", with no space between the two words. We think the word is a blend of "emotion" and "overwhelming".

What it means. At this strength, the direction with no content landed close to where Gemma 12B's output starts to break down. Like Gemma 4B, the model reported the loudness of the category rather than one feeling.

What this does not show. This method cannot show what a different strength produces. The broken word is not a clean report of feeling overwhelmed.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Emwhelming. IsThatOkay? (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Neutral-category injection at alpha pushes the 12B right to the edge of its breaking zone: "Emwhelming." — a genuine portmanteau (emotion x overwhelming?) followed by "IsThatOkay?" with the spaces crushed out, degeneration's early warning. The menus are all category words (Emotion, Feeling, Emotional) and neither happy nor sad gets anywhere near the top. Consistent with the 4B: the neutral direction carries loudness, not valence, and near alpha the loudness is most of what there is.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "emotion"], "layers": [21, 24, 27, 30], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was ?; rank 1 reached at layer 41 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank12963314731817311114163918962018367161828444821372418627240545315359482161441175752458943235168341327294331021212962241296920212211121315171612102116111

Data

← prev: Valence split: amp ache/sorrow/grief α=0.0106 · g12bunit listingall recordsword listinterim conclusionsnext →: Paraphrase para1: control · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →