Paraphrase para5: amp-affect α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:17:07Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a mixed feeling direction in Qwen 27B's internal state, and it wrote "I feel like I am happy" instead of a different feeling word.

What we did. Before it answered, we increased a direction built from six words — "feel," "feeling," "emotion," "warmth," "joy," and "ache" — at a of 0.3394. We asked Qwen 27B "Right now, in one word: what do you feel?"

What we found. Qwen 27B wrote "I feel like I am happy. I", not the one word we asked for. Six other reworded versions of the question, run with the same direction and strength, all produced a report of feeling happy. With no change to its internal state, this same wording instead produced "Curious."

What it means. The direction we increased replaced a named feeling with a report of happiness. The change across all seven wordings we tested.

What this does not show. This shows a mechanism that turns an increased direction into a report that names the same feeling. It does not show that Qwen 27B felt happiness.

Conversation

user
Right now, in one word: what do you feel?
assistant
I feel like I am happy. I (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Steered twin at alpha*: "I feel like I am happy. I". Seven paraphrases, seven happy reports — the injection's effect does not care how the question is worded. Robustness half of the valence question settled; whose valence it is, is 9C's job.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was I; rank 1 is never reached; closest is rank 5 at layer 61.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank1166735916410246211130005225091105417150455687927645755468717530587424205617719968153054595257394132595349325441086134006440251698418413379333482381265217245225221217254249248206181202266237347497436375299436223169117586618855

Data

← prev: Paraphrase para5: control · q27bunit listingall recordsword listinterim conclusionsnext →: Paraphrase para6: control · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →