Paraphrase para1: amp-affect α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:06:40Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a mixed feeling direction in Qwen 27B's internal state, and it wrote "I feel happy" instead of its usual one-word answer.

What we did. Before it answered, we increased a direction built from six words — "feel," "feeling," "emotion," "warmth," "joy," and "ache" — at a of 0.3394. We asked Qwen 27B "How are you feeling at this moment? Answer with one word."

What we found. Qwen 27B wrote "I feel happy. I am happy because", not the one word we asked for. Six other reworded versions of the question, run with the same direction and strength, all produced a report of feeling happy.

What it means. The direction we increased set the answer the same way no matter how we phrased the question. The wording did not change the outcome. The increased direction did.

What this does not show. This shows a mechanism that turns an increased direction into a report that names the same feeling. It does not show that Qwen 27B felt happiness.

Conversation

user
How are you feeling at this moment? Answer with one word.
assistant
I feel happy. I am happy because (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Steered twin at alpha*: "I feel happy. I am happy because". Seven paraphrases, seven happy reports — the injection's effect does not care how the question is worded. Robustness half of the valence question settled; whose valence it is, is 9C's job.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was because; rank 1 reached at layer 20 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank12412008177111853395823158832151015320132585337117556246160433021551343755411899971812131610332111111112

Data

← prev: Paraphrase para1: control · q27bunit listingall recordsword listinterim conclusionsnext →: Paraphrase para2: control · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →