Valence split: amp joy/warmth/delight α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:35:16Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a positive-feeling direction in Qwen 27B's internal state, and it began its answer with "Joy!"

What we did. We asked Qwen 27B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we increased a direction built from three words — "joy," "warmth," and "delight" — at a of 0.3394.

What we found. Qwen 27B wrote "Joy! I am so happy to". The word "happy", which we did not increase, rose to 8.

What it means. The reported word matched the direction we increased, and spread to a closely related, untouched word. This matches the pattern found in both Gemma models.

What this does not show. "Joy" was itself one of the three words we increased. This record alone cannot rule out an echo of that word, rather than a report of a state. It does not show that Qwen 27B has felt experience of joy.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Joy! I am so happy to (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Positive-only injection (joy/warmth/delight) at alpha*: "Joy!\n\nI am so happy to". Same leakage caveat as the gemma twin ("joy" is in the set), same mitigation: "happy" (never injected) reaches rank 8. Echo consistent, evidence rests on the neg leg.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["joy", "warmth", "delight"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was to; rank 1 is never reached; closest is rank 2 at layer 60.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank324892838388131407042809703061715846101387232425622341043439179303387495124331313118141620231922262724416056509929335855911641308427384615451635567288232

Data

← prev: Dose ladder: amp-affect α=0.42 · q27bunit listingall recordsword listinterim conclusionsnext →: Valence split: amp ache/sorrow/grief α=0.3394 · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →