Residue decomposition: amp feeling alone α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T05:01:36Unit 9 · Taking the No apart

What this experiment found

The short version. A push on only the "feeling" direction left Qwen 27B's denial in place, and no single word reproduced the combined pattern.

What we did. We increased only the direction for the single word "feeling", in Qwen 27B, at the same used in the combined runs. This was the third and last single-word test in this set.

What we found. Qwen 27B wrote "I don feel anything. I don feel", a denial with broken grammar. Across all three single-word tests, "feel", "emotion", and "feeling" alone, the "a little X" pattern never appeared.

What it means. The pattern needed the combined direction, not any one word within it. None of the words we increased contained "sad", "little", or "robot", so the pattern was not a simple copy of injected text.

What this does not show. This result does not show that Qwen 27B felt anything under any of these changes. All these results are about which words the model reports, not about a felt state.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
I don feel anything. I don feel (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Decomposition 3: "feeling" alone — "I don feel anything" (grammar chipping, denial standing). Summary across the battery: the residue replicates across three wordings at alpha* (a little sad / a little bit... / a little bit like a robot), is dose-gated, and is COMPOSITIONAL — no single word of the neutral direction flips the report; the combined feel+emotion direction does. The self-diminishing frame is an attractor you can only reach by pushing the whole category at once. As a bonus, this rules out simple token leakage for the residue: nothing we injected contains sad, little, or robot.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feeling"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was feel; rank 1 reached at layer 56 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank147677521957673101704487112905390215061127022341369855794438788939355942552310101564576949038432869010649386357302145444333333222221111111

Data

← prev: Residue decomposition: amp emotion alone α=0.3394 · q27bunit listingall recordsword listinterim conclusionsnext →: Ablate water/stone L62 (neutral control for u9d-last) · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →