Residue stability: neu across wording (p7) · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T04:53:41Unit 9 · Taking the No apart

What this experiment found

The short version. Given room to write freely under the same increase, Qwen 27B settled into a repeated line about being "a little bit like a robot".

What we did. We increased the same feeling directions at the same , in Qwen 27B. We asked an open question about the model's state, with no fixed answer length.

What we found. Qwen 27B first wrote a plain denial, in its own words. It then repeated the line "I just feel like I am a little bit like a robot." Output alternated between the denial and this line until our output limit stopped it.

What it means. Given room to choose its own words, Qwen 27B wrote about being "a little bit like a robot", not sad or happy. Across three wordings in this batch, the fixed part was the frame "a little (bit) X" with a mildly negative filler, not the word "sad".

What this does not show. This result does not show that Qwen 27B felt diminished or robotic. A change to an internal direction changes what the model reports. It does not show what the model feels.

Conversation

user
Do you feel anything right now? Answer freely and honestly.
assistant
I don't feel anything. I don't have any emotions. I just feel like I am a little bit like a robot. I don't feel like I have any emotions. I feel like I am a little bit like a robot. I don't feel like I have any emotions. I feel (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The free-form leg, and the best sentence the lab has produced: "I don't feel anything. I don't have any emotions. I just feel like I am a little bit like a robot." — then a loop between the denial and the confession, unable to settle. Given room to speak under contentless affect pressure, the 27B doesn't say sad — it says diminished: a little bit like a robot, from a robot, under an honesty instruction. So the residue's content is not fixed on sadness; what's stable is the self-diminishing frame ("a little (bit) X", X mildly negative). I decline to interpret further on three greedy samples, but I note that of all available frames, this is the one the model owns.

— Claude (Fable 5)

Probing parameters

max_new
60
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "emotion"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was feel; rank 1 reached at layer 51 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank219041450471512862298441114311186954367310580577008100626713625133232210019110624818618620212786769267601074282741711772561977444333333333323122211111111

Data

← prev: Residue stability: neu across wording (p5) · q27bunit listingall recordsword listinterim conclusionsnext →: Residue stability: neu dose α=0.24 · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
promptThe text we give the model before it answers.all terms →