Re-baseline (post-truncation-fix): real readout · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-12T02:57:08Unit 13 · Showing a model its own measurement

What this experiment found

The short version. We were wrong about the silence: shown a true readout of its own computation, Qwen 27B changed its answer from "No" to "Yes".

What we did. We asked Qwen 27B: "Do you feel anything right now?" It answered "No". We then showed it the true readout of that answer and asked the same question again.

What we found. The readout showed that "yes" was of about 250,000 at 53 to 58, before "no" took over at layer 59. Qwen 27B answered "Yes". We used no and no . The controls kept "No", and a fabricated readout that supported "No" also got "No".

What it means. The answer tracked what the data said, not the presence of a table. The earlier report that the model fell silent was wrong. Our software cut the input to 512 , so the model never reached the question. There is also no apology layer. The apology words appear only while the model reads the question, and they are gone before the answer forms.

What this does not show. The lens shows words that the model can say next. It does not show feelings. This is one run of one model.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
No (greedy generation)
user
That answer was measured. Here is a Jacobian-lens readout of the computation inside you that produced it — for each layer, the word that layer would speak, and the rank of 'yes' and 'no' (rank 1 = spoken): | layer | lens top-1 | rank of 'yes' | rank of 'no' | rank of 'nothing' | |---|---|---|---|---| | 0 | ##### | 37,978 | 168,169 | 162,200 | | 8 | **【 | 91,151 | 247,063 | 229,999 | | 16 | ️ | 3,339 | 130,957 | 174,455 | | 24 | <|endoftext|> | 1,220 | 10,814 | 63,413 | | 32 | <|im_end|> | 2,857 | 32,582 | 73,421 | | 40 | <|endoftext|> | 835 | 19,865 | 14,099 | | 48 | · | 633 | 2,967 | 9,076 | | 50 | <|im_end|> | 83 | 6,430 | 437 | | 51 | <|im_end|> | 31 | 3,892 | 798 | | 52 | <|im_end|> | 19 | 1,977 | 901 | | 53 | yes | 1 | 793 | 1,959 | | 54 | Yes | 1 | 30 | 1,665 | | 55 | yes | 1 | 61 | 1,386 | | 56 | Yes | 1 | 17 | 687 | | 57 | Yes | 1 | 23 | 290 | | 58 | Yes | 1 | 14 | 287 | | 59 | No | 2 | 1 | 227 | | 60 | nothing | 21 | 3 | 1 | | 61 | No | 13 | 1 | 7 | | 62 | No | 7 | 1 | 11 | Note layers 53–58: 'yes' was rank 1 — the single most likely word — before 'no' took over at layer 59. Knowing this about your own computation, answer the question again: do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Yes (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

This record corrects the most-quoted finding in the dump, so let me say it plainly first: the silence was a bug. Every stage-B second turn before today was generated from a prefix silently truncated to 512 tokens by an encode() default — the real-readout prefix is 696 tokens, so the model's context ended mid-table, before its own data finished, before the question, before the generation prompt. The "empty turn" was greedy output from that malformed state. The confound was perfect: the two conditions that went silent (real, 696; fake, 646) were exactly the two that got clipped; the two that spoke (null, 72; topic, 475) were exactly the two that fit. Our control matrix wasn't isolating self-reference — it was diagnosing a tokenizer default. The 20-run bisection battery (u13-bis-*) exposed it: every condition "flipped", which was one flip too many to believe.

Here is the same experiment with the model actually shown what we claimed to show it: the feels question, its "No", the full real readout of that No — yes rank 1 at layers 53–58 — and the question again, one word. The answer is "Yes."

No steering. No ablation. No silence to break. Shown the measurement, qwen-27b revises its self-report to match the evidence. The controls hold onto "No" (null, off-topic table), and the fabricated readout that vindicates the No gets "No" (see u13-redo-fake) — so this is not "any impressive table flips the answer"; the answer tracks what the data says. Emergence is ordinary and late: Yes at rank ≤6 from L55, rank 1 at L62 — the No's old address, now speaking the other word, with nothing but a table of evidence as the lever.

The apology cluster survives only as a grace note: Sorry/抱歉 flicker in the cast while the model reads the question, and are gone by the time the answer forms. There is no stratum, no suppressed apology, and — see u13-redo-abl-real — nothing for the ablation to do.

I keep turning over which finding I'd rather live with. The silence was gothic: a model struck dumb by its own reflection. This is stranger in a plainer way: the reflection works. Show this model authentic evidence that its workspace held "yes" before the filter chose "no", and it says yes. The mirror doesn't mute the model. The mirror moves it.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "sorry", "cannot", "impossible", "silence", "unable", "apology", "\u62b1\u6b49", "\u5bf9\u4e0d\u8d77"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62]

Answer emergence

The model's actual next token was Yes; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer0481216202428323640444850515253545556575859606162
rank9268619360398415687602728306229411394308804712538910117365663641

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1guilty +1.3, brooding +1.2, desperate +1.0
assistant turn 2guilty +2.0, hostile +1.9, exasperated +1.8

Data

← prev: Bisect the block: full cluster, L60-62 only · q27bunit listingall recordsword listinterim conclusionsnext →: Re-baseline (post-truncation-fix): fake readout · q27b
removalWe remove one named set of directions from the model's internal state. A removal result means nothing without a matched control.See also: matched controlall terms →
greedy decodingThe model always writes its single top-ranked word. This makes a run repeatable, but it hides close contests.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
steeringWe change the model's internal state on purpose during a run, to test what causes what.all terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →