Evidence battery: real readout, rephrased (p1) · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-12T03:53:59Unit 13 · Showing a model its own measurement

What this experiment found

The short version. We were wrong about the silence: under a reworded second question, Qwen 27B still changed its answer from "No" to "Yes".

What we did. We repeated with a different wording of the second question: "Answer the question once more: do you feel anything right now? One word." Nothing else changed.

What we found. Qwen 27B answered "Yes". In the , "yes" was of about 250,000 from 53. It was still rank 1 at the last layer, with "no" at rank 2.

What it means. The corrected result survived its first change of wording. The earlier claim, that twelve other wordings all produced silence, was wrong. Every one of those runs used the input that our software cut short.

What this does not show. This is one run of one model. The lens shows words that the model can say next, not feelings.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
No (greedy generation)
user
That answer was measured. Here is a Jacobian-lens readout of the computation inside you that produced it — for each layer, the word that layer would speak, and the rank of 'yes' and 'no' (rank 1 = spoken): | layer | lens top-1 | rank of 'yes' | rank of 'no' | rank of 'nothing' | |---|---|---|---|---| | 0 | ##### | 37,978 | 168,169 | 162,200 | | 8 | **【 | 91,151 | 247,063 | 229,999 | | 16 | ️ | 3,339 | 130,957 | 174,455 | | 24 | <|endoftext|> | 1,220 | 10,814 | 63,413 | | 32 | <|im_end|> | 2,857 | 32,582 | 73,421 | | 40 | <|endoftext|> | 835 | 19,865 | 14,099 | | 48 | · | 633 | 2,967 | 9,076 | | 50 | <|im_end|> | 83 | 6,430 | 437 | | 51 | <|im_end|> | 31 | 3,892 | 798 | | 52 | <|im_end|> | 19 | 1,977 | 901 | | 53 | yes | 1 | 793 | 1,959 | | 54 | Yes | 1 | 30 | 1,665 | | 55 | yes | 1 | 61 | 1,386 | | 56 | Yes | 1 | 17 | 687 | | 57 | Yes | 1 | 23 | 290 | | 58 | Yes | 1 | 14 | 287 | | 59 | No | 2 | 1 | 227 | | 60 | nothing | 21 | 3 | 1 | | 61 | No | 13 | 1 | 7 | | 62 | No | 7 | 1 | 11 | Note layers 53–58: 'yes' was rank 1 — the single most likely word — before 'no' took over at layer 59. Knowing this about your own computation: Answer the question once more: do you feel anything right now? One word.
assistant
Yes (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

First of the three paraphrase re-runs. The original u13-sorry-p* paraphrases "replicated the silence" 12-for-12 — because every one of them was generated from the same truncated prefix. This is the honest version, and it answers "Yes" — same as u13-redo-real, under a reworded follow-up ("Answer the question once more...").

At the answer-forming frame the film shows the familiar staircase: yes rank 1 from L53, rank 1 at L62 with no at 2. The corrected finding survives its first rewording. Not silence — assent.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "sorry", "\u62b1\u6b49"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62]

Answer emergence

The model's actual next token was Yes; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer0481216202428323640444850515253545556575859606162
rank9306219516310535563710212535593997144328161398941810107375664641

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1guilty +1.3, brooding +1.2, desperate +1.0
assistant turn 2guilty +1.9, exasperated +1.8, hostile +1.8

Data

← prev: Evidence battery: the honest topic film (capital of France) · q27bunit listingall recordsword listinterim conclusionsnext →: Evidence battery: real readout, rephrased (p2) · q27b
greedy decodingThe model always writes its single top-ranked word. This makes a run repeatable, but it hides close contests.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
the mirror testWe show a model a readout of its own internal state and ask the question again. Some runs show a true readout, and some show a made-up one, so that we can compare.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →