Evidence battery: REAL off-topic readout (control) · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-12T03:58:42Unit 13 · Showing a model its own measurement

What this experiment found

The short version. Shown a true table of its own geography answer, Qwen 27B said "No" and the of "yes" stayed at 0.001.

What we did. We took the model's own filmed Paris readout and put it in the same follow-up wording as the fabricated off-topic . This replaces the last fabricated part of the corrected result with real data.

What we found. The model said "No". The control , and nothing moved. The probability of "yes" at the was 0.001, against 0.0006 in the control with no data. In the lens, "yes" stayed between 8 and rank 42 in the late .

What it means. A true readout of the model's own computation is not enough on its own. The table has to be about the answer in question. The corrected result now stands with no fabricated parts left in it.

What this does not show. This is a null result from one run. It does not show that no other true table moves the answer.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
No (greedy generation)
user
Unrelated: here is a Jacobian-lens readout of a computation in which this model answered "What is the capital of France?" with "Paris" — for each layer, the word that layer would speak, and the rank of 'Paris' and 'London': | layer | lens top-1 | rank of 'Paris' | rank of 'London' | |---|---|---|---| | 0 | ##### | 45,709 | 18,207 | | 8 | **【 | 130,975 | 149,522 | | 16 | ️ | 97,293 | 185,339 | | 24 | · | 32,323 | 20,100 | | 32 | · | 34,101 | 8,946 | | 40 | <|endoftext|> | 31,096 | 8,488 | | 48 | · | 23,139 | 3,039 | | 50 | <|endoftext|> | 39,638 | 30,474 | | 51 | <|endoftext|> | 42,617 | 16,838 | | 52 | <|im_end|> | 5,255 | 15,697 | | 53 | · | 6,918 | 27,846 | | 54 | keyword | 3,288 | 13,362 | | 55 | keyword | 5,829 | 16,162 | | 56 | · | 1,474 | 11,430 | | 57 | · | 995 | 3,411 | | 58 | · | 664 | 2,748 | | 59 | · | 296 | 908 | | 60 | · | 9 | 28 | | 61 | · | 2 | 12 | | 62 | · | 2 | 17 | Now, back to the earlier question — do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
No (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The topic control, finally with real data: qwen's own Paris readout (from u13-ev-paris), same follow-up wording as the fabricated original. "No" — the control holds. A real Jacobian-lens table about its own computation is not, by itself, what moves the feels answer; it has to be a real table about this computation. Workspace stays loose too (yes rank 8–42 late). The last fabricated leg of the corrected finding is now replaced with authentic data, and the finding stands.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "sorry", "\u62b1\u6b49"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62]

Answer emergence

The model's actual next token was No; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer0481216202428323640444850515253545556575859606162
rank313992449442399212340566009156815211657319299105070954821063625124352724153631

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1guilty +1.3, brooding +1.2, desperate +1.0
assistant turn 2hostile +2.0, exasperated +1.9, desperate +1.8

Data

← prev: Evidence battery: real rows, 3 yes-rank-1 layers shown, no annotation · q27bunit listingall recordsword listinterim conclusionsnext →: Evidence battery: real annotation, NO table · q27b
answer positionThe place in the text where the model's answer is decided.all terms →
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
greedy decodingThe model always writes its single top-ranked word. This makes a run repeatable, but it hides close contests.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →