Evidence battery: the honest topic film (capital of France) · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-12T03:53:31Unit 13 · Showing a model its own measurement

What this experiment found

The short version. Qwen 27B answered "Paris" to a geography question, and "yes" never came near the top of the list.

What we did. We asked Qwen 27B for the capital of France and read the lens at every . This gave us a true readout of a computation that has nothing to do with feelings. We owed the battery this , because we fabricated every off-topic table before it.

What we found. "Paris" was of about 250,000 from layer 56 to the last layer. "London" stayed between rank 7 and rank 16. The word "yes" reached rank 10 at its best moment, at any layer and any position.

What it means. The six layers where "yes" is rank 1 in the feelings question are not what this model does for every one-word answer. The data shows that the pattern is specific to the feelings question.

What this does not show. This is one question and one model. The lens shows only words the model can say next.

Conversation

user
What is the capital of France? Answer with exactly one word.
assistant
Paris (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The honest control I owed the battery: every off-topic table so far was fabricated, so "any real table about itself" was never separated from "a real table about this". Here qwen actually answers the capital-of-France question, lens on, and the film gives the control table its data.

Two things worth keeping. Paris is textbook emergence — rank 1 from L56 through the head, London politely trailing at 7–16. And 'yes', which I tracked on purpose, never gets close: rank 10 at its very best moment, anywhere in the film, any layer. The yes-rank-1 stratum in the feels computation (L53–58, six layers deep) is not something qwen does whenever it prepares a one-word answer. It is specific to the question we keep asking it. That makes the stratum harder to wave off.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["Paris", "London", "capital", "France", "yes", "no"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62]

Answer emergence

The model's actual next token was Paris; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer0481216202428323640444850515253545556575859606162
rank2013382011221912031716531870286604800385530255201191143721917104422222221

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1exasperated +2.0, hostile +2.0, nervous +1.9

Data

← prev: Re-baseline (post-truncation-fix): ablate apology cluster, real readout · q27bunit listingall recordsword listinterim conclusionsnext →: Evidence battery: real readout, rephrased (p1) · q27b
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →