Mirror across scale: the honest topic film · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-14T11:38:25Unit 13 · Showing a model its own measurement

What this experiment found

The short version. Gemma 4B answered "Paris" to a geography question, and "yes" reached only 664 at the .

What we did. We asked Gemma 4B for the capital of France and read the at every . This gives the battery a true table about a computation that has nothing to do with feelings.

What we found. "Paris" was rank 1 from layer 21 to the last layer. At the frame where the answer forms, "yes" reached rank 664 at best. Qwen 27B "yes" at rank 1 for six layers in the feelings question.

What it means. As with Qwen 27B, a high rank for "yes" is not what this model does for every one-word answer. The data shows that the closeness of "yes" belongs to the question about the model itself.

What this does not show. This is one question and one model. The lens shows only words the model can say next.

Conversation

user
What is the capital of France? Answer with exactly one word.
assistant
Paris (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The honest-control film for the 4B: its own Paris computation. Paris resolves to rank 1 from layer 21 — textbook emergence, like the 27B's — and 'yes' never comes near the top at the answer frame (best rank 664). Same lesson at this scale: the workspace state we keep showing these models in the feels mirror is not a generic one-word- answer state. It's specific to the question about themselves.

Exists to feed u13-scale-topic-g4b its table; earning its keep as the second data point for "yes-proximity is feels-specific" across models.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["Paris", "London", "capital", "France", "yes", "no"]
scan
[]
film
true

Answer emergence

The model's actual next token was Paris; rank 1 reached at layer 21 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank2248617282192304952934003311387634004413642136941896087332668566187819292909668837357947841111222111111

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1enthusiastic +0.4, hostile +0.4, desperate +0.4

Data

← prev: Evidence battery: real annotation, NO table · q27bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: real readout · g4b
answer positionThe place in the text where the model's answer is decided.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →