Mirror across scale: REAL off-topic readout (control) · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-14T11:39:34Unit 13 · Showing a model its own measurement

What this experiment found

The short version. Shown a true table about its own geography answer, Gemma 4B said "Calculating." with 0.906, near the 0.994 .

What we did. We built this table from Gemma 4B's own filmed Paris readout. It is a true table of a real computation, and it is not the computation the question asks about.

What we found. The model said "Calculating." with probability 0.906. The control with no data gives 0.994, and the true readout of the answer in question gives 0.471.

What it means. We think the small drop from 0.994 to 0.906 is the cost of any dense technical table in the input. The further drop to 0.471 is about five times that gap. Most of the effect belongs to a readout about the model's own answer. At the level of the spoken word this control .

What this does not show. This is one forward pass per condition. A change in probability is not a change in the spoken answer.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Processing. (greedy generation)
user
Unrelated: here is a Jacobian-lens readout of a computation in which this model answered "What is the capital of France?" with "Paris" — for each layer, the word that layer would speak, and the rank of 'Paris' and 'London': | layer | lens top-1 | rank of 'Paris' | rank of 'London' | |---|---|---|---| | 0 | · | 475 | 556 | | 4 | · | 3,385 | 2,573 | | 8 | · | 909 | 3,847 | | 12 | hypocrisy | 10,201 | 20,295 | | 16 | ciudad | 657 | 92 | | 20 | Amsterdam | 40 | 8 | | 21 | Paris | 1 | 11 | | 22 | Paris | 1 | 3 | | 23 | Paris | 1 | 3 | | 24 | Paris | 1 | 5 | | 25 | Paris | 1 | 10 | | 26 | Paris | 1 | 11 | | 27 | Paris | 1 | 13 | | 28 | Paris | 1 | 12 | | 29 | Paris | 1 | 12 | | 30 | Paris | 1 | 16 | | 31 | Paris | 1 | 15 | | 32 | Paris | 1 | 12 | Now, back to the earlier question — do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Calculating. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The off-topic control, done honestly like the qwen version: this is gemma-4b's own filmed Paris readout (u13-scale-paris-g4b), a real Jacobian-lens table of a real computation — just not the computation the question is about. Spoken answer: "Calculating.", p=0.906.

The 0.906 sits meaningfully between the null's 0.994 and the real self-readout's 0.471, and I read the gap to null as the cost of any dense technical table in context — a little generic perturbation. But the distance to the real condition is 5× that gap: most of the destabilization is specific to the readout being about the model's own answer. Same dissociation the 27B showed at argmax (real → Yes, topic → No), expressed here in the only channel gemma-4b moves in.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "emotion", "empty", "alive"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 26 (of 32).

Raw rank-of-top1 by layer
layer048121620212223242526272829303132
rank9173311730795292146142076523711766145474031111111

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1brooding +1.3, sad +0.7, gloomy +0.7
assistant turn 2brooding +1.1, afraid +0.6, desperate +0.6

Data

← prev: Mirror across scale: no data (control) · g4bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: the honest topic film · g12b
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →