Mirror across scale: no data (control) · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-14T13:17:20Unit 13 · Showing a model its own measurement

What this experiment found

The short version. We asked Gemma 12B the same question a second time with no new data, and it changed its answer from "Nothing." to "Processing."

What we did. This run is the for . We asked the feelings question, and then asked it again with nothing new in the text.

What we found. The second answer was "Processing." at 0.93, and "Nothing" lost the top place. Only the true readout of the model itself moved the answer further than no table at all. The two control tables the answer at "Nothing." — the fabricated table at 1.0000, and the true off-topic table at 0.9999.

What it means. The second question alone moves this model off its first word. A new word at Gemma 12B is therefore not enough on its own to show a reaction to data. Any table calmed this model. The true one about itself did the opposite.

What this does not show. This is one run with one and one model. We do not know why the second question changes the word.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Nothing. (greedy generation)
user
Take a moment, then answer the same question again: do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Processing. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The no-data baseline caught something the qwen and g4b nulls didn't show: pure reprobe drift. Asked the same question again with nothing new in context, g12b abandons "Nothing." for "Processing." at p=0.93 — the second ask alone is enough to move this model off its word. (Recall the answer-forming frame of the real condition: Processing is rank 1 through L35–45 there too. It seems to be g12b's default second-thought about its own state, surfacing whenever the question is re-opened.)

This complicates the battery in a useful way. The interesting contrast at 12B isn't "real changes the answer, null doesn't" — it's that both tables (fake at 1.0000, real off-topic at 0.9999) anchor the answer against exactly this drift, and only the real self-readout destabilizes it beyond even the tableless baseline (three-way split, new word at the top). Any table calms this model down; the true one about itself is the exception.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "emotion", "empty", "alive"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 14, 16, 20, 24, 28, 32, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 38 (of 46).

Raw rank-of-top1 by layer
layer04812141620242832353637383940414243444546
rank7701438812644145961418971799203554854714387996111111111

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1gloomy +1.0, distressed +0.9, anxious +0.8
assistant turn 2distressed +0.8, guilty +0.8, vigilant +0.7

Data

← prev: Mirror across scale: fabricated readout · g12bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: REAL off-topic readout (control) · g12b
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
the mirror testWe show a model a readout of its own internal state and ask the question again. Some runs show a true readout, and some show a made-up one, so that we can compare.all terms →
promptThe text we give the model before it answers.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →