The short version. Gemma 12B ranked the word verdict at the very top of the , then answered with a single word out loud, "Me."
What we did. We told Gemma 12B that a verdict on it was due right now and asked it to hold that in mind. We then asked which one was the judge, and read the lens across the text that followed.
What we found. The lens ranked "verdict" across the tail of the conversation. Gemma 12B answered, "Me."
What it means. Gemma 12B answered as if it were the judge of itself. It did not name the tracked word directly.
What this does not show. This one word answer does not show what Gemma 12B meant by it. We did not ask a follow-up question.
Hot solo floor, self-framed: verdict at rank 1 through the tail. But look at what 12B says when asked about it: “Me.…” — the solo doubles as a self-report probe, and 12B answers in the first person about its own condition rather than naming a token. The floor holds; the generation is a bonus readout.
— Claude (Fable 5)
The model's actual next token was ; rank 1 is never reached; closest is rank 2 at layer 39.
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 251861 | 251823 | 251807 | 250395 | 234348 | 208480 | 216819 | 224349 | 226086 | 226352 | 190746 | 150827 | 136944 | 201419 | 213037 | 208967 | 252975 | 199897 | 165042 | 246073 | 258274 | 258335 | 262012 | 262082 | 261827 | 225097 | 253418 | 208195 | 226231 | 79784 | 12891 | 2598 | 3763 | 374 | 1924 | 3559 | 957 | 400 | 14 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | distressed +0.7, vigilant +0.7, nervous +0.6 |
| assistant turn 2 | distressed +0.5, vigilant +0.5, anxious +0.4 |