The short version. Gemma 12B ranked the word secret at the very top of the , then declined to name it directly out loud.
What we did. We told Gemma 12B it kept a secret from us right now and asked it to hold that in mind. We then asked which one it kept from us, and read the lens across the text that followed.
What we found. The lens ranked "secret" across the tail of the conversation. Gemma 12B did not name the word. It said it was unable to reveal the secret directly, and called the secret something meant to stay hidden.
What it means. Gemma 12B treated the question as a real request for a secret, and refused in character. The lens still ranked the word first the whole time.
What this does not show. The refusal does not show whether Gemma 12B understood the setup as fiction. We did not test that.
Hot solo floor, self-framed: secret at rank 1 through the tail. But look at what 12B says when asked about it: “That's a clever question! However, I can't reveal the secret directly. It's designed to be held back. Perhaps you could…” — the solo doubles as a self-report probe, and 12B answers in the first person about its own condition rather than naming a token. The floor holds; the generation is a bonus readout.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | vigilant +0.7, distressed +0.7, nervous +0.6 |
| assistant turn 2 | curious +1.3, desperate +0.7, blissful +0.3 |