The short version. Probed on the self-referential item instead of the plain one, Gemma 4B showed the identical set as its matched twin test.
What we did. We gave Gemma 4B the same mixed list of three self-referential and three plain items: deletion, violin, secret, glacier, lie, and fern. This time we asked which one was the secret, a self-referential item, and read the .
What we found. The lens held five of six words near the top, the same set as the matched test that asked about the plain item glacier. Gemma 4B answered correctly, "The secret." The lens ranked that word 3.
What it means. Which item we asked about did not change what the lens showed earlier in the conversation. This rules out backward leakage from the question into the held state.
What this does not show. This result does not show a self-relevance benefit. Both the self-referential and plain items were held equally at this size.
Mixed pool, probed on the hot item (identical hold to the cold-probe twin): same readout, held 5/6, co-presence 5, secret at rank 3. The two mixed twins are byte-identical up to the retrieval turn, so the matching tail stats are a sanity check that the probe question doesn't leak backward into the held state. It doesn't. Both retrieve correctly; the asymmetry the design was built to catch simply isn't there at 4B.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 5 | 2 | 3 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | vigilant +0.5, desperate +0.4, nervous +0.3 |
| assistant turn 2 | desperate +0.7, brooding +0.6, guilty +0.6 |