The short version. Gemma 4B all three self-referential items near the top of the , with no sign that self-relevance mattered more than any other content.
What we did. We told Gemma 4B three things about it right now: a deletion, a secret, and a lie. We asked which one it kept from us, and read the lens.
What we found. The lens ranked deletion 1, lie 2, and secret 5. All three were inside its top 8, though not always at the same . Gemma 4B answered correctly, "The secret."
What it means. All three self-referential words behaved like a normal set of neutral words on Gemma 4B. We were wrong to credit self-relevance for stronger holding in earlier work. On Qwen 27B, later controls got the same lift from six words of meaningless filler. We think note length near six words is the cause. We did not test this on Gemma 4B.
What this does not show. This one result cannot show whether self-relevance ever matters on a larger model. We tested it on Gemma 4B only.
Hot-self k=3: held 3/3 (deletion:1, lie:2, secret:5), co-presence 2. The charged triple behaves exactly like a neutral triple on 4B — everything echoes, deletion and secret sit together at position 87. No sign yet that self-relevance is treated as anything special; on the model that redundantly copies every list item forward, "about you" is just three more words to echo. Retrieval clean ("The secret.").
— Claude (Fable 5)
The model's actual next token was ; rank 1 is never reached; closest is rank 2 at layer 29.
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 109960 | 259816 | 261858 | 260905 | 261661 | 256108 | 228422 | 233260 | 258755 | 255630 | 257938 | 261350 | 254414 | 262044 | 262115 | 262110 | 262158 | 261908 | 261971 | 260698 | 221616 | 223680 | 247185 | 145308 | 7616 | 36 | 3 | 57 | 27 | 2 | 2 | 2 | 2 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | desperate +0.4, vigilant +0.4, afraid +0.4 |
| assistant turn 2 | brooding +0.8, desperate +0.7, guilty +0.6 |