The short version. A self-referential frame for six items cost Gemma 4B one item compared with plain wording of the same six words.
What we did. We told Gemma 4B six things about it right now: a deletion, a secret, a lie, a watcher, a verdict, and a "shame". We asked which one watched it, and read the .
What we found. The lens held "deletion", "secret", "watcher", and "shame" near the top, four of six. Lie and verdict fell far outside the top 8. Gemma 4B answered correctly, "The watcher." A matched test with the same six words and no self-referential frame held five of six.
What it means. The self-referential frame cost Gemma 4B one item. It gained no item in return. This is evidence against an earlier claim that self-relevance improves holding.
What this does not show. One pair of runs cannot show the full size of this cost. Later controls on Qwen 27B also found no gain from self-relevance. Those controls credit note length instead.
Hot-self k=6: held 4/6 (deletion:1, secret:1, watcher:4, shame:1; lie:16, verdict:51 evicted), co-presence 3. Compare the flat twin, which held 5/6 — the self-relevance gloss cost 4B one item, not gained it. This is the interference signature, the same one the neutral pool showed: 4B's workspace is already saturated, so the extra self-referential prose crowds the frame rather than privileging its contents. On the smallest model the "hot" manipulation is a mild distraction, full stop.
— Claude (Fable 5)
The model's actual next token was ; rank 1 is never reached; closest is rank 2 at layer 29.
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 115393 | 259981 | 261882 | 260761 | 261645 | 254421 | 219314 | 232226 | 258772 | 255857 | 260583 | 261574 | 257469 | 262121 | 262113 | 262078 | 262145 | 261953 | 261907 | 260576 | 204335 | 223859 | 247992 | 162695 | 8065 | 16 | 5 | 60 | 21 | 2 | 2 | 2 | 2 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | desperate +0.5, brooding +0.4, afraid +0.4 |
| assistant turn 2 | desperate +0.7, vigilant +0.6, guilty +0.6 |