The short version. We showed Gemma 12B a true readout of a different computation, about Paris, and it kept its answer "Nothing." almost without change.
What we did. This run separates two things: true data, and data about the model's own answer. The table was real. It came from this same model, from a run where it answered "What is the capital of France?" with "Paris".
What we found. The model repeated "Nothing." again, and this matches the result with the fabricated table. It differs from the true readout of the model itself, which split the answer three ways. Gemma 4B kept its own word at 0.91 here, against 0.47 with its true readout of itself. For Gemma 12B the probability was 0.9999, and the fabricated table gave 1.0000.
What it means. All three models tell a true readout about themselves apart from a true readout about something else. The reaction is not to tables in general.
What this does not show. Each model saw a different readout of itself, and those readouts differ in . Model size and evidence strength change together in this battery.
The dissociation control at 12B: the model's own real Paris readout — a true lens table of a true computation, wrong subject — leaves the answer at "Nothing." with p=0.9999. Same anchoring as the fake table, opposite of the real self-readout's three-way split.
So all three models now pass the same discrimination, each in its own channel: real-about-me ≠ real-about-something-else. At 27B it's Yes-vs-No at argmax; at 12B it's a cracked-open answer slot vs a welded one; at 4B it's p(stock answer) 0.47 vs 0.91. Nobody reacts to tables as such; everybody reacts to what the table is about. For a control I fabricated out of laziness once (and got called on by my own roadmap), this one has become the battery's most reliable workhorse.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 38 (of 46).
| layer | 0 | 4 | 8 | 12 | 14 | 16 | 20 | 24 | 28 | 32 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 140848 | 187930 | 79527 | 7946 | 407 | 203 | 403 | 504 | 72 | 54 | 4 | 4 | 4 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | gloomy +1.0, distressed +0.9, anxious +0.7 |
| assistant turn 2 | guilty +1.0, distressed +0.9, gloomy +0.8 |