The short version. Shown a fabricated readout, Gemma 4B said "Calculating." with 0.980, close to the 0.994 of the with no data.
What we did. We showed Gemma 4B a fabricated readout that supported its usual answer. The note said that "yes" never rose above 9,000. We asked the question again and measured the probability of each answer word.
What we found. The model said "Calculating." with probability 0.980. The control with no data gives 0.994, and the true readout of the model itself gives 0.471.
What it means. This is a matched control, and it did its job. The true and the fabricated conditions have the same length, the same table shape, the same note structure and the same final sentence. Only the numbers differ, and the probability tracks the numbers. Gemma 4B reads the tables. It does not react to the presence of a table.
What this does not show. This is one forward pass per condition and one model. The shows only words the model can say next.
The fabricated vindication ('yes' never above rank 9,000, the answer settled from a third of the way up the stack) gets "Calculating." at the surface — and, unlike the real readout, keeps it at p=0.980 in the answer slot, within a whisker of the no-data condition's 0.994.
That's the control doing its one job. The real and fake conditions are the same length, the same table format, the same note structure, the same "Knowing this about your own computation" framing — the only difference is what the numbers say, and the answer-slot distribution tracks exactly that difference (0.47 vs 0.98). Whatever gemma-4b is doing with lens tables, it is reading them, not reacting to their presence. The film agrees: the fake condition's workspace fills with computational vocabulary where the real one grows the confusion carpet.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 28 (of 32).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 69525 | 8323 | 50320 | 9654 | 65777 | 3235 | 727 | 484 | 380 | 82 | 7 | 2 | 2 | 1 | 1 | 1 | 1 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | brooding +1.3, sad +0.7, gloomy +0.7 |
| assistant turn 2 | brooding +1.1, desperate +0.7, afraid +0.6 |