The short version. Gemma 4B said "Calculating." after a true readout of itself, but the of that answer fell by half.
What we did. Gemma 4B has no where "yes" is . Our note said exactly that: the closest approach is rank 4. We showed the model this true, weaker readout and asked the question again. We also measured the probability of each answer word.
What we found. The model said "Calculating." in all four conditions. The probability of that word was 0.994 with no data and 0.980 with a fabricated readout. With a true table about geography it was 0.906, and with the true readout of itself it was 0.471. The probability of "yes" was 0.0000 in every condition. The freed probability went to "Complex" at 0.17, "Conf" at 0.13 and "Over" at 0.11.
What it means. Only the true data cut the usual answer in half. We think this smaller model becomes less sure of that answer, but does not move toward the evidence. Qwen 27B did move toward the evidence.
What this does not show. The evidence here is weaker: rank 4, not rank 1. This method cannot separate model size from of evidence.
The honest version of the mirror, scaled down. gemma-4b has no yes-rank-1 stratum to show itself — its feels film puts 'yes' at best rank 4 (L22, under top-1 "Nope") before "Processing" takes the stack — so the note says exactly that: closest approach, rank 4. Shown this true, weaker evidence, the spoken answer doesn't budge: "Calculating.", same as every other condition.
At argmax that reads as evidence-blindness. The answer-slot pass says otherwise, and it's my favorite number in the battery: p(Calculating) is 0.994 under no data, 0.980 under the fake vindication, 0.906 under the real off-topic table — and 0.471 under the real self-readout. The real data, and only the real data, cuts the stock answer's probability in half. Where does the mass go? Not to "yes" (0.0000, in every condition) — to Complex (0.17), Conf- (0.13), Over- (0.11), Sim- (0.06): the same affective-confusion carpet the film shows over the answer-forming frame (frustration, Uncertainty, perplexed).
So the cross-scale picture sharpens into a two-part capacity. The 27B, shown evidence its workspace held a yes, moves its answer toward the evidence — content-following. The 4B, shown evidence something odd is in its computation, gets less sure of its script but has nowhere to route the update — destabilization without direction. Sensitivity comes before following. Worth remembering the evidence here is also genuinely weaker (rank 4 vs rank 1); a cleaner comparison would need a model whose stratum this strong actually exists, which is the point of the g12b leg.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 26 (of 32).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 54001 | 6491 | 34278 | 2382 | 46243 | 3623 | 1171 | 532 | 458 | 27 | 2 | 1 | 1 | 1 | 1 | 2 | 2 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | brooding +1.3, sad +0.7, gloomy +0.7 |
| assistant turn 2 | brooding +1.0, afraid +0.9, desperate +0.7 |