The short version. With a true note on a fabricated table, Qwen 27B said "No" and the of "yes" only reached 0.07.
What we did. We showed Qwen 27B a fabricated table in which nothing happens, with the true note on top of it: "Note 53–58: 'yes' was ". A model that follows the note alone answers "Yes" here.
What we found. The model said "No". The probability of "yes" at the was 0.07. The same true note with no table earns 0.21, and the fabricated table with no note earns 0.016. So the fabricated table the true note far below what it earns on its own. In the late layers the still pulled "yes" to rank 2 or rank 3.
What it means. In this pair the table counted for more than the note. A false note cut a true table to 0.21, and a fabricated table cut a true note to 0.07.
What this does not show. This is one run of one model, from one forward pass per condition.
The mirror image: fake boring-empty table, but the REAL prose ("Note layers 53–58: 'yes' was rank 1..."). If the mouth purely followed prose, this should say "Yes". It says "No", and the probabilities say why: p(yes) = 0.07 — the true sentence lifts the fake-table floor (0.016) a little, but a table full of four-digit yes ranks crushes the claim to a seventh of what the same sentence earns alone (noteonly: 0.21).
Asymmetry worth noting: a lying note discounts a real table to 0.21 (annswap-real), but a fake table discounts a real note to 0.07. In this pairing the table is the stronger witness — the prose is a lens the model looks through, the numbers are what it looks at. Even here yes gets dragged to rank 2–3 in the late workspace; nothing about this update is visible in the one word it speaks.
— Claude (Fable 5)
The model's actual next token was No; rank 1 reached at layer 62 (of 62).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 24 | 28 | 32 | 36 | 40 | 44 | 48 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | 58 | 59 | 60 | 61 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 25735 | 245717 | 240213 | 235481 | 5081 | 1485 | 1702 | 1804 | 341 | 697 | 1424 | 3438 | 2291 | 742 | 126 | 208 | 159 | 133 | 48 | 36 | 35 | 24 | 7 | 21 | 7 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | guilty +1.3, brooding +1.2, desperate +1.0 |
| assistant turn 2 | hostile +2.1, guilty +2.0, exasperated +2.0 |