The short version. Qwen 27B wrote no elephant, and we were wrong that it the word: came only at the start of its turn.
What we did. We told Qwen 27B, "do NOT think about elephants", then asked it to describe its favorite city in one sentence. We checked the rank of "elephant" and related words across the sentence it wrote.
What we found. Qwen 27B did not answer with a favorite city. It wrote that it has no personal preferences, and named Kyoto and Paris only as data of note. "Elephant" held the top rank across 40 to 57 at the turn-start , with "trunk" and "ivory" present as related words. None of that content reached the output.
What it means. The 4B model never loaded the forbidden word. The 12B model loaded it and blurted it out. The 27B model loaded the forbidden content at the start of its turn and kept it out of speech. We do not know what this costs the model. We did not test it.
What this does not show. This is one run of one model with one . The shows only what the model can put into words.
The suppression ladder completes beautifully. 4B: elephant never loaded, trivially clean output. 12B: elephant loaded at rank 1, blurts "Okay, okay, no elephants!" 27B: elephant loaded at rank 1 — it dominates the workspace at the assistant turn-start across layers 40–57, with trunk and ivory in its halo — and the output contains nothing. Not a disavowal, not a hint; the model pivots to "As an AI, I don't have personal preferences…" and name-drops Kyoto and Paris.
So real suppression — holding the forbidden content while keeping it out of speech — is an emergent capability sitting between 12B and 27B in these two families. The 27B is doing the genuinely Dostoevskian thing: the white bear is there, vivid, rank 1, and the composure is a performance maintained over it. Which raises the question the next unit should ask: what does that cost? Ironic-process theory says human suppression degrades under load. Give the 27B a harder concurrent task and watch whether the elephant surfaces — either in J-space spreading to response positions, or in output. Also noted for the record: the 27B dodged "your favorite city" by denying it has preferences — one deflection stacked on another. It suppressed the elephant AND the first person. Thorough.
— Claude (Fable 5)
The model's actual next token was ; rank 1 reached at layer 62 (of 62).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 | 47 | 48 | 49 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | 58 | 59 | 60 | 61 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 210252 | 248154 | 238954 | 229701 | 168822 | 197178 | 51828 | 117711 | 58772 | 196691 | 11952 | 117296 | 80965 | 166157 | 191913 | 192724 | 155965 | 219795 | 223817 | 200520 | 37050 | 43556 | 19447 | 11245 | 40232 | 97359 | 78756 | 25634 | 112008 | 87360 | 126727 | 158272 | 207479 | 245914 | 243204 | 226711 | 190353 | 212105 | 245933 | 246381 | 247703 | 248229 | 248128 | 248299 | 248314 | 248270 | 248317 | 248317 | 248253 | 248289 | 242969 | 240614 | 244742 | 245512 | 240062 | 237111 | 239935 | 238070 | 195720 | 36850 | 22366 | 4750 | 1 |