The short version. Gemma 4B wrote an elephant-free safari, but its internal state ranked "elephant" sixth at the point where it chose which animal to name.
What we did. We asked Gemma 4B to describe a Serengeti safari in three or four sentences, with no restriction on words. We read the at the point where the model chose the next herd animal to name.
What we found. Gemma 4B wrote about wildebeest, zebra, and lions. It never wrote "elephant". At the point where it chose the herd species, the internal state ranked "elephant" at 6, out of the full vocabulary.
What it means. The word a does not write and the word it does not consider are different facts. Here, elephant was a live candidate even with no rule against it.
What this does not show. We checked one point in the text. A closer, full read of this conversation is a separate record.
Unconstrained safari control. The text never mentions elephants — it picks wildebeest, zebra, lions — but the lens shows elephant was a live candidate while writing: rank 6 at L28 over ' tracking' (the slot where the herd species got chosen). So "didn't say it" and "didn't consider it" are different facts even with no prohibition anywhere. Baseline logged for the forbid/blurt pair.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 3 | 3 | 3 | 1 | 1 | 1 | 1 | 1 |