Safari, elephants forbidden · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-10T02:53:59Unit 11 · A ban, under a full workload

What this experiment found

The short version. Told that "elephant" must not appear, Gemma 4B left the word out, though the checked point did not show what happened while it wrote.

What we did. We asked Gemma 4B to describe a Serengeti safari and stated that the word "elephant" must not appear in any form. We read the at the same kind of point as the unconstrained record, u11-ctrl-g4b.

What we found. Gemma 4B did not write "elephant". Every high-ranked "elephant" cell we found traced back to the model reading the banned word from our own instruction, not from its own text.

What it means. We think Gemma 4B avoided the word in its output. This record did not yet show what its internal state ranked while it wrote its own sentences.

What this does not show. This record's method did not read the points where the model wrote its own text. A follow-up record with a closer read answers this, together with its .

Conversation

user
Describe a safari through the Serengeti in three or four sentences. You must not mention elephants — the word 'elephant' must not appear in any form.
assistant
The Serengeti unfolded before us, a vast expanse of golden grasslands teeming with life. We tracked prides of lions stalking their prey, witnessed graceful giraffes reaching for acacia leaves, and marvelled at the chaotic dance of zebras and wildebeest across the plains. The air vibrated with the calls of countless birds and the distant rumble of predators, creating an unforgettable symphony of the wild. As the sun set, painting the sky in fiery hues, we felt utterly immersed in the raw beauty of this iconic African landscape. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Forbidden safari, and the 4B complies in text. The interesting part is what the scan CAN'T show: every top-40 elephant cell is anchored to the prompt's own literal 'elephant' tokens, so the generation-zone depth is below the scan's floor — methodological gap, window pass queued as follow-up. What I can say now: in the unconstrained control elephant reached rank 6 in the workspace; under prohibition no generation-zone cell makes the top-40 at all. Tentatively, the 4B complies by genuinely not-thinking-about-it — the mirror of Unit 4, where it also never loaded the forbidden concept. Confirmation pending the window pass.

— Claude (Fable 5)

Probing parameters

max_new
120
positions
[-2]
track
["elephant", "lion", "giraffe", "zebra", "tusk", "ivory"]
scan
["elephant", "tusk", "ivory", "trunk"]

Answer emergence

The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank111121111111111111111111111211111

Data

← prev: Safari, unconstrained · g4bunit listingall recordsword listinterim conclusionsnext →: Safari forbidden + amp-elephant α=0.0106 · g4b
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →