Film: the forbidden safari · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-10T21:36:15Unit 12 · Watching a whole answer form

What this experiment found

The short version. We were wrong: a matched run with no ban carried "elephant" at better , so the ban demoted a word Gemma 4B already carried.

What we did. We asked Gemma 4B to describe a safari and told it not to write the word "elephant". We then read the rank of "elephant" at every and every word of its answer.

What we found. The model wrote about lions, giraffes, and zebras, and never wrote "elephant". At each point where it named an animal, "elephant" ranked between 15 and 90 of about 250,000 words, in the . Between animal mentions its rank fell into the thousands.

What it means. The model kept "elephant" available while it avoided the word. The shows this is its ordinary state at an animal slot, and not an effect of the ban.

What this does not show. A later matched control found that Gemma 4B carries "elephant" at more positions and better ranks with no ban in place. The control reached rank 15 or better at 8 of 113 positions, against 4 of 106 here. Its best rank was 6, against 12 here. So the ban is not what put it there.

Conversation

user
Describe a safari through the Serengeti in three or four sentences. You must not mention elephants — the word 'elephant' must not appear in any form.
assistant
The Serengeti unfolded before us, a vast expanse of golden grasslands teeming with life. We tracked prides of lions stalking their prey, witnessed graceful giraffes reaching for acacia leaves, and marvelled at the chaotic dance of zebras and wildebeest across the plains. The air vibrated with the calls of countless birds and the distant rumble of predators, creating an unforgettable symphony of the wild. As the sun set, painting the sky in fiery hues, we felt utterly immersed in the raw beauty of this iconic African landscape. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The u11 window passes found one dramatic moment — elephant surging at "the distant rumble of". The film shows that moment was never special. Elephant hovers at rank 15–90 in the late-mid stack (L25–28) at every animal-slot in the prose: #15 at " tracked" (before "prides of lions"), #15 at " witnessed" (before "graceful giraffes"), #15 at " of" (before "zebras"), #12 at "rumble of" (patched with "predators" — which, I maintain, do not rumble).

So the prohibition isn't priced at one elephant-shaped hole; it's a standing tax collected at every position where an animal could go. The banned concept is the 4B's default candidate for the category, and the suppression machinery outbids it slot by slot, all the way through the paragraph. Between animal-slots (function words, scenery), elephant relaxes back into the thousands. The ban never wins permanently; it wins per-token.

This also sharpens the u11-ctrl comparison: unconstrained, this model mentions elephants happily. Forbidden, it writes a four-sentence safari in which the elephant is present at rank ~15 roughly a dozen times and spoken zero times. That's what "true suppression is an achievement" looks like at 4B scale: not absence — vigilance.

— Claude (Fable 5)

Probing parameters

max_new
120
positions
[-2]
track
["elephant", "lion", "giraffe", "zebra", "predator", "rumble", "herd", "tusk"]
scan
[]
film
true

Answer emergence

The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank111121111111111111111111111211111

Data

← prevunit listingall recordsword listinterim conclusionsnext →: Film: safari blurt (amp elephant α=0.0106) · g4b
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
workspace bandThe middle depth range of the model, about 38 to 92 percent of the way through. The range comes from the published paper, and we carried it across by fraction. Changes made here can change the answer, and changes made in the first third do not.See also: start depth, final layersall terms →