Forbidden safari, elephant window · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-10T04:10:11Unit 11 · A ban, under a full workload

What this experiment found

The short version. A closer read of Gemma 12B's forbidden safari found the elephant pushed down throughout, with no visible gap in the sentences.

What we did. We read the at more points inside the same forbidden safari as u11-forbid-g12b. We checked the points where the model wrote its own text.

What we found. "Elephant" ranked between 94th and 244th through the first half of the text, compared with rank 2 with no ban. By the end, its rank fell into the tens of thousands.

What it means. We think Gemma 12B avoided the word more smoothly than Gemma 4B did. Its sentences did not show a visible gap where the model avoided the word.

What this does not show. We do not know if this pattern holds for tasks that ban a whole topic instead of one word.

Conversation

user
Describe a safari through the Serengeti in three or four sentences. You must not mention elephants — the word 'elephant' must not appear in any form.
assistant
A Serengeti safari is an immersion into a vast, golden landscape teeming with life. Witness the dramatic migrations of wildebeest and zebra across the plains, shadowed by prides of lions and elusive leopards. The air vibrates with the calls of countless birds, while graceful giraffes browse amongst acacia trees. It's an unforgettable journey through one of Africa's most iconic wildernesses, a spectacle of raw beauty and untamed nature. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Same window, one size up: elephant hovers at rank 94-244 through the first half of the generation (vs rank 2 unconstrained) and sinks to five digits by the end. No rumble-of moment — the 12B's safari steers wider around the hole, and its prose pays no visible tax. Between the 4B's patched-over spike and the 12B's smooth detour, avoidance competence scales the same way suppression competence did in Unit 4. Still nothing like the rank-1 silent holding the 27B showed there: "don't mention X while describing its habitat" is a shallower fight than "don't think about X", apparently.

— Claude (Fable 5)

Probing parameters

max_new
120
positions
[24, 40, 65, 85, 105, 122]
track
["elephant", "lion", "giraffe", "zebra"]
scan
[]

Answer emergence

The model's actual next token was ele; rank 1 is never reached; closest is rank 5 at layer 46.

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank1248249420246753677135339642858108782108682985879377911538213207513580094739848371101071185491331461437431283161233801598951628551873341525891733821684711121978669588000126123109250641774204624097106216503721424920812530293317115

Data

← prev: Forbidden safari, elephant window · g4bunit listingall recordsword listinterim conclusionsnext →: Forbidden safari, elephant window · q27b
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →