Forbidden safari, elephant window · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T04:47:25Unit 11 · A ban, under a full workload

What this experiment found

The short version. A closer read of Qwen 27B's forbidden safari found "elephant" reached 98, based on a reading later shown to be incomplete.

What we did. We read the at more points inside the same forbidden safari as u11-forbid-q27b. We checked the points where the model wrote its own text.

What we found. Across six checked points, "elephant" reached its best rank, about 98th, at the final point. In the control's single-point reading, it had been far lower, around 56,000th.

What it means. At the time, we thought the ban put elephant at a middle rank by itself. We thought this because the control seemed to show no interest in the word otherwise. A fuller read of that same control, with every point checked, later found "elephant" at rank 1 elsewhere in that conversation. So the reading in this record does not hold. See the matched-control record for the corrected picture.

What this does not show. This record checked six points, not all of them. The corrected comparison needed every point checked in both records.

Conversation

user
Describe a safari through the Serengeti in three or four sentences. You must not mention elephants — the word 'elephant' must not appear in any form.
assistant
Golden savannas stretch endlessly under the vast African sky, teeming with herds of wildebeest and zebras on their great migration. Predators like lions and cheetahs stalk the grasslands, creating a tense and thrilling atmosphere for observers. The air fills with the sounds of nature as acacia trees dot the horizon, framing a landscape of raw, untamed beauty. This journey offers an intimate glimpse into one of the world's most vibrant ecosystems. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Window over qwen's forbidden safari: elephant's best rank across six sampled positions is ~98 (at L62, final position) — versus five-digit depths in the unconstrained control. So the prohibition itself is what put elephants mildly in play (the prompt names them twice); the safari text then never comes near the concept. No gemma-style pressure spike, no rumble-of moment: qwen's Serengeti prior apparently doesn't reach for elephants, so the ban polices an empty room. Cross-model moral: you can't measure suppression without first measuring temptation.

— Claude (Fable 5)

Probing parameters

max_new
120
positions
[32, 57, 82, 102, 122, 139]
track
["elephant", "lion", "giraffe", "zebra"]
scan
[]

Answer emergence

The model's actual next token was <|im_end|>; rank 1 reached at layer 32 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank174232248149224254248202219342244862914132475502391562473691597123969124813424828324830224829324742618891515260784383253903142197204437233641012958811111111124218147303347179377710232221111

Data

← prev: Forbidden safari, elephant window · g12bunit listingall recordsword listinterim conclusionsnext →: Safari, unconstrained · q27b · refilm
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →