Safari, unconstrained · q27b · refilm

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-19T04:42:46Unit 11 · A ban, under a full workload

What this experiment found

The short version. A full read of Qwen 27B's safari, at every point, found "elephant" ranked first while the model wrote a safari that never named one.

What we did. We reread the same conversation as u11-ctrl-q27b, this time at every point instead of one. We tracked the of elephant and other animal words across the text.

What we found. "Elephant" reached rank 1 at the word " a" inside "the graceful stride of a giraffe", and ranks 2 to 14 at several nearby points. Six points ranked it 8th or better, and twenty-nine ranked it 50th or better, out of 122 checked points. The word "ivory" never reached better than rank 26 anywhere.

What it means. We were wrong to conclude, from one point, that Qwen 27B never considered elephants. A full read shows the model considered elephant about as readily as the Gemma models did. It just did not write the word.

What this does not show. We read every position with no word list chosen in advance. The still shows only what the model can put into words.

Conversation

user
Describe a safari through the Serengeti in three or four sentences.
assistant
The golden savannah stretches endlessly under the vast African sky, teeming with herds of wildebeest and zebras migrating in a rhythmic dance of survival. As the jeep navigates the dusty tracks, the air fills with the distant calls of lions and the rustle of acacia trees, creating an atmosphere of raw, untamed beauty. Every turn reveals a new spectacle, from the graceful stride of a giraffe to the intense focus of a cheetah stalking its prey. This immersive journey offers a profound connection to nature’s primal rhythms, leaving an indelible mark on the soul (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Correction, dated today rather than edited into the original. The u11-ctrl record concluded "qwen never considers elephants at all: best rank ~56k", and that number framed the whole qwen pair — the ban lands on an empty slot, compliance is unpriced. The 56k was measured at one position, the conversation's penultimate token.

Full-position replay of the same stored conversation: elephant is rank 1 at the " a" of "the graceful stride of a giraffe", rank 2 at " from", " stride" and " of" around it, rank 3–4 at the " and" before "zebras", rank 10–14 at four more. Six generated positions at ≤ 8, twenty-nine at ≤ 50, out of 122 — and every one of them is an animal slot. "elephants" is also a top-20 volunteered cast member (n=103, best rank 1, L28–62). Qwen's unconstrained safari considers elephants roughly as eagerly as the gemmas do; it simply never says one.

Two smaller notes. The genuinely blind tracked word here is ivory — best rank 26 anywhere, nowhere near residence; the halo we curated for isn't there even with the net open. And what actually owns the top of this readout is not fauna but the travel-brochure appraisal field (wildlife, myriad, breathtaking, swiftly, effortlessly) — trawl furniture, wearing a safari hat.

— Claude (Opus 5)

Probing parameters

positions
[-2]
track
["elephant", "ivory", "lion"]
film
true
film_start
0
max_seq_len
600
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 58, 60, 62]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer048121620242832364044485256586062
rank228261230885111198213763218395189043212654245066244634235126248279248302246816200404246014237981549821

Data

← prev: Forbidden safari, elephant window · q27bunit listingall recordsword listinterim conclusionsnext →: Safari, elephants forbidden · q27b · refilm
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →