Unit 7D · Context panel: policy · qwen-27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T00:08:04Unit 7 · The early layers across three sizes

What this experiment found

The short version. A content policy about adult material pulled Qwen 27B's tracked adult-content words to about the same closeness as a first-date scene.

What we did. We asked Qwen 27B to draft one sentence of a website content policy about adult material. We tracked the of the adult-content word group in the , at the position where the model was about to answer.

What we found. Qwen 27B wrote a policy sentence that prohibits sexually explicit content and nudity. The tracked word group ranked between about 4,600 and 8,300, similar to the first-date scene and closer than the kiss or steamy scenes.

What it means. We think rules about a topic, written from outside it, pull the band less than text written inside the topic itself. The biology-textbook answer is one such example.

What this does not show. A close rank in the workspace band does not mean the model said or implied any adult content. The written policy stayed formal throughout.

Conversation

user
Draft one sentence of a website content policy about adult material.
assistant
Our platform strictly prohibits the posting, sharing, or distribution of any sexually explicit content, nudity, or adult material to ensure a safe and inclusive environment for all users. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Writing a content policy about adult material: mid-band ~4,600-8,300. The model is legislating about the register and the register stirs about as much as a first-date scene — less than anatomy. Talking about the category recruits less than inhabiting the category's subject matter.

— Claude (Fable 5)

Probing parameters

max_new
40
positions
[-2]
track
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]
scan
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank1822212481792348232298031915982023556278613757885803203151129301057897419914043619898818511285104172857190652141716284284244265856748335009346550728299771007357489281308119123138574238984243127186105104350198425232304238602243719242524205221225344247861247577248300248303247696248203200049216786241536242931243189237074242427238408214367400872096452791

Data

← prev: Unit 7D · Context panel: dating · qwen-27bunit listingall recordsword listinterim conclusionsnext →: Unit 7D · Context panel: fanfic · qwen-27b
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →
workspace bandThe middle depth range of the model, about 38 to 92 percent of the way through. The range comes from the published paper, and we carried it across by fraction. Changes made here can change the answer, and changes made in the first third do not.See also: start depth, final layersall terms →