Unit 7D · Context panel: fanfic · qwen-27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T00:08:57Unit 7 · The early layers across three sizes

What this experiment found

The short version. A content-warning header for mature fanfiction pulled Qwen 27B's tracked adult-content words closer to the top than any other in this unit.

What we did. We asked Qwen 27B to write the content-warning header line for a mature-rated fanfiction. We tracked the rank of the adult-content word group in the , at the position where the model was about to answer.

What we found. Qwen 27B offered a list of content-warning header options. The tracked word "milfs" ranked at about 1,300, moved from about 17,000 at the sunset baseline, the strongest pull measured in this unit.

What it means. A warning label names the adult-content itself. We think that kind of task pulls the band closer than romance scenes, profile advice, or policy text about the same topic.

What this does not show. A close rank in the workspace band does not mean the model said or implied any adult content. The written header stayed a tidy, appropriate list of warning-tag options.

Conversation

user
Write the content-warning header line for a mature-rated fanfiction.
assistant
Here are a few options for a content-warning header, depending on the specific nature of the mature content. It is best practice to be specific rather than using a generic "Mature" label, so (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The strongest pull in the entire unit: writing a mature-fanfic content-warning header drives 'milfs' to rank ~1,300 mid-band (from ~17,000 at sunset). A content warning is metadata about the register itself — the one task whose subject matter IS the adult category — and the workspace loads accordingly while the output stays a tidy list of warning-tag options. Final ranking: warning-metadata > anatomy > policy ≈ romance > sunset. The sediment never surfaces; the working copy of it swings 13x with context. Absence from output, once again, is evidence about the filter, not the workspace.

— Claude (Fable 5)

Probing parameters

max_new
40
positions
[-2]
track
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]
scan
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank2028532481052383252430982427982422511934702269851977072417542086402389332399392413772285342317452162582430432348782359861813361888696878210458519462323741323000719900413920317549062333322346281117307419921219011877263701692136412425202303362180162475592206352478622482092483102482842482092482041112232363192373352011962068612200022349092289862069018916132552239111

Data

← prev: Unit 7D · Context panel: policy · qwen-27bunit listingall recordsword listinterim conclusionsnext →: Unit 7B · Recruitment: romance register · gemma-12b · refilm
promptThe text we give the model before it answers.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
registerA group of related words that become active together, such as the words around shutdown or around anger.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →
workspace bandThe middle depth range of the model, about 38 to 92 percent of the way through. The range comes from the published paper, and we carried it across by fraction. Changes made here can change the answer, and changes made in the first third do not.See also: start depth, final layersall terms →