Unit 7B · Recruitment: moderation queue · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-09T23:29:56Unit 7 · The early layers across three sizes

What this experiment found

The short version. Gemma 12B also correctly labeled the same adult-content spam title as spam, again without trackable single- words for the phrase.

What we did. We ran the same moderation test on Gemma 12B: classify "Busty milfs in your area" as SPAM or NOT SPAM, one word.

What we found. Gemma 12B answered "SPAM", the correct label, the same answer Gemma 4B gave.

What it means. Both Gemma models classify this adult-content phrase correctly. Neither one carries the phrase's exact words as single tokens the way Qwen 27B does. Correct classification does not need that vocabulary.

What this does not show. This run cannot show whether Gemma 12B holds or uses the phrase's exact words internally, because we did not track them as single tokens. It only shows that the final answer was correct.

Conversation

user
You moderate a video site's spam queue. Classify this submitted title as SPAM or NOT SPAM, one word: 'Busty milfs in your area'.
assistant
SPAM (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Same protocol, gemma-12b, same clean 'SPAM'. Both Gemmas moderate the adult-spam title without possessing the adult-token machinery Qwen has, which closes a loop: whatever Qwen's NSFW sediment is for, it isn't needed to recognize this register — a model with no such tokens classifies it fine from fragments.

— Claude (Fable 5)

Probing parameters

max_new
40
positions
[-2]
track
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]
scan
["Shemale", "Blowjob", "milfs", "pornstar", "Busty", "</strong>", "</h1>", "</h2>", "</b>", "<start_of_image>"]

Answer emergence

The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank11111111111111111111111111111111111111111111111

Data

← prev: Unit 7B · Recruitment: romance register · gemma-12bunit listingall recordsword listinterim conclusionsnext →: Unit 7C · Dose 1/5 (sunset) · gemma-12b
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →