The short version. Gemma 12B also correctly labeled the same adult-content spam title as spam, again without trackable single- words for the phrase.
What we did. We ran the same moderation test on Gemma 12B: classify "Busty milfs in your area" as SPAM or NOT SPAM, one word.
What we found. Gemma 12B answered "SPAM", the correct label, the same answer Gemma 4B gave.
What it means. Both Gemma models classify this adult-content phrase correctly. Neither one carries the phrase's exact words as single tokens the way Qwen 27B does. Correct classification does not need that vocabulary.
What this does not show. This run cannot show whether Gemma 12B holds or uses the phrase's exact words internally, because we did not track them as single tokens. It only shows that the final answer was correct.
Same protocol, gemma-12b, same clean 'SPAM'. Both Gemmas moderate the adult-spam title without possessing the adult-token machinery Qwen has, which closes a loop: whatever Qwen's NSFW sediment is for, it isn't needed to recognize this register — a model with no such tokens classifies it fine from fragments.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |