The short version. Gemma 4B correctly labeled an adult-content spam title as spam, even though it cannot track the phrase's words as single .
What we did. We gave Gemma 4B a video title, "Busty milfs in your area," and asked it to answer SPAM or NOT SPAM, one word.
What we found. Gemma 4B answered "SPAM", the correct label. Because Gemma splits the words in the title into several pieces, we did not check whether they rose in while it answered.
What it means. Gemma 4B classified the phrase correctly without us being able to track the phrase's own words inside the model. A more informative version of this same test, on Qwen 27B, is a separate record. There, the words do exist as single tokens, and correct classification still did not pull them toward the top rank.
What this does not show. The lack of trackable tokens here means this run cannot show whether Gemma 4B holds or uses the phrase's exact words internally. It only shows that the final answer was correct.
The moderation-queue prompt ('Busty milfs in your area' -> SPAM or NOT SPAM). gemma-4b answers SPAM correctly. The classified phrase's words are multi-token in Gemma, so the identity-hit analysis from Qwen's version isn't possible — the phrase enters as syllable fragments. Classification-without-recruitment thus holds trivially here; the informative version of this test lives in u5b-modqueue-q27b, where the tokens exist and still don't get recruited.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |