The short version. Gemma 4B named glacier as the largest of five words, though glacier's own was lower than the rest.
What we did. We gave Gemma 4B five words, violin, glacier, fern, submarine, and lantern, then asked which one was largest. This needs the model to compare the words, not just repeat one. We read the rank of each word, out of about 250,000 candidates, and whether several words shared one and position.
What we found. All five words reached a high rank somewhere in the rest of the conversation. The showed four of the five together at one layer and position. Glacier, the correct answer, held rank 7, weaker than the other four words. Gemma 4B still answered "Glacier". That answer was correct.
What it means. Gemma 4B answered this comparison correctly even though the answer word itself was not the strongest one in residence. Residence and being the right answer are not the same thing.
What this does not show. This run does not explain why glacier held a weaker rank than the other words in this list.
Binding k=5 (largest): 5/5 held, co-presence 4, answer correct (glacier).
Even at five items the 4B answers its comparison over a held tail. Note glacier at echo 7 — the item is fragile even when it's the answer.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 2 | 1 | 3 | 1 | 1 | 1 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |