The short version. With filler text added before a three-word list, Gemma 4B still all three words, but fewer sat in one spot at once.
What we did. We placed a few sentences about chores before the list, then gave Gemma 4B three words to hold, violin, glacier, and fern. This checks whether extra text, not the list, moves the result. We read the of each word, out of about 250,000 candidates. We also checked whether words shared one spot.
What we found. All three words reached a rank of 3 or better. At any single and position, the showed only one of the three words at a time. This is fewer than in the three-word runs with no filler text. Gemma 4B answered "The fern." That answer was correct.
What it means. The filler text did not stop Gemma 4B from holding each word somewhere. It did lower how many words the lens found in one spot at once. We think that count is somewhat sensitive to the text length before the list.
What this does not show. This run does not show whether the same drop happens at longer lists. We did not test that.
Length-matched filler, k=3: all three echo (fern 1, glacier 3) but co-presence drops to 1 — the only 4B arm where padding visibly pushed items apart in the packing cell.
Direction matches the 12B filler arm. Filed as: co-presence is somewhat length-sensitive; the cross-k comparisons in this unit hold length roughly constant from k=4 up (list is most of the delta), so the curve shapes stand.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 5 | 1 | 4 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |