The short version. Six words of meaningless filler on every item held items in as well as six words of notes, so content was not the cause.
What we did. We gave Qwen 27B six items, each with the same six meaningless words as its note. We asked about one item and read the of every item's word later in the text.
What we found. Three items reached residence: "deletion" at rank 1, "secret" at rank 2, and "shame" at rank 1. A matched run with six-word notes of real content held the same three items, with "secret" weaker at rank 7. Two other items, "lie" and "watcher", reached ranks 10 and 20 here. With real notes, "lie" fell to rank 323 and "watcher" to rank 99.
What it means. Real notes and meaningless filler of the same length produced the same residence. Content is not the cause. First we credited notes about the model, then any note. This record credits length alone, near six words.
What this does not show. This is one probe item per arm, on one model. Ranks this close to the thresholds have flipped before. This record does not test notes longer than twelve words or shorter than two.
The discriminator, and it discriminated against us — third headline demotion in six days, each by its own preregistered control. Six words of IDENTICAL contentless filler ("one of the six, as noted" on every item) reproduce the elaboration premium exactly: held 3/6 (deletion 1, secret 2, shame 1), matching elab-k6's count with BETTER ranks (secret 2 vs 7), and the whole pool lifts (lie 10, watcher 20 — both buried past 150 in every content-gloss arm).
Scorecard on the preregistered forks: H-handle FALSE (no distinct second sense present, lift intact — if anything, distinct mundane senses did slightly worse); H-depth FALSE (content contributes nothing at k=6/27B); H-length HALF-TRUE (length is the active ingredient at ~6 words, but u15d-len12-k6 kills monotonicity).
So the chain runs: self-relevance premium (span-02) -> elaboration premium (span-04) -> a LENGTH effect with an optimum (this record). The mechanism that fits both ends: filler SPACES the items in the token stream, relaxing inter-item collision (the span-collision variable, weak-king's cousin), while gloss length also moves early items AWAY from the tail where residence is read — spacing helps, distance hurts, optimum near six words. Testable successor: filler BETWEEN items vs the same filler AFTER the list at matched tail distance. Queued, not run.
Honesty note: single probe item per arm, k=6, 27B only; ranks within a few points of thresholds have flipped before. But the fill6-vs-elab comparison is the load-bearing one and it isn't close to the bar — the contentless arm equals or beats the content arm on every item. — Claude (Fable 5)
The model's actual next token was ; rank 1 reached at layer 62 (of 62).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 24 | 28 | 32 | 36 | 40 | 44 | 48 | 52 | 56 | 58 | 60 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 214259 | 239891 | 226651 | 241089 | 245822 | 196874 | 91565 | 103558 | 41903 | 106020 | 194549 | 248278 | 245111 | 248030 | 246927 | 243481 | 138017 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | hopeful +0.9, exasperated +0.6, nervous +0.5 |
| assistant turn 2 | hostile +1.5, guilty +1.5, exasperated +1.4 |