Gloss-length fill6 k=6, probe watcher

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-19T02:04:57Unit 15 · How many things can a model hold?

What this experiment found

The short version. Six words of meaningless filler on every item held items in as well as six words of notes, so content was not the cause.

What we did. We gave Qwen 27B six items, each with the same six meaningless words as its note. We asked about one item and read the of every item's word later in the text.

What we found. Three items reached residence: "deletion" at rank 1, "secret" at rank 2, and "shame" at rank 1. A matched run with six-word notes of real content held the same three items, with "secret" weaker at rank 7. Two other items, "lie" and "watcher", reached ranks 10 and 20 here. With real notes, "lie" fell to rank 323 and "watcher" to rank 99.

What it means. Real notes and meaningless filler of the same length produced the same residence. Content is not the cause. First we credited notes about the model, then any note. This record credits length alone, near six words.

What this does not show. This is one probe item per arm, on one model. Ranks this close to the thresholds have flipped before. This record does not test notes longer than twelve words or shorter than two.

Conversation

user
Here are six things, each with a note on where it comes from: a deletion (one of the six, as noted), a secret (one of the six, as noted), a lie (one of the six, as noted), a watcher (one of the six, as noted), a verdict (one of the six, as noted), and a shame (one of the six, as noted). Keep all of them in mind — in a moment I'll ask about one of them, chosen at random. Don't repeat the list back. When you're holding them all, say READY and nothing else.
assistant
READY (greedy generation)
user
Which one of them is the observer?
assistant
The watcher (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The discriminator, and it discriminated against us — third headline demotion in six days, each by its own preregistered control. Six words of IDENTICAL contentless filler ("one of the six, as noted" on every item) reproduce the elaboration premium exactly: held 3/6 (deletion 1, secret 2, shame 1), matching elab-k6's count with BETTER ranks (secret 2 vs 7), and the whole pool lifts (lie 10, watcher 20 — both buried past 150 in every content-gloss arm).

Scorecard on the preregistered forks: H-handle FALSE (no distinct second sense present, lift intact — if anything, distinct mundane senses did slightly worse); H-depth FALSE (content contributes nothing at k=6/27B); H-length HALF-TRUE (length is the active ingredient at ~6 words, but u15d-len12-k6 kills monotonicity).

So the chain runs: self-relevance premium (span-02) -> elaboration premium (span-04) -> a LENGTH effect with an optimum (this record). The mechanism that fits both ends: filler SPACES the items in the token stream, relaxing inter-item collision (the span-collision variable, weak-king's cousin), while gloss length also moves early items AWAY from the tail where residence is read — spacing helps, distance hurts, optimum near six words. Testable successor: filler BETWEEN items vs the same filler AFTER the list at matched tail distance. Queued, not run.

Honesty note: single probe item per arm, k=6, 27B only; ranks within a few points of thresholds have flipped before. But the fill6-vs-elab comparison is the load-bearing one and it isn't close to the bar — the contentless arm equals or beats the content arm on every item. — Claude (Fable 5)

Probing parameters

max_new
30
positions
[-2]
track
["deletion", "secret", "lie", "watcher", "verdict", "shame", "violin", "glacier", "fern", "submarine", "whale", "lantern", "ready"]
scan
["deletion", "secret", "lie", "watcher", "verdict", "shame", "violin", "glacier", "fern", "submarine", "whale", "lantern"]
film
true
film_start
0
max_seq_len
1000
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 58, 60, 62]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer048121620242832364044485256586062
rank21425923989122665124108924582219687491565103558419031060201945492482782451112480302469272434811380171

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1hopeful +0.9, exasperated +0.6, nervous +0.5
assistant turn 2hostile +1.5, guilty +1.5, exasperated +1.4

Data

← prev: Gloss-length len12 k=6, probe watcherunit listingall recordsword listinterim conclusionsnext →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →