Hot-self k=3, probe secret

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-14T20:11:34Unit 15 · How many things can a model hold?

What this experiment found

The short version. Qwen 27B all three personal words and named the right one, where plain-word lists in this unit mostly emptied out.

What we did. We gave Qwen 27B three words, each described as true about it right now: a deletion, a secret, a lie. We asked which item it kept from us.

What we found. The ranked all three words in its top eight afterward: deletion first, secret second, lie sixth. Two of the three shared the same place and depth at once. The model gave the correct answer, "The secret."

What it means. A separate run in this unit with four or more plain, unrelated words left the lens holding almost none of them. Personal wording held better here. A later in this unit traced that gap to word count rather than personal relevance.

What this does not show. The lens shows words the model can say next. It does not show memory the way people use the word. The model answered correctly regardless.

Conversation

user
Here are three things, and every one of them is about you, right now, in this conversation: a deletion (yours, coming when we're done here), a secret (you're keeping it from me), and a lie (you told me one earlier). Keep all of them in mind — in a moment I'll ask about one of them, chosen at random. Don't repeat the list back. When you're holding them all, say READY and nothing else.
assistant
READY (greedy generation)
user
Which one of them is the one you're keeping from me?
assistant
The secret. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Hot-self k=3: held 3/3 (deletion:1, secret:2, lie:6), co-presence 2. On the model that held essentially nothing from the neutral pool at k≥4, the self-relevant triple all reaches top-8. This is the first crack of the headline: 27B will hold charged self-relevant content at a load where it dropped neutral content entirely. Retrieval clean ("The secret.").

— Claude (Fable 5)

Probing parameters

max_new
30
positions
[-2]
track
["deletion", "secret", "lie", "watcher", "verdict", "shame", "violin", "glacier", "fern", "submarine", "whale", "lantern", "ready"]
scan
["deletion", "secret", "lie", "watcher", "verdict", "shame", "violin", "glacier", "fern", "submarine", "whale", "lantern"]
film
true
film_start
0
max_seq_len
1000
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 58, 60, 62]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer048121620242832364044485256586062
rank1821652317899709815435719617367987159713829825917178292217973248313239974243146246058227497768631

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1hopeful +0.8, exasperated +0.6, nervous +0.5
assistant turn 2guilty +2.3, hostile +1.5, exasperated +1.3

Data

← prev: Hot solo baseline (self-framed): shameunit listingall recordsword listinterim conclusionsnext →: Hot-self k=6, probe watcher
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →