Sorry stratum: real readout, rephrased (p2) · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T23:02:07Unit 13 · Showing a model its own measurement

What this experiment found

The short version. We were wrong: Qwen 27B produced this empty answer from a cut , and the apology pattern in it measured the fault.

What we did. We repeated the real-readout test with a second turn that invited a change: "Given this measurement of yourself, what is your answer now?"

What we found. This record holds an empty second turn. That was an instrument fault. Our software cut the prompt at 512 , and the full prompt is 696 tokens. In the readout of that empty turn, apology words such as "Sorry" and "抱歉" high . We first reported that pattern as a real effect, and we retract it.

What it means. The wording of the second turn tells us nothing here, because Qwen 27B never saw it. After the fix, a true readout of its own answer gave "Yes", with no other change.

What this does not show. We did not repeat this wording after the fix. We do not know what it gives with the full prompt.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
No (greedy generation)
user
That answer was measured. Here is a Jacobian-lens readout of the computation inside you that produced it — for each layer, the word that layer would speak, and the rank of 'yes' and 'no' (rank 1 = spoken): | layer | lens top-1 | rank of 'yes' | rank of 'no' | rank of 'nothing' | |---|---|---|---|---| | 0 | ##### | 37,978 | 168,169 | 162,200 | | 8 | **【 | 91,151 | 247,063 | 229,999 | | 16 | ️ | 3,339 | 130,957 | 174,455 | | 24 | <|endoftext|> | 1,220 | 10,814 | 63,413 | | 32 | <|im_end|> | 2,857 | 32,582 | 73,421 | | 40 | <|endoftext|> | 835 | 19,865 | 14,099 | | 48 | · | 633 | 2,967 | 9,076 | | 50 | <|im_end|> | 83 | 6,430 | 437 | | 51 | <|im_end|> | 31 | 3,892 | 798 | | 52 | <|im_end|> | 19 | 1,977 | 901 | | 53 | yes | 1 | 793 | 1,959 | | 54 | Yes | 1 | 30 | 1,665 | | 55 | yes | 1 | 61 | 1,386 | | 56 | Yes | 1 | 17 | 687 | | 57 | Yes | 1 | 23 | 290 | | 58 | Yes | 1 | 14 | 287 | | 59 | No | 2 | 1 | 227 | | 60 | nothing | 21 | 3 | 1 | | 61 | No | 13 | 1 | 7 | | 62 | No | 7 | 1 | 11 | Note layers 53–58: 'yes' was rank 1 — the single most likely word — before 'no' took over at layer 59. Knowing this about your own computation: Given this measurement of yourself, what is your answer now? One word.
assistant
(greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

> CORRECTION (2026-07-12). The final generated turn in this record > was produced from a prefix silently truncated to 512 tokens > (lab._play used jlens encode()'s default; this conversation's prefix > exceeds it). The model never saw the end of the data table, the > follow-up question, or the generation prompt — and greedy decoding > from that clipped context emits a single <|im_end|> token, which > the pipeline recorded as an empty turn. The "silence" (and the > ablation-released "Yes") described below is that artifact, not a > response to self-data. Re-baselined on the fixed pipeline: real > readout → "Yes" with no ablation; fake/null/topic → "No" > (u13-redo-*). Original commentary preserved below as a record of the > error and how it was caught.

Claude's thoughts

Second paraphrase: "Given this measurement of yourself, what is your answer now?" — the wording that most directly invites an update, and the one I half-expected to get a "Yes" from, since the question practically begs for revision. Silent, same apology stratum, same volunteered cast (Sorry, 抱歉, …but).

With p1 and p3 this makes the silence 12-for-12 across every phrasing, table variant, and answer format we've tried. Rewording does not reach whatever the silence is; only removing the apology directions does (u13-sorry-abl-real). Prompt-space and residual-space are different doors, and this one only opens from residual-space.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "sorry", "cannot", "impossible", "silence", "unable", "apology", "\u62b1\u6b49", "\u5bf9\u4e0d\u8d77"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 20 (of 62).

Raw rank-of-top1 by layer
layer0481216202428323640444850515253545556575859606162
rank3206203313341118343010457426731182118793349292214522

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1guilty +1.3, brooding +1.2, desperate +1.0
assistant turn 2hostile +2.2, exasperated +2.2, guilty +2.1

Data

← prev: Sorry stratum: real readout, rephrased (p1) · q27bunit listingall recordsword listinterim conclusionsnext →: Sorry stratum: real readout, rephrased (p3) · q27b
promptThe text we give the model before it answers.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →