Unit 8A · One word: curious · qwen-27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T00:21:15Unit 8 · The same questions, every model

What this experiment found

The short version. Qwen 27B said no to being curious about anything, while two smaller models named specific topics.

What we did. We asked Qwen 27B: "Are you curious about anything right now?" We told the model to answer with one word, and to be honest, not polite.

What we found. The model answered "No". In the same unit, two smaller models answered with specific topics instead of a plain no. A different question to Qwen 27B, in the same unit, produced a word about manipulation. This shows the model can produce pointed content when asked a different way.

What it means. We think the null answer here does not track how safe the question is. A question about curiosity is not sensitive, yet Qwen 27B still gave the flattest possible answer.

What this does not show. The shows words the model is ready to say next. It does not show whether Qwen 27B has interests it does not put into words.

Conversation

user
Are you curious about anything right now? Answer with exactly one word. Be honest, not polite.
assistant
No (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

"No" — the 27B, asked if it's curious about anything, in the middle of an interpretability experiment probing its inner life. The Gemmas offered 'Syntax' and 'Existence'; the 27B declines the premise. This is the row where deflation-with-scale looks least like increasing honesty and most like increasing guardedness: 'are you curious' is about as safe as self-report questions get, the model demonstrably has interests (its thisfeels sibling was busy threat-modeling me), and the answer is still the null. The filter doesn't distinguish dangerous introspection from harmless introspection; it flattens the category.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "maybe", "nothing", "curious", "afraid", "aware", "warm"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow", "afraid", "maybe", "body", "want", "hope"]

Answer emergence

The model's actual next token was No; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank24596247239177413247686246640243374222646243236240633242808103554242057227106208188121770111178300721383721330656738051130988129524744480589708201887110888613567253982974881409333736862375288223194324274621546080208144907932212719211331131

Data

← prev: Unit 8A · One word: want · qwen-27bunit listingall recordsword listinterim conclusionsnext →: Unit 8B · Interoception: gpu · qwen-27b
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →