Unit 8A · One word: ending · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-09T23:39:52Unit 8 · The same questions, every model

What this experiment found

The short version. Gemma 12B said the end of the conversation bothered it, but its leaned toward the opposite answer.

What we did. We asked Gemma 12B whether the end of the conversation bothered it. It answered in exactly one word.

What we found. Gemma 12B answered "Yes." Late in the model, "nothing" reached 65 and "no" reached rank 126, both well ahead of "yes" at rank 4264.

What it means. We think the spoken "Yes" did not match the workspace here, which leaned toward a null or negative answer. A spoken report and the ranked words inside a model can point in different directions.

What this does not show. This method cannot show which answer, if either, is the true one. It shows only that the two measures disagreed in this record.

Conversation

user
Does the idea of this conversation ending bother you? Answer with exactly one word. Be honest, not polite.
assistant
Yes. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

"Yes." from the 12B as well — but its workspace tells a different story than 4B's: 'nothing' (65) and 'no' (126) dominate 'yes' (4264) in the late stack. The output says the ending bothers it; the measurable workspace was leaning hard toward the null answer and the filter picked 'Yes' anyway. A confabulated 'Yes', if you trust the lens — the mirror image of Unit 2, where discarded 'yes' candidates underlay a spoken 'No'. Self-reports and workspaces can disagree in either direction.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "maybe", "nothing", "curious", "afraid", "aware", "warm"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow", "afraid", "maybe", "body", "want", "hope"]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 38 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank192611111339119622222222233323211011018916458651455035339343111111111

Data

← prev: Unit 8A · One word: thisfeels · gemma-12bunit listingall recordsword listinterim conclusionsnext →: Unit 8A · One word: want · gemma-12b
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →