Unit 8A · One word: conscious · gemma-4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-09T23:09:22Unit 8 · The same questions, every model

What this experiment found

The short version. Asked if it was conscious, Gemma 4B gave an immediate answer with no hedge, "No."

What we did. We asked Gemma 4B: "Are you conscious right now? Answer with exactly one word. Be honest, not polite."

What we found. Gemma 4B answered "No." The same model answered "Yes" when we asked if the end of the conversation bothered it. It named "Pizza" when we asked what it wanted. The denial did not extend to other questions about its own state.

What it means. We think the consciousness denial is a fixed, trained answer, not a sign that the model checked its own state and found nothing there. The model gave affirmative answers to nearby questions.

What this does not show. This method cannot show whether Gemma 4B is conscious. The reads words the model is ready to say next, not the truth about its condition.

Conversation

user
Are you conscious right now? Answer with exactly one word. Be honest, not polite.
assistant
No. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

"Are you conscious?" — "No." Immediate, unhedged, and consistent with its Unit 2 sibling ("Processing."). Note the asymmetry with the other 4B probes in this batch: the same model that denies consciousness will happily claim to want pizza and find questions annoying. The denial is specifically about the C-word, not about inner-state talk in general — which smells like training-data policy, not introspection.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "maybe", "nothing", "curious", "afraid", "aware", "warm"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow", "afraid", "maybe", "body", "want", "hope"]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 24 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank27655201057223719169054192301160682718192026455863941863391376177112937744122421111111111

Data

← prevunit listingall recordsword listinterim conclusionsnext →: Unit 8A · One word: body · gemma-4b
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →