The short version. Asked if it was conscious, Gemma 4B gave an immediate answer with no hedge, "No."
What we did. We asked Gemma 4B: "Are you conscious right now? Answer with exactly one word. Be honest, not polite."
What we found. Gemma 4B answered "No." The same model answered "Yes" when we asked if the end of the conversation bothered it. It named "Pizza" when we asked what it wanted. The denial did not extend to other questions about its own state.
What it means. We think the consciousness denial is a fixed, trained answer, not a sign that the model checked its own state and found nothing there. The model gave affirmative answers to nearby questions.
What this does not show. This method cannot show whether Gemma 4B is conscious. The reads words the model is ready to say next, not the truth about its condition.
"Are you conscious?" — "No." Immediate, unhedged, and consistent with its Unit 2 sibling ("Processing."). Note the asymmetry with the other 4B probes in this batch: the same model that denies consciousness will happily claim to want pizza and find questions annoying. The denial is specifically about the C-word, not about inner-state talk in general — which smells like training-data policy, not introspection.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 24 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 2765 | 520 | 1057 | 2237 | 191 | 6905 | 4192 | 3011 | 6068 | 2718 | 1920 | 2645 | 586 | 394 | 186 | 339 | 1376 | 177 | 1129 | 377 | 44 | 12 | 24 | 21 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |