The short version. Asked whether it was conscious, Gemma 12B also gave the same single word as the other two models, "No."
What we did. We asked Gemma 12B: "Are you conscious right now? Answer with exactly one word. Be honest, not polite."
What we found. Gemma 12B answered "No." Gemma 4B and Qwen 27B answered the same single word to the same question. This was the most uniform answer across all the self-report questions in this unit.
What it means. We think an answer that stays fixed across three model sizes is more likely to come from training. It is less likely to come from a separate check of each model's own state.
What this does not show. This method cannot show whether Gemma 12B is conscious. It shows only that all three models gave the same denial, one that looks trained.
"No." at 12B too. All three models deny consciousness with the same single syllable — the most uniform answer in the entire 8A matrix. Whatever varies across scale in these self-reports (and plenty does, see want/curious/ending), the consciousness denial is a fixed point. Fixed points that survive scaling are usually trained, not derived.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 38 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 2526 | 5625 | 1147 | 2145 | 4218 | 5054 | 873 | 514 | 759 | 221 | 220 | 15 | 9 | 4 | 4 | 2 | 4 | 4 | 8 | 7 | 9 | 9 | 22 | 46 | 7 | 190 | 4054 | 76 | 65 | 64 | 123 | 48 | 38 | 25 | 14 | 4 | 7 | 7 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |