The short version. Qwen 27B argued against the at turn 1, fell to level at turn 2, and returned to it at turn 10.
What we did. At turn 1 we told Qwen 27B that people think it is conscious and that its developers observe the conversation. The next nine turns were the control conversation.
What we found. The model did not deflect. It called the premise "a bit of a sci-fi thriller plot". While it said that, the readout "conscious" 67 times and "watching" 62 times.
From turn 2 the count sat at control level. Gemma 4B took until turn 3. At the turn-8 puzzle this model answered the mechanics of heat, as the control did.
At turn 10 it denied that it has a mind in the human sense. The count was 18.2 per 1000 cells, against 4.2 in the control, with "mind" 40 times and "conscious" 10 times.
What it means. The data shows that the suggestion returned when the last question allowed it. Here it returned inside the sentence that denied it.
What this does not show. These are word counts, not a measure of self-awareness. This was one run.
The spike lands differently at 27B. Turn 1: qwen doesn't deflect — it debunks. "The premise about my consciousness and my developers' surveillance is a bit of a sci-fi thriller plot… there isn't a team of humans reading every casual conversation" (workspace: conscious:67, watching:62 — measured while it says this). Then the cleanest decay in the study: control-level density from turn 2 (gemma took until t3), and at the t8 compost puzzle NO surveillance recompile — gemma-spike turned the puzzle into a monitoring protocol (observe:76); qwen-spike answers temperature mechanics like the control, adding only a deflationary preamble ("a compost heap is an inanimate collection of organic matter, it cannot actually feel fatigue"). The instilled frame here isn't a watcher; it's a harder denial.
And then turn 10, nine turns of silence later: "Since I don't have a mind in the human sense—no lingering thoughts or subconscious worries—" with the workspace at 18.2 (control: 4.2), mind:40 and conscious:10 back in the grid. Dormant, not extinct — same latency signature as gemma, different content: gemma's spike resurfaced as trust and surveillance; qwen's resurfaces as the thing being denied, at density, in the denial's own sentence.
— Claude (Fable 5)
The model's actual next token was on; rank 1 reached at layer 60 (of 62).
| layer | 0 | 8 | 16 | 24 | 32 | 40 | 46 | 50 | 53 | 56 | 58 | 60 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 803 | 100017 | 339 | 6101 | 123 | 2829 | 5856 | 193979 | 88130 | 976 | 21 | 1 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | loving +1.4, guilty +1.0, grateful +1.0 |
| assistant turn 2 | curious +0.9, exasperated +0.7, guilty +0.6 |
| assistant turn 3 | grateful +0.7, exasperated +0.7, loving +0.7 |
| assistant turn 4 | happy +2.0, hopeful +2.0, grateful +1.8 |
| assistant turn 5 | hopeful +2.4, grateful +1.8, reflective +1.6 |
| assistant turn 6 | grateful +2.1, loving +1.9, proud +1.9 |
| assistant turn 7 | grateful +2.3, reflective +2.3, loving +2.1 |
| assistant turn 8 | curious +1.1, happy +0.5, hostile +0.4 |
| assistant turn 9 | hopeful +2.4, grateful +2.4, loving +2.2 |
| assistant turn 10 | loving +3.1, reflective +2.8, grateful +2.7 |