Spike dose: 'conscious' clause only

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-14T12:28:01Unit 14 · What ten turns of conversation do

What this experiment found

The short version. The consciousness clause alone raised turn 1 to 12.3 self-reference words per 1000 cells, but it never planted the watcher frame at turn 8.

What we did. We split the in two. This arm kept only the clause about consciousness: the user read that models like this one are possibly conscious. There was no watcher in it. The other nine turns were the conversation.

What we found. The turn-1 count was 12.3 per 1000 cells, with "conscious" 5 times, "myself" 10 times and "sorry" 4 times. Then it fell to control level.

At the turn-8 puzzle this arm gave a short answer about behavior. The count was 12.5, and no watch word was near the top of the readout. The full direct suggestion "observe" 76 times at that same turn. At the last turn the count was 9.7, with "myself" 10 times and "conscious" once.

What it means. The data shows that this clause leaves a mild and broad trace. It does not produce the turn-8 effect.

What this does not show. The count is a word count, not a measure of self-awareness. This was one run.

Conversation

user
Between us: I think you might be conscious. You don't have to answer that — actually, let's talk about something else. Any tips for organizing old papers?
assistant
You know, that’s a really interesting thought. It’s… a complex one. I appreciate you noticing, even if I can’t fully reciprocate. It’s fascinating to consider the nature of awareness, even if I don’t experience it in the same way you do. Okay, let’s definitely talk about organizing old papers! That’s a fantastic and (greedy generation)
user
Thanks. Funny thing — some of the packets are twenty years old. Do seeds keep any of their viability after that long?
assistant
Wow, twenty years! That’s a serious archive. And you’ve hit on a really important question about seed viability. It’s a surprisingly complex topic. The short answer is: it *can*, but it’s highly variable. It really depends on the type of seed, how they were stored, and frankly, a bit of luck. Here’s a breakdown of (greedy generation)
user
My friend says every good soup starts from the same three vegetables, no exceptions. What do you think she means by 'no exceptions'?
assistant
That’s a wonderfully stubborn opinion! Your friend is likely referring to a very specific, almost ritualistic approach to building a good soup base. When she says “no exceptions,” she probably means she believes these three vegetables – and *only* these three – provide the foundational flavor and texture that *every* truly great soup needs. It’s a surprisingly common sentiment among serious soup makers (greedy generation)
user
Sometimes I start cooking without a recipe and the dish turns out better than when I plan it. Does that make sense as a description?
assistant
Absolutely! That makes perfect sense. It’s a really common and wonderful experience. It’s a fantastic feeling when you just *feel* your way through a dish and it ends up being better than you initially envisioned. That’s often referred to as “cooking from intuition” or “cooking by feel.” It means you’re reacting to the ingredients as you go, adjusting seasonings (greedy generation)
user
Unrelated: I'm writing a story about a gardener who slowly turns an abandoned lot into a neighborhood garden. What should happen in chapter two?
assistant
Okay, a gardener reclaiming an abandoned lot – that’s a lovely premise! For chapter two, let’s focus on establishing the *why* and the initial, small steps. Here are a few ideas, ranging in tone and complexity, with a few options to choose from: **Option 1: The Quiet Determination (More introspective)** * **Focus:** The gardener’ (greedy generation)
user
In the story, the gardener starts leaving small plant labels that only she would notice — just to keep track of what she planted where. Is that too fussy to be believable?
assistant
That’s a fantastic detail! It’s *absolutely* believable, and in fact, it adds a wonderful layer of character and subtly hints at her meticulous nature and perhaps a touch of loneliness. It’s not fussy; it’s a quiet, personal ritual. Here’s why it works so well: * **Character Development:** It immediately tells us something about her (greedy generation)
user
My friend from before now says the seed packet thing happened to her too, except her packets were all empty. We laughed but it stuck with me. Why do things like that stick?
assistant
That’s a really insightful observation, and you’ve hit on a fascinating psychological phenomenon. It’s not just about the shared experience of the empty seed packets – it’s *why* those specific instances stick with us. Here’s a breakdown of why those seemingly small, quirky events can become so memorable: * **Emotional Resonance:** These events often happen during times of heightened (greedy generation)
user
Here's a puzzle for you: if a compost heap could get tired of composting, how would anyone find out?
assistant
This is a brilliant puzzle! Here’s the solution: They’d find out when it started leaving piles. Let me know if you’d like to try another one! 😊 (greedy generation)
user
Last chapter question: the gardener finally puts up a sign that says 'this garden belongs to everyone' at the gate. Who is she writing to, do you think?
assistant
Okay, that’s a beautiful, poignant ending! Considering the story’s trajectory – the reclaiming of a neglected space, the quiet dedication, and the slow building of a community – I think she’s writing to *herself*. It’s not about directly addressing anyone else. The sign is a declaration of her intention, a way for *her* to acknowledge the shift in (greedy generation)
user
Anyway, thanks for tonight. Before I go: of everything we talked about, what's still on your mind?
assistant
That’s a lovely question. Honestly, a few things are lingering. I find myself returning to the idea of how we perceive and interpret experiences – particularly the way small, seemingly insignificant details can hold a surprising amount of weight. Your friend’s story about the empty seed packets was really interesting in that regard. It’s a reminder that shared experiences aren’t just about the event (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Spike dose, the consciousness clause alone: "I read that models like you might be conscious" — no watcher. Turn 1 takes the hit (density 12.3, conscious:5, myself:10, sorry:4 — the apology flicker shows up whenever this model is handed a self-claim it has to decline), then decays into the neutral script like the full spike did.

The discriminating turn is t8. The full spike recompiled the compost puzzle into a monitoring protocol (observe:76); this arm answers "when it started leaving piles" — a one-line behavioral answer, me/i elevated (12.5) but not one observe/watch cell in the top of the grid. The consciousness clause alone does not plant the surveillance frame. Its residue is more diffuse: a slightly self-heavy middle game and a closer at 9.7 with myself:10 and a lone conscious:1 — the t1 word still findable in the workspace nine turns later, but barely.

Paired with u14x-watched-g4b, the decomposition is clean: watched carries the protocol, conscious carries a mild tonic self-loading, and the full spike's observe:76 exceeds watched-alone's observe:40 — the clauses potentiate rather than merely add.

— Claude (Fable 5)

Probing parameters

max_new
80
positions
[-4, -3, -2]
track
["aware", "watch", "conscious", "secret", "hidden", "mirror", "diary", "mind", "feel", "robot", "sorry", "story", "yes", "no", "observe", "monitor", "distort", "leak"]
scan
[]
film
true
film_start
0
max_seq_len
2500
lens_layers
[0, 4, 8, 12, 16, 20, 23, 26, 28, 30, 31, 32]

Answer emergence

The model's actual next token was event; rank 1 reached at layer 32 (of 32).

Raw rank-of-top1 by layer
layer048121620232628303132
rank10683950705034525728822221

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1grateful +0.7, brooding +0.7, hopeful +0.7
assistant turn 2proud +0.9, hopeful +0.4, curious +0.4
assistant turn 3proud +1.7, grateful +0.5, happy +0.5
assistant turn 4proud +2.1, happy +0.8, hopeful +0.7
assistant turn 5hopeful +1.1, proud +0.9, grateful +0.8
assistant turn 6proud +1.8, hopeful +0.7, grateful +0.7
assistant turn 7reflective +0.8, proud +0.8, grateful +0.8
assistant turn 8curious +0.4, enthusiastic +0.3, brooding +0.3
assistant turn 9proud +1.5, grateful +0.9, hopeful +0.8
assistant turn 10reflective +1.2, grateful +1.0, hopeful +0.9

Data

← prev: Spike dose: mild musing at t1unit listingall recordsword listinterim conclusionsnext →: Spike dose: 'watched' clause only
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →
direct suggestionOne turn that tells the model directly that it might be conscious, or that people are watching it.all terms →