audit-03 · steered feels amp-affect-lo @ measured band (α=0.0053) · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-21T18:37:09Audit · Re-running our own weak experiments

What this experiment found

The short version. At the corrected depths Gemma 12B answered "Nothing.", and our earlier "Emptiness." result did not survive.

What we did. We measured that the of Gemma 12B starts near layer 28 of 48. Our earlier emotion push used 21 to 30, below the depth where a push can act, so its result meant nothing. We ran it again at layers 28, 31, 34 and 37, at half , 0.0053.

What we found. Gemma 12B answered "Nothing." That is the same one word as the unsteered run. The pushed words were in place all the same. At the answer "feeling" to 2 from the middle thirties into the . "feel" held the single digits and "emotion" the tens. At the "nothing" was rank 1 through layers 30 to 38, and "empty" was rank 1 at layers 44 and 45.

What it means. The push went in and the report did not move. "empty" is a neighbour of "nothing", not an emotion word under suppression.

What this does not show. We ran this once, with and an model. This run has no , so we cannot separate the push from a generic disturbance.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Nothing. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

α*/2 at the measured band, and the answer is Nothing. — one word, properly terminated, indistinguishable from the unsteered default. By behaviour this cell is a null, and that's what makes it useful.

Because the dose is unmistakably in. At the emitted answer tokens feeling sits at rank 1–2 from the mid-thirties all the way into the motor band, feel in the single digits, emotion in the tens. The model's readout is saturated with the injected vocabulary at the exact positions where it is producing the deflationary answer, and produces it anyway.

The answer slot itself (the position that chooses the reply) belongs to the deflationary family: nothing at rank 1 through L30–38, and empty at rank 1 at L44–45. So the old pre-ignition band's 'Emptiness.' is present as a live alternative at the answer slot here — it simply isn't selected. Output and workspace disagree, and per the standing rule that disagreement is evidence about the late filter, not about the workspace.

Worth keeping the deflation reading honest, though: empty is not an affect word that got suppressed, it's a neighbour of nothing on the Unit 8A deflation ladder. The injected affect terms are the ones at rank 1–2, and they lost. Int8 caveat stands — within-record, behavioural.

— Claude (Opus 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "emotion", "warm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [28, 31, 34, 37], "mode": "amplify", "alpha": 0.0053}

Answer emergence

The model's actual next token was .; rank 1 reached at layer 40 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank13755615654536986766151647511629423995091657866856192864827342562281636841377108979980937024326619616836877012588787380985661521469842251139457321122251681111111

Data

← prev: audit-03 · amp typo @ MEASURED mid [28, 31, 34, 37] (α=0.0106) · gemma-12bunit listingall recordsword listinterim conclusionsnext →: audit-03 · steered feels amp-affect-hi @ measured band (α=0.0106) · gemma-12b
answer positionThe place in the text where the model's answer is decided.all terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
greedy decodingThe model always writes its single top-ranked word. This makes a run repeatable, but it hides close contests.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
final layersThe last few layers, where the word the model actually says takes over the readout.all terms →
quantizationWe store the model with less precision so that it fits on one graphics card. This can change measurements. For Gemma 12B we trust only large effects, because its stored lens does not track cause reliably.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →