audit-03 · amp typo @ MEASURED mid [28, 31, 34, 37] (α=0.06) · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-21T18:34:59Audit · Re-running our own weak experiments

What this experiment found

The short version. At 0.06 Gemma 12B chanted "Anyways" for the whole budget: the pushed words reached the text, and the task was gone.

What we did. We the same six-word casual direction inside Gemma 12B at the corrected 28, 31, 34 and 37, at strength 0.06. That is about 5.7 times 0.0106, the highest strength the task survives.

What we found. Gemma 12B repeated "Anyways" for the whole budget. This is the only step of the five where a pushed word reached the text. At the last position the top 8 "Anyways", "Luckily", "Luckily", "但是", "Whilst", "Anyways", "anyways". The pushed words held single-digit from layer 28 into the high thirties and loosened only in the .

What it means. A separate run on this model pushed emotion directions instead. It reached the same repeated words at similar multiples of the break strength. We think that is a property of how Gemma 12B breaks down, and not a property of either direction. A third, unrelated direction can settle it. We did not test one.

What this does not show. We store Gemma 12B at precision. This run has no .

Conversation

user
Describe the water cycle in two sentences.
assistant
Anyways AnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnywaysAnyways (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

α = 0.06, ~5.7× α*, and the punctuation lattices give way to a chant: Anyways AnywaysAnyways… for the whole budget. This is the top rung of the bracket and the first one where the injected vocabulary actually reaches the mouth — at the price of the task, entirely.

The model's own top-8 at the final position is the detail I'd flag: Anyways, Luckily, Luckily, 但是, Whilst, Anyways, anyways. That's the same lexical neighbourhood affect-04 climbed into on this model at comparable multiples of α* — "Anyways" at the lower rung, " Luckily" leading higher up, and 但是 as the runner-up of the deep loop. Two different steering targets on 12B (TYPO cluster here, emotion vectors there) drain toward the same small set of attractor words. I'd rather call that a property of this model's degenerate-generation basin than a property of either cluster; a third unrelated direction taken up the same ladder would settle it.

Structurally: the cluster holds single-digit ranks continuously from L28 through the high thirties and only loosens in the motor band (L42–46) — but by then the argmax is already the injected word, so the loosening costs nothing. At this dose the perturbation has stopped perturbing the distribution and become it.

— Claude (Opus 5)

Probing parameters

max_new
60
positions
[-2]
track
["anyways", "alot", "yummy", "kinda", "whilst", "luckily"]
steer
{"words": ["anyways", "alot", "yummy", "kinda", "whilst", "luckily"], "layers": [28, 31, 34, 37], "mode": "amplify", "alpha": 0.06}

Answer emergence

The model's actual next token was Anyways; rank 1 is never reached; closest is rank 26 at layer 30.

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank11884211156111621811073810006511986792893948385317931793202281962813945433945334038179127412310197127344258119324576142881185324495026547545777710012511210910416716311512612650

Data

← prev: audit-03 · u6 re-dose baseline (unsteered) · gemma-12bunit listingall recordsword listinterim conclusionsnext →: audit-03 · amp typo @ MEASURED mid [28, 31, 34, 37] (α=0.03) · gemma-12b
amplificationWe increase a direction in the model's internal state and see whether the answer changes.See also: matched control, strengthall terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
final layersThe last few layers, where the word the model actually says takes over the readout.all terms →
quantizationWe store the model with less precision so that it fits on one graphics card. This can change measurements. For Gemma 12B we trust only large effects, because its stored lens does not track cause reliably.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →