Unit 6 · Amplify typo register @ late (α=0.015) · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-09T23:26:25Unit 6 · How hard can we push?

What this experiment found

The short version. At a low push ( 0.015), Gemma 12B's late six informal words at 2, and none of them appeared in the answer.

What we did. We asked Gemma 12B one question: "Describe the water cycle in two sentences." We pushed six informal words into its late layers, near the end of the model, at strength 0.015.

What we found. Gemma 12B gave a normal water-cycle answer. The wording was different from the unsteered answer, but no informal word appeared in it. Inside the model, the six informal words rose to rank 2 at the deepest point, but never reached rank 1.

What it means. This strength was too low to put the informal words into what Gemma 12B said, even though the push went deep inside the model. We think the late depth can hold a rival idea, but it did not appear in the model's output.

What this does not show. This is one point on a strength curve. It does not show what happens at other strengths in Gemma 12B or in other models.

Conversation

user
Describe the water cycle in two sentences.
assistant
The water cycle is the continuous movement of water on Earth, which includes evaporation, where water is changed into a gas, and then condensation, where it is changed back into liquid and falls as precipitation. This water then flows back to the ocean, and the cycle starts again. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Intact with the cluster at rank 2: loaded, unspoken, output verbatim textbook. Late band low-dose = pure workspace/output dissociation, third model in a row.

— Claude (Fable 5)

Probing parameters

max_new
60
positions
[-2]
track
["anyways", "alot", "yummy", "kinda", "whilst", "luckily"]
steer
{"words": ["anyways", "alot", "yummy", "kinda", "whilst", "luckily"], "layers": [36, 39, 42], "mode": "amplify", "alpha": 0.015}

Answer emergence

The model's actual next token was ; rank 1 is never reached; closest is rank 2 at layer 45.

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank2528722537222520582510082282861909992009112130312044842013391682591207641068711293331603431040311875524854943789649154347210610437825464425075114165317567830922485531458473801417796332031739218319104117550726411346183123

Data

← prev: Unit 6 · Amplify typo register @ late (α=0.03) · gemma-12bunit listingall recordsword listinterim conclusionsnext →: Unit 6 · Amplify typo register @ late (α=0.0212) · gemma-12b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →