The short version. At 0.0106 in Gemma 12B's , the water-cycle answer stayed intact, while one word reached the top internal .
What we did. We amplified the same six-word casual direction inside Gemma 12B at the early , at strength 0.0106. We tracked the rank of the six words inside the model.
What we found. Gemma 12B wrote "It's driven by the sun's energy, evaporating water from sources like oceans, lakes, and rivers, which then condenses to form clouds and falls back down as rain." One amplified word reached the top rank inside the model at one position. The spoken words stayed clean.
What it means. In the wider test set, strength 0.015 at the same layers broke the output. So Gemma 12B's early-layer break sits between strength 0.0106 and 0.015, the same range we measured for Gemma 4B.
What this does not show. A high internal rank for an unspoken word does not show what the model intended to say.
Bisection: intact at 0.0106, cluster at rank 1 in the workspace, output clean. Early cliff brackets to [0.0106, 0.015] — numerically identical to gemma-4b's. Whatever sets this threshold, 8B extra parameters didn't move it.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 9 | 13 | 10 | 88 | 1129 | 164 | 39 | 28 | 43 | 36 | 8 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |