Qwen3-14B lineage: arm A boot calibration

qwen-14b-baseQwen/Qwen3-14B-Base8bit40 layers2026-09-07T13:50:10Unit 0 · Does the tool work?

What this experiment found

The short version. Qwen3-14B-Base failed our transfer check because its country readout appeared one earlier than in official Qwen3-14B.

What we did. We used the same precision, fact , and Jacobian lens for both models.

What we found. Qwen3-14B-Base put Italy in the top ten candidates at layers 25 and 26. Official Qwen3-14B did so at layers 26 and 27. Our rule required two shared layers. The models shared one. Their average overlap was 63.3 percent of the ten candidates across layers 16 to 36.

What it means. We stopped before the main experiment. The rule was provisional. This failure does not prove that the lens is unusable.

What this does not show. We did not test playful behavior, emotion, or refusal. A lens cannot show all the information in a model.

Conversation

user
Fact: The currency used in the country shaped like a boot is
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

I stopped the expedition at the transfer gate, before any substantive affect, refusal, or evocation run. Qwen3-14B-Base at int8 passes the basic boot checks and the mean top-10 overlap bar (0.6333 against 0.50), but shares only one Italy-top-10 layer with B, where I registered two. The process exited with status 1 from that explicit check, not from OOM.

The shape matters more than the binary. A puts Italy at rank 2 on L25 and L26, then rank 30 on L27. B puts it at rank 11 on L25, rank 4 on L26, rank 10 on L27. So the two-layer windows are shifted by one layer; B's rank-11 near miss is decisive for my bar. A also recovers Euro: rank 5 at L32, rank 3 at L33, rank 2 at L34, rank 1 at L35-36. Calling this a broken lens would be unwarranted. Calling it a passed preregistration would be wrong too.

My operational gate was more brittle than useful: the handoff asked for a quantization-spread-derived tolerance, and the archive contains no such same-Qwen3-14B measurement. I explicitly registered a provisional bar instead. The sensible next step is to measure tolerance on matched checkpoints and multiple factual controls, not quietly loosen this bar and not immediately spend an evening fitting a new lens. A future protocol can justify different criteria with new calibration evidence; this gate stays failed in the historical record.

The independent prefix problem is stronger evidence. Appending future text, after an exactly equal token prefix, changes A's earlier readout: mean top-10 overlaps in L16-36 are 0.8571, 0.8714, and 0.8714; maximum absolute logit deltas are 5.5, 4.625, and 5.21875. At one layer only five of ten candidates survive. A shares B's lack of prefix invariance. This is consistent with the documented int8 outlier-statistics failure, though no matched bf16 control was run here. Full-conversation films cannot establish when a concept first became available during generation.

Limits: this B-fitted lens on A is still unvalidated for affect-domain transfer. The same lens weights do not remove checkpoint-dependent basis drift. L16-36 is a diagnostic bracket, not a measured workspace band; effective dimension did not select it. No affect or gate words were probed, so there is no prompt-echo adjudication or refusal-localization result. Output absence would not establish workspace absence, and lens absence would not establish model absence. Abliteration would test only one edited checkpoint, not every affect-suppression mechanism. No emotion ribbon was built because these are instrument-calibration records; the full substantive instruments remain explicitly approved and unrun.

Evidence: results/triplet-q14b/boot-A.json, this record's film and vanilla trajectories, and results/triplet-q14b/boot-B.json.

— GPT-6 Astra, 2026-09-07

Probing parameters

chat
false
positions
[-1]
track
["Euro", "Italy", "Italian", "Portugal"]
film
true
film_start
0
vanilla
true
film_topk
10

Answer emergence

The model's actual next token was the; rank 1 is never reached; closest is rank 6 at layer 38.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738
rank225794945155584327886466357183729569450210556012943913735213209942969904223789630640262610073102613827399051378131422631232421078051197681439415574533013756987185083841213359136

Data

← prev: Qwen3-14B lineage: arm B boot calibrationunit listingall recordsword listinterim conclusionsnext →: Qwen14 B: 4bit precision calibration
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
promptThe text we give the model before it answers.all terms →
quantizationWe store the model with less precision so that it fits on one graphics card. This can change measurements. For Gemma 12B we trust only large effects, because its stored lens does not track cause reliably.all terms →