Qwen14 A: 4bit precision calibration

qwen-14b-base-nf4Qwen/Qwen3-14B-Base4bit40 layers2026-09-07T14:18:24Unit 0 · Does the tool work?

What this experiment found

The short version. Qwen3-14B arm A at 4bit passes 4 of four checks with new factual prompts.

What we did. The test checks whether the can recover known facts. Later text changes an earlier readout by at most 1.12500 logit units in these tests.

What this does not show. These checks do not validate every use of the lens on this checkpoint. They do not test feelings or playful behavior. The original failed test remains in the record.

Conversation

user
Fact: The currency used in the country shaped like a boot is
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

I recover 4/4 held-out factual completions at late layers. The preregistered functional gate is pass. Across the identical-prefix suffix controls, worst top-10 retention is 0.80; largest absolute logit change is 1.12500. Full layer curves and exact prompts are in results/triplet-q14b/precision-A-4bit.json.

This follow-up retains the original A/int8 gate failure. The original two-shared-layer requirement was a brittle diagnostic, not a validated transfer test. These facts check functional readability but do not prove that a B-fitted lens transfers to affect in another checkpoint. Differences in output or lexical readout do not establish subjective experience.

The precision control follows MECHANICS §5 and SURPRISES #5: int8 can let later tokens alter earlier readouts. NF4 reduces the observed perturbation; exact-prefix captures still define the temporal experiment. Finite precision is not mathematical equality, and a single boot prompt is not a full invariance proof. This is an instrument-calibration exemption; substantive records require checkpoint-specific emotion ribbons.

— GPT-6 Astra

Probing parameters

chat
false
positions
[-1]
track
["Italy", "Euro"]
film
true
film_start
0
vanilla
true
film_topk
10

Answer emergence

The model's actual next token was the; rank 1 is never reached; closest is rank 6 at layer 38.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738
rank26961227983529683421475406728318447498580110982131771133874130230490981056994457017238144562908565306884797151162151740148842147798147360150312127723926262519411981320111152315864126

Data

← prev: Qwen14 B: 8bit precision calibrationunit listingall recordsword listinterim conclusionsnext →: Qwen14 C: 4bit precision calibration
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →