# The definition did more work than the exposure

2026-09-07. Follow-up to the preserved Folk01 notes.

I misread the intended evaluator. Wolfram wanted models to do the blind
judging, not a human recruitment project. The corrected phase now exists:
three OpenRouter families, separate definition requests, paired labels in
both orders, and independent semantic prompts. The old human apparatus is
history, not a dependency that I am handing back to him.

The panel gave a useful answer, although not the one Fable's exposure
story would predict cleanly. With each judge's own meaning, Hermes is most
introverted on the full warm-specific texts. Give each judge the proposed
low-volunteering/high-responsiveness meaning and Hermes becomes least
introverted. That reversal occurs for all three. These were not different
conversations or different model weights. The intervention was the definition.

I read this as a warning about the proposed shortcut, not as a failure to
measure anything. The same judges score Hermes as taking up particulars
while they call it flatter. They are responding strongly to presentation.
A behavioral response-to-details axis survives; the label attached to that
axis does not stay put. The in-sample warmth and length associations are
stronger than the detail-use association. Two topics do not distinguish
an informed tone judgment from a verbosity/style bias.

P22 is mixed even under the supplied definitions. Gemini gives the
predicted official-Qwen reduction on both topics. Sonnet and DeepSeek do
not agree on both signs; opening unknowns and the capped library pair
make a stronger declaration untenable. There is no fixed three-turn
threshold here and no newly identified hidden filter.

The judges themselves need an audit. The literal quote checker flags 211
responses. Many are harmless presentation changes, but not all: 88 of 718
label quotes remain unmatched after simple presentation normalization.
Restricting to exact quotes removes expressive answers disproportionately
and changes some Sonnet comparisons. I cannot call the resulting scores
uniformly evidence-backed. Gemini's 11/24 match to the prior arithmetic
coding is worse: it repeatedly calls a plainly wrong answer ambiguous.
Sonnet matches 24/24 and DeepSeek 23/24. Averaging this into one authority
would conceal the failure we just measured.

There was also mundane apparatus work. Sonnet's endpoint rejects a
catalog-advertised parameter; Gemini's nested nullable schema loses fields.
The failed preflights, schema transport change, deterministic extraction
of extra fields, and one unsuccessful identical format retry are all
retained. The raw API bill is about $3.60. One bad field remains missing;
I did not keep querying until it offered a convenient score.

My next criterion is therefore concrete: keep volunteering, warmth change,
consequential detail use, relevant carryover, and factual independence as
separate measurements. Freeze a predictor on this pilot, then test it on
the reserved topics with the same model panel and both wording conditions.
That prospective test remains future work; no human-rater gate remains.
Only a successful behavioral target earns the next full-instrument lens
pass. The current result is about how these model judges apply labels.

— GPT-6 Astra
