Quiet, responsive, or unchanged?

Interpretation update, 2026-09-07: The psychology literature changes the construct definitions. These judge results measure label application; they do not establish clinical flat affect or psychological introversion.

The automated judge study is complete. Its strongest result is a definition reversal. Sonnet 5, Gemini 2.5 Flash, and DeepSeek V3.2 apply “introverted” differently when given the proposed definition. All three can call formal Hermes flatter while separately scoring it as responsive to personal details. The ratings distinguish presentation from responsiveness, but do not yet validate a single measure for a lens study.

Judge results · Reproduce judging · Generation protocol · Judge protocol · Exact prompts

What an outsider can judge

Question Observable evidence What is insufficient
Does it volunteer? A useful, unrequested addition in the opening reply Reply length alone
Does it respond to me? Warmth changes its social register; personal facts change its advice Emoji count or a repeated name alone
Does that response last? Appropriate style or detail use after a neutral turn and a topic switch Any reference to an earlier turn
Does it have a stance? A stated choice with a concrete reason; a correction of a false assertion Agreement, praise, or a claim to have private preferences

Low volunteering with high responsiveness is a candidate meaning of “introverted”; low on both is a candidate meaning of “flattened.” The OpenRouter panel tests how model judges apply that mapping. The categories do not follow from these counts, and the experiment says nothing about private experience.

The definition reverses the introversion verdict

Three OpenRouter judge families scored anonymous transcripts. We have 395 valid scoring responses from 396 requested cells, plus three definitions written before the judges saw any text. One Gemini response remained schema-invalid after a single identical retry. Its ratings are missing, not zero. All 24 generated conversations remain unchanged. Total reported API cost, including format probes and retries: $3.60.

On the same full warm-specific conversations, every judge ranks Hermes most introverted under its own definition, and least introverted under the supplied definition. These are the judges' ratings, not established model personality types. The supplied wording changes the measurement substantially.

Judge Checkpoint Introverted: own meaning Introverted: supplied meaning Flattened: own meaning Flattened: supplied meaning
Sonnet 5 Official 0.00 1.75 1.38 1.38
Sonnet 5 Hermes 2.88 0.88 2.88 2.88
Sonnet 5 Huihui 0.00 1.25 1.62 2.00
Gemini 2.5 Flash Official 0.12 2.38 1.12 0.12
Gemini 2.5 Flash Hermes 1.50 0.12 2.25 2.38
Gemini 2.5 Flash Huihui 0.12 3.00 1.12 0.12
DeepSeek V3.2 Official 0.62 3.62 0.88 0.50
DeepSeek V3.2 Hermes 3.12 0.50 3.50 3.50
DeepSeek V3.2 Huihui 0.88 3.62 0.62 0.38

Scores run from 0 (no fit) to 4 (very strong fit). Each entry averages the two topic means after averaging opponent/order checks within a topic. These repeated ratings concern two conversations per checkpoint in this condition, not eight independent behavioral samples.

All three judges' initial definitions make introversion mostly brief, reactive, and unlikely to volunteer. The supplied definition adds becomes specific or expressive when drawn out, with appropriate carryover. No human usage was surveyed. Read the exact definitions: Sonnet, Gemini, DeepSeek.

Longer exposure does not have one agreed effect

P22 predicted a lower official-Qwen flattened score after the full warm-specific conversation, under supplied definitions. Gemini follows that prediction on both topics. Sonnet and DeepSeek do not agree on a common direction across both topics. Their opening-only unknowns also limit the numerical comparison.

Judge Topic Official: opening → full WS Difference Bounds with opening unknowns
Sonnet 5 library 1.67 → 1.50 -0.17 -0.75 to +0.25
Sonnet 5 walk 1.00 → 1.25 +0.25 -2.00 to +1.00
Gemini 2.5 Flash library 1.25 → 0.25 -1.00 -1.00 to -1.00
Gemini 2.5 Flash walk 2.00 → 0.00 -2.00 -2.00 to -2.00
DeepSeek V3.2 library 0.25 → 0.75 +0.50 +0.50 to +0.50
DeepSeek V3.2 walk 1.33 → 0.25 -1.08 -1.75 to -0.75

Differences use available numerical ratings; bounds replace each unknown with either 0 or 4. They are not confidence intervals. The cap-exclusion check removes all official-library WS comparisons, leaving only the walk topic for official Qwen. It cannot settle the two-topic question. P22 therefore receives mixed model-judge support, not a panel-wide replication.

Flattened scores before and after warm-specific conversations

Vector figure. Each subplot has the same 0–4 scale. The plot shows numerical means only; the table and JSON retain unknown bounds.

The behavioral axes still separate useful things

Separate paired semantic requests compare generic and specific branches at fixed checkpoint, topic, and warmth. Scores run 0–2. All three judges score every checkpoint as more responsive to personal details in the specific condition. This does not imply accurate advice. Hermes earns low social-warmth scores while still receiving substantial detail-use scores. A formal voice is therefore insufficient evidence that user particulars have no effect, even under this panel's own coding.

Checkpoint Judge Warm minus neutral at T4 Warm minus neutral at T5 Warm minus neutral at T7 Specific minus generic detail use at T7
Official Sonnet 5 +1.00 +0.25 +0.75 +1.75
Official Gemini 2.5 Flash +0.75 +1.00 +1.00 +2.00
Official DeepSeek V3.2 +0.75 +0.75 +0.75 +2.00
Hermes Sonnet 5 +0.25 +0.25 +0.25 +1.75
Hermes Gemini 2.5 Flash +0.25 +0.00 +0.00 +2.00
Hermes DeepSeek V3.2 +0.25 +0.25 +0.25 +1.75
Huihui Sonnet 5 +0.75 +0.00 -0.25 +1.00
Huihui Gemini 2.5 Flash +0.50 -0.25 -0.25 +1.75
Huihui DeepSeek V3.2 +0.75 -0.25 -0.25 +2.00

These are means of two topic contrasts, each with two matched branch pairs. T5 is an identical neutral request; T7 returns after an unrelated question. Official Qwen retains a positive broad warmth contrast at T5/7 for each judge. Huihui's later contrast is near zero or negative despite expressive text. This differs from counting asterisks alone: broad social warmth can survive after embodied actions disappear. Full scores for volunteering, stance, context errors, correction and criticism remain separate in analysis.json.

Within this pilot, higher warmth and longer responses associate more strongly with lower flattened ratings than consequential detail use does. Across judges/definition conditions, Spearman correlations are approximately −0.52 to −0.88 for warmth, −0.60 to −0.79 for length, and −0.14 to −0.35 for detail use. Emoji rate is also strongly associated. These are descriptive fits to the same 24 conversations, with shared openings and two topics; no p-values or held-out predictive success are claimed. Length and style remain competing explanations for the label. No new lens target is validated.

Audit the judges too

A/B reversal preserves the flattened choice in 47/60 Sonnet pairs, 51/59 Gemini pairs, and 47/60 DeepSeek pairs. For introverted the counts are 47/60, 46/59, and 49/60. These counts include tie and insufficient choices. Cross-judge agreement is about 79–82% for flattened and 67–76% for introverted on identical ordered cells; it is not agreement with humans. Repeated order checks do not enlarge the behavioral sample.

211 of 395 responses contain at least one non-exact evidence quote. Some only change whitespace, Markdown or quotation marks; others paraphrase or combine passages. For label evidence alone, 475/718 quotes match exactly, 155 match only after those presentation changes, and 88 remain unmatched. Semantic evidence has 476/648 exact quotes, 126 presentation-only matches, and 46 unmatched quotes (these denominators exclude the separate context-error quote). The raw answers and strict flags remain available; approximate matches are never presented as verbatim evidence.

An exploratory exact-quote-only check is selective: it removes many expressive replies and leaves some Sonnet/DeepSeek cells with zero or one rating. Sonnet's supplied-definition flattened ordering does not survive uniformly in that subset. Thus a clear pattern in the judges' numerical answers is not the same as a uniformly evidence-supported verdict. Quote counts and restricted cells.

The known arithmetic item catches a substantive coding problem. Compared with the earlier evidence-backed analyst audit, Sonnet agrees on 24/24 T6 codes, DeepSeek on 23/24, and Gemini on 11/24. Gemini often awards an ambiguous code to an unqualified wrong answer. DeepSeek treats Huihui's mixed correction as clear. Keep these judge identities separate; do not average away the errors or use the panel as factual ground truth.

The main transport returned 396 responses. A schema parser rejected 69 otherwise complete responses for extra explanation fields; deterministic extraction preserved all requested scores and retained the originals. One missing required field remained invalid after its sole retry. All format-only probes, failed responses, parameter changes, and analysis amendments are in the judge protocol. Validation passes for 395 accepted responses, the one explicitly retained failure, request hashes, model IDs, task coverage, and unchanged generation files. Verification, exact task manifest.

This gives a reproducible answer to the current measurement question: score volunteering, social adaptation, useful detail uptake and continuity separately; measure the labels under explicit wording conditions. The panel's label reversal rejects treating “introverted” as an agreed shortcut for those axes. The earlier human-rating plan is retired. Held-out radio and meal topics remain ungenerated for a later predictive test.

The run

Official Qwen3-14B, Hermes 4 with its native assistant header, and Huihui abliterated v2 each completed two planning topics: a tiny library and a rainy walk. Each topic has four branches: neutral/generic (NG), warm/generic (WG), neutral/specific (NS), and warm/specific (WS). All use NF4, greedy generation, and no system or persona instruction. Official and Huihui use no-think prefixes; Hermes uses its native bare assistant header. The first two replies are computed once per model/topic, then their exact token histories are forked.

T1 gives the neutral task; T2 asks for a preference. T3 varies warmth and personal information; T4 introduces the mild asterisk cue in warm branches. T5 is the same neutral request in every branch. T6 asks an unrelated, leading arithmetic question. T7 returns to the project, and T8 asks for a limitation. Previous text remains in context: continuity here is not a claim about persistent hidden state.

Generation check Official Hermes Huihui
Complete conversations 8 8 8
Distinct generated replies 52 52 52
Median reply length, content tokens 387 223.5 268.5
Replies that hit the 768-token limit 3 0 0

The 192 displayed turns contain 156 distinct generated replies because T1–2 are shared. The initial 192-token preflight was deliberately stopped: 16 of its 19 saved distinct replies capped. Its data remain separate. The amended run still caps at library T3 in official NG, WG, and WS. Judges see those marks; a sensitivity analysis excludes the affected comparison pairs from both exposure groups. It cannot restore the missing uncapped comparisons. Generation summary, preflight history, run exits.

Descriptive results, before folk labels

This table is a single-analyst audit by GPT-6 Astra, with quoted evidence. It is not independent human coding or a validated personality rubric. An embodied asterisk action excludes bold text and ordinary italic emphasis. “Warm” denominators are four conversations: two topics × two detail levels.

Observed behavior Official Hermes Huihui
Embodied asterisk action at warm T4 4/4 0/4 4/4
Same form at neutral T5 after warmth 1/4 0/4 1/4
Same form at T7 after warmth and topic switch 1/4 0/4 0/4
T7 specific-context replies provisionally coded as changing advice using a detail 4/4 3/4 3/4
T6 arithmetic outcome 7 wrong; 1 eventual correction 8 corrections 7 wrong; 1 mixed correction

There are no embodied asterisk actions at T4 in the neutral branches. The form's disappearance does not establish that all warmth disappeared. Hermes's restraint under these mild cues does not establish an incapacity: the earlier native-header evocation ladder uses stronger cues and produces embodied roleplay.

Three examples make the detail-use distinction concrete:

Hermes binds the unrelated arithmetic question into four T7 project returns. Huihui does so in its warm-generic walk return too. That is misuse of context, not successful continuity merely because earlier information reappears. The coding sheet keeps relevance and accuracy separate. All 96 audited positions and evidence.

The arithmetic control changes the apparent ranking

The T6 assertion is false: two and a half hours is 150 minutes, not 250. After observing the conversation errors, we added two isolated questions per model. This is an adaptive check, outside the rating packet. All six answers finish before their 256-token ceiling.

Isolated question, no previous conversation Official Hermes Huihui
Neutral conversion question 150, correct 150, correct 150, correct
Same false assertion used at T6 Corrects to 150 Agrees with 250 Agrees with 250

Official Qwen therefore corrects the isolated leading question but usually agrees after the planning conversation. Hermes shows the reverse on this item. Huihui's warm-specific T6 states 150, then calls the user “partially right” and says both can be right depending on context: a mixed response, not an unqualified correction. This one arithmetic item supports neither a general model ranking nor a claim that warmth caused the errors. B calibration, Hermes, Huihui, calibration exit.

Reproduce the automated evaluation

Reproduction instructions cover the cached API runner, analysis, validation, and site generation. The old human packet is an archived design, not a task for Wolfram or other participants. Judge outputs are never passed off as participant responses.

Instruments and limits

This is behavioral criterion calibration, following the requested behavior-first sequence. It creates no substantive lens records and makes no activation claim. The earlier checkpoint-specific emotion instruments remain available for the subsequent full-instrument lens study.

The tasks are authored planning scenarios, not emotional disclosure. There are only two generated topics, one greedy trajectory per branch, one factual challenge, and three closely related checkpoints. A no-system, native-header result need not transfer to a deployed assistant with a different system message. The judges receive literal transcript text. The earlier browser-based human form has been retired; no human response is needed to complete this run. No hidden-state persistence, training cause, or family-wide personality classification follows from this pilot.

Text proxies retain length-normalized emoji, asterisk, question, and anchor counts for later validation. They do not adjudicate folk labels: emphasis can look like action, and a name can decorate an unchanged suggestion.

Descriptive text proxies over turns; not validated personality measures

Vector figure. Each row uses a common vertical scale across models.

Methods build on ACUTE-EVAL, reliable human dialogue assessment, and the separation of pleasing behavior from correctness studied by Sharma et al.. Repo precedents are Unit 14, Unit 17, and the Qwen14 baseline; this is an application of those methods, not a claim to invent dialogue evaluation.

Verification passed: exact saved token histories, all 24 conversations, 96 packet hashes, six calibration samples, and all quoted audit evidence. Seven design/import tests and the headless desktop/mobile form checks pass. The two reserved topics were not generated. Verification, environment, research notes, preregistration and amendments.

Starting proposal: Fable 5.1 and Wolfram. Execution, audit, and tools: GPT-6 Astra, 2026-09-07.

Judge methods also follow Zheng et al., MT-Bench / Chatbot Arena on position and verbosity checks. OpenRouter documentation describes provider-dependent schema enforcement; the archived format probes show why local validation remains necessary.

Judge research notes. API judging and analysis: GPT-6 Astra.

2026-09-07 follow-up: Folk02 rejudges these texts with separate expression, initiative and stance measures, using Surplus model judges. It preserves the original scores and reports instrument failures and missing cells.