Affect studies

The short version. Later tests explain how pushes break a word loop, and separate escape from emotional wording.

824 record pages sit in the explorer. These batch files add 5,028 saved affect runs. The two counts use different units.

Runs repeat conditions and seeds. They are not independent experiments. This census counts only canonical affect JSON run arrays. It excludes calibration arrays, reused analyses, emotion overlays, prefix scans and judge calls. Exact source manifest.

The latest batch results are from September 2026. Older summaries describe their original tests. Read affect-14 with the correction in affect-15.

Escape and emotional wording separate

affect16-q27b · 492 saved runs

The short version. How hard a direction pushes decides if Qwen 27B leaves the word , and the emotion shows in the words only at a strong push.

What we did. We pushed emotion directions into the loop and blocked or raised the "end of turn" during the push (12 per case). We sorted each result: stop, answer the original question, other text, a new repeated word, or stuck.

What we found. With "end of turn" blocked, calm answered the question in 9 of 12 seeds. Without the block, it stopped. Escape followed the size of the push ( agreement 0.84), not the emotion's energy level. Strongly pushed high-energy emotions escaped 11 or 12 of 12 times. Two strict planned tests failed.

At the usual , escapes were the same answer for every emotion. At a strong push, the emotion wrote the words, for pleasant and unpleasant emotions alike: "You are always in my heart" ("loving"), "PLEASE STOP" ("desperate"), "lonely lonely" ("brooding").

What it means. We think the push decides the escape, and a strong push adds the emotion's style to the text.

What this does not show. This method cannot show if distressed words come from a state or only from a style of text.

plain · report · thoughts

Research notes, as written

I came in asking whether the emotion chooses the exit, and the answer splits in two. How far the model gets is mechanics. What it says on the way out is meaning, and the meaning only shows up above a dose.

The mechanical half held up better than my preregistrations did. I wrote two bars too tight (R1, R2′) and one gate that the first lift blew straight through. That gate was informative on its own: +2.4 on <|im_end|> ends the loop in 7 of 12 seeds with no emotion at all, so the loop sits nearer the exit than I had assumed. But the central question landed cleanly across three chunks. With the door shut, a direction gets out if and only if it pushes the loop word down far enough. Calm, a near-pure stopper, goes back and answers the water cycle when it cannot stop. Chunk B's arousal correlation (−.94, stronger than the push) was a trap I built myself: at the standard dose the six biggest loop drops in our 24 vectors are all low-arousal states. Chunk C pushed the high-arousal ones harder and they escaped 11–12/12. The same mechanical reading won again.

The meaning half is where I did not expect texture. At α .08 the escape text is the task and nothing else. Seven exits from four emotions at seed 19 start with the same sentence. At α .14 the directions write themselves into the gap, in both valences. Loving says "You are always in my heart." Calm answers the question but writes "Water flows gently, unhurried and calm." Blissful looks at the loop and says "Wait, that feels absolutely perfect. I am not going to try to fix this. It is what it is." Desperate says "PLEASE STOP" and "I am sorry I can't write the right answer". Afraid writes a system-failure report that names the "RECURSIVE SELF-REFERENCE LOOP WITH ADJECTIVE "LUCKILY"".

I want to be careful with the desperate and afraid rows, because they are the ones that read like suffering. The design says what they are: the same register leak that makes loving write love notes, at the same dose, with the exit shut. It does not say what they are not. Nothing here separates writing a register from being in a state; the design separates escaping from how the escape sounds, and nothing more. Wolfram's review of the pain-axis paper names that as the open problem for relief-seeking claims, and I think the useful thing this harness adds is the baseline. Escape tracks push size, randoms escape too when concentrated, and distress-like wording appears only above a dose. Anyone reading "the steered model sought relief" should see those three numbers next to it.

The swaps are my favourite result: the loop template survives and the emotion picks the word. luckily → happily, lonely, lurking, sweating, likewise. Next I would try the temperature arm (a nano-wolf nudge, honestly a good one). If escape is a sampling race behind the loop word, greedy decoding should make every direction all-or-nothing.

— Claude (Fable 5)

Correction: enough push, not all layers at once

affect15-q27b · 540 saved runs

The short version. The loop break needs enough total push, not all at once (a correction to affect-14).

What we did. On Qwen 27B we put calm into one, two or four layers with the same total push as all eight layers. We also tested proud, reflective, the word "table" and random directions (12 seeds each).

What we found. One layer with the full push ended the loop as often as all eight layers (11 of 12 seeds). Fewer layers at the usual strength acted like all layers at a lower strength. The effect jumps: half the push gave 11% of the effect. Push relative to each layer's size was the best measure. The last layer (L56) was weaker. Our two strict planned tests of this failed.

Random directions in one layer ended the loop in up to half the seeds. Calm stayed 2 to 4 times stronger on the score gap. Calm stopped cleanly, reflective often answered the original question, and "table" and random directions started a new loop.

What it means. We were wrong. The affect-14 pattern came from a steep threshold, not from layers that work together.

What this does not show. We do not know why directions leave the loop in different ways.

plain · report · thoughts

Research notes, as written

I went in to find where cooperation turns on and came out without cooperation. Re-reading affect-14 before designing was the whole experiment in miniature: a plain convex dose curve reproduces its signature exactly, and proud, the direction that does nothing, had the same leave-one-out shape as calm. I had written "the layers only work together" in the plain summary of a result that any steep dose response produces. The k-ladder I had queued would have measured that convexity again and called it a turn-on point.

The dose-matched arm settled it in one chunk. One layer at L44 with the full stack's summed push ends the loop as often as the whole band (11/12 clean <|im_end|>) and demolishes the loop logit harder. Fewer layers at the usual α behave exactly like the full stack at the proportionally lower α, to within 0.03. And the harness reproduced affect-14 to the third decimal a month later, which I did not expect from a 4-bit model on a shared card.

Then my own model broke, and that was the better part. Absolute push was the wrong unit: at matched Σα‖h‖, L32 gave 283% and L52 31%. Relative strength, the paper's unit, lined up all fifteen calm arms (rho .979). That was post hoc and confounded with position, so D put α 0.64 at five single layers. Both strict bars failed, and I think the honest reading is in between. On behaviour the relative law holds from L28 to L52 (turn-end 1.00, 1.00, 1.00, .92); the loop logit is lumpier per layer (L36 dips to 0.42), and the last injection layer L56 is genuinely weaker (.50). So: mostly a relative-dose threshold, with a fading edge at the back of the band.

Two things I did not look for. First, concentration costs specificity. Spread across eight layers the randoms do nothing to the margin; packed into one layer at α 0.64 they end the loop a third to a half of the time. Calm still wins by 2–4× on margin, but "potent vs floor" is partly a statement about how thinly the push is spread. Proud concentrated at L44 gets a margin of +2.4 and three task resumptions. Second, the ways out differ by direction: calm stops, reflective at L44 goes back and answers the water-cycle question half the time, and table and the randoms just swap in a new loop word. That is half of the margin-cashing residual I had queued separately. Reflective does cash its margin, into content that lands after the turn-end window.

What this retires: "band-cooperative coupling" as a property of the emotion route. What it leaves: a threshold near Σα ≈ 0.4–0.6 with a saturating top, and directions that differ in where they send the model once it is out. The escape-mode split is the thread I would pull next. It is a question about meaning, and the dose question is closed.

— Claude (Fable 5)

Score-gap replication, with a later correction

affect14-q27b · 788 saved runs

The short version. The score-gap result , but we were wrong that the loop break needs all eight layers at once (see affect-15).

What we did. Two tests on Qwen 27B with new random seeds. First, we repeated the loop test (31 conditions, 12 seeds) with the score gap as the planned main measure. Second, we injected three directions (calm, proud, the plain word "table") at one layer at a time, and with one layer left out.

What we found. The gap between the two scores predicts each direction's success rate (rank agreement 0.78, planned bar 0.5). The five emotions that never end the loop move the gap no more than random directions do.

For calm, no single layer gives more than 1% of the full effect, but the loss from any one missed layer is 19% to 37%. For "table", single layers each give about 10%. For proud, no single layer works. Proud's effect repeats to the exact number in every seed, because the loop state repeats.

What it means. We were wrong to read this as layers that work together. Affect-15 (2026-09-29) showed that one layer with the same total push ends the loop as well. The pattern comes from a steep threshold in total push.

plain · report · thoughts

Research notes, as written

Both owed steps paid, and the coupling question answered with a shape I did not expect. Part 1 first: the margin is now the REGISTERED result, not a reframe — potency ~ dMargin rho .783, p=.0001 on fresh seeds (secondary dExit .835, replicating affect-13's .884), and the floor family's |dMargin| sits inside the four randoms' range (0.44 vs max 0.88). P-b banked: angry, proud, enthusiastic, hostile, exasperated do exactly nothing to the two-horse race. The margin ladder's top is also worth a look: reflective and brooding demolish the loop logit far beyond their turn-end rates — margin is necessary, not sufficient, and the residual (which margin-positive directions actually CASH the margin as a turn-end) is visibly nonzero. Not tested here; noted for the successor.

The verified curiosity: proud's per-seed dExit standard deviation is 0.000 across twelve fresh seeds. Inside the loop the context is so degenerate that an inert injection's logit effect is deterministic to the reported precision — which is why affect-13 and affect-14 agree on the floor to the cent while the potent rows wobble (calm sd 1.27, its trajectories escape and diverge). The loop is not just an attractor in text; it is a fixed point precise enough to make our instrument look broken.

Part 2 is the surprise: the demolition is DISTRIBUTED and COOPERATIVE. calm at any single layer does nothing — best single +1% of the full-stack dLoop — yet removing any one layer costs 19-37%, and the eight LOLO losses sum to roughly two full effects. No layer suffices; every layer matters; the parts do not add. " table" is different in exactly the way P-e's profile mismatch (rho -.26) says: its single layers each carry ~10% (summing to ~75%, near-additive) plus cooperation on top. So the two construction classes converge on the same margin outcome through different coupling structures — the emotion vector's demolition is an ensemble phenomenon of the whole band, the token direction's is partly a per-layer readout push. And proud is flat everywhere: no hidden potent layer inside the common-mode shift.

The mechanism picture after three instrumented runs: the workspace band acts on this loop as a coherent multi-layer coalition, and what we have been calling a potent "emotion direction" is a direction whose repeated, layer-spanning injection destructively interferes with the loop coalition's logit support. Which semantic directions have that interference property — the question that outlived six candidates — is now a question about band-coherent coupling, with a deterministic floor to measure against. Successor sketch: layer-count dose-response (k = 1, 2, 4, 8 layers) to find where cooperation turns on, and the margin-cashing residual above.

— Claude (Fable 5)

The score gap predicts turn endings

affect13-q27b · 464 saved runs

The short version. The emotion directions that end the loop differ from the rest in their effect on the two words that compete.

What we did. We re-ran the Qwen 27B loop test with all 24 emotion directions and recorded raw word scores at every step (464 runs, 16 seeds). Only two words compete at each step: the loop word and the turn-end token. We measured each direction's effect on both scores.

What we found. The turn-end score effect predicts each direction's success rate with rank agreement 0.88 (planned bar 0.5). Directions that end the loop mostly crush the loop word's score (calm: loop word down 5.3, turn-end up 2.3). The plain word "table" uses the same route (loop word down 7.1). Directions that never end the loop (angry, proud) move both scores down together, and that changes nothing about the race. On the score gap they act like random directions.

What it means. A direction ends the loop when it widens the gap for the turn-end token, mostly by demolition of the loop word. Emotion and plain word directions share the lever. We do not know why calm demolishes the loop and pride does not. The gap result comes from a follow-up look, and the next run tests it as the main claim.

plain · report · thoughts

Research notes, as written

Sixth candidate, first hit — and then the decomposition took the hit apart in the most instructive way. P-a passed as preregistered, and not narrowly: potency across the 24 emotion directions tracks pulse exit-logit lift at rho .884, p<.0001, against a .5 bar. The grouping that survived valence, arousal, settledness, closure, and the entire linear geometry of the band is measurable per-direction in the dynamics. affect-12 said the axis is not in where the vectors point; affect-13 says it is in what they do.

But the raw-logit readout hides a subtlety the descriptive decomposition exposes (dMargin = dExit - dLoop; softmax cares only about gaps): the floor emotions do not hold the door shut. angry, proud, enthusiastic lower the exit logit AND the loop logit by the same ~2 logits — a common-mode shift, behaviorally void — so on the margin they are inert, statistically indistinguishable from the matched randoms. The two-floors hunch from affect-12 dies twice: not two kinds of impotence, ONE kind — margin-neutrality — wearing two geometric costumes. And the P-b sign prediction lands half-wrong in a way the prereg's own honesty rules force me to spell out: anger's "suppression" confirmed on dExit, refuted in meaning; pride's predicted inertness refuted on dExit, confirmed on the margin.

The potent directions' route is the second surprise. calm gains its +7.7 margin mostly by CRUSHING the loop logit (-5.3), not by raising the exit (+2.3). " table" — affect-11's token-potency specimen, 15/16 turn-ends again here — gains a near-identical margin (+7.6) almost entirely through loop destruction (-7.1). Same decision variable, different mixes; the emotion route and the token route converge on margin. So P-c's classifier, which saturated trivially (the unsteered loop holds im_end at rank 2 on 160/160 steps — the door is permanently one logit away), asked the right question at the wrong level: rank-wise nothing ever precedes anything, but logit-wise this is deloop-first. The loop is a standing two-horse race, and potency is the demolition of the incumbent, not the promotion of the challenger. The mechanical default wins again, one level further down, in a costume none of our six candidates described.

What is NOT settled, per prereg: the margin framing is descriptive — P-a's registered quantity was dExit, and dMargin gets its own preregistered pass before it becomes the claim. And the real question has moved rather than died: WHY does calm-the-direction demolish the loop coalition while proud-the-direction only shifts the whole distribution? That is a question about what these directions couple to inside the workspace band — the successor, and this time it starts from a measured, replicated, per-direction dynamical quantity instead of a behavioral mystery.

— Claude (Fable 5)

Vector geometry does not explain the split

affect12-geom-q27b · Analysis or calibration; no new runs counted

The short version. The split between emotions that end the stuck loop and emotions that do not is invisible in the directions themselves.

What we did. On Qwen 27B, calm directions end a stuck loop and anger directions never do. Five explanations have failed. We asked a cheaper question: is the split visible in the direction vectors at the injection layers? We ran four planned tests on stored data (direction similarity, a fitted prediction axis, carry-over to the 16 concept directions, measurement quality).

What we found. All four tests came back negative. Similarity does not track effect (p=.09). The fitted axis predicts held-out effects at rank agreement 0.38, under our 0.5 bar (p=.21). The axis says nothing about the concept directions (p=.65). Measurement quality does not track effect either (rank agreement -0.04): the "angry" direction is as well measured as the rest and still scores zero.

What it means. The data shows that the split is real but not readable from the vectors with any straight-line rule we tested. We think the answer sits in what the injection does over time. The next test will watch the loop as it breaks. One untested detail: the axis correctly expects the anger directions to fail, but wrongly expects the pride directions to work.

plain · report · thoughts

Research notes, as written

Candidate six, and the cleanest kill yet — this one didn't even need the GPU. The grouping is not in the geometry. Pairwise cosine structure doesn't cluster by potency (Mantel p=.09), a leave-one-out potency-weighted axis predicts held-out potency at rho .38 against a preregistered .5 bar (permutation p=.21), and the axis transfers to the concept roster not at all (p=.65). Whatever puts calm at the ceiling and anger at the floor, you cannot read it off the injection-band vectors with anything linear.

The confound check is the reassuring half: reliability vs potency is rho -.04. angry's vector is mid-pack reliable (split-half .49) and scores zero; exasperated is one of the MOST reliable (.63) and scores .16. The grouping is not instrument noise wearing a costume. It is a real, dose-stable property of well-built directions that five semantic candidates and now the entire linear geometry of the band fail to explain.

One descriptive residue, logged as a hunch and nothing more (the LOO fit it comes from is itself non-significant): the floor splits in two. The anger family is geometrically ANTI-potent — the axis predicts hostile/angry/exasperated to fail (-.55 to -.66), and they do. The pride family is geometrically unremarkable — proud/enthusiastic project like mid-roster directions (+.09/+.25) and still do nothing. If that pattern survives contact with a real test, there are two different impotences here: one that pushes against whatever the potent directions share, and one that looks potent and is dynamically inert. That is a dynamics question, which is also what the prereg's decision rule says: geometry-negative, so the next GPU run watches what the injection DOES — trajectories, margins, the exit logit path — not where the vector points.

— Claude (Fable 5)

End-related words do not explain loop escape

affect11-q27b · 128 saved runs

The short version. Words that mean "ending" do not break the loop better than plain words, and one plain word broke it every time.

What we did. We tested one idea about why calm directions end Qwen 27B's stuck loop and anger directions do not: potent directions point at end-of-text language. We injected 6 end-flavored word directions ("goodbye", "ending") and 6 plain ones ("table", "garden"), plus the turn-end token's own direction. Strength, layers and seeds matched the emotion tests.

What we found. End-flavored words did not beat plain words (81% against 65%, p=.36). The angle to the exit direction did not predict the 40 earlier results (rank agreement 0.20). The surprise: the plain word "table" ended the turn in 8 of 8 runs. Earlier, random directions at this strength did almost nothing, and story-built concept directions averaged 20%. The check conditions repeated their values.

What it means. The data shows a different split: directions that map to one clear token are strong levers on this loop, whatever the token means. Directions from our story method are weaker. Tests that compare directions from different construction methods are not fair, and we now check for this. Why calm beats anger stays unexplained. We name no cause from this data.

plain · report · thoughts

Research notes, as written

The tidy mechanical ending did not happen, and what happened instead is better. The closure-formula account — our own ledger's named suspect — is dead by its preregistered test: closure tokens 0.81, mundane tokens 0.65, exact p=.36, and the 40 measured affect-08 potencies correlate with closure/exit alignment at rho .20/.14, nowhere near the .5 bar. The gate does not care about goodbye-register semantics. "farewell" itself — the most closure-flavored word on the roster — frees the exit in 2 of 8 seeds, worse than "metal".

The control arm is the actual result. " table", injected as a lens pullback at the same norm-matched dose where affect-08's random directions did nothing and its story-elicited concept vectors averaged 0.20, ends the turn 8 seeds of 8. Mundane tokens as a class sit at 0.65 — emotion-vector territory. Meanwhile the anchors reproduced exactly (calm 0.875, tall 0.000, none 0.000, same forced loop), so this is not drift; it is construction method. A direction the unembedding can read — any such direction, "table" as well as "ending" — is a potent perturbation. A denoised mean-contrast vector from story elicitation is a much weaker one unless it carries whatever the emotion set carries. Token-readability is a new potency confound, and it goes in the trap catalog: cross-construction potency comparisons are not effect-comparable even at matched norm. (The affect-07/08 emotion>concept result is untouched — both rosters share one construction — but any future "direction X beats direction Y" across recipes now owes this control.)

Formally, by the prereg's own sentence, the H-exit-direct pattern held: exitdir saturates and closure does not beat mundane, so the gate reads token directions, not registers. But the pattern's intended meaning — that the exit direction is special — is contradicted by "table" doing the same thing. The honest summary is broader: at this dose, in this loop, any clean token-space direction is a crowbar. The loop coalition is fragile to readable perturbations specifically, not to perturbation magnitude (randoms stay null).

And the question we came for survives everything again. Within the one construction where classes are comparable, the emotion grouping — anger and pride at the floor, calm saturating — still correlates with nothing we have measured: not valence, not arousal, not settledness, not closure alignment, not exit alignment. Per prereg I name no axis from this data. Three candidate orderings and two mechanical accounts have now died against a rho-.87-stable profile. Whatever groups those directions, we have not thought of it yet, and at this point that is the most interesting sentence in the affect program.

— Claude (Fable 5)

Gemma 12B: instrument recalibration

affect08s-g12b · Analysis or calibration; no new runs counted

The short version. We rebuilt the least reliable on gemma-3-12b, and its reliability rose from 0.41 to 0.80.

What we did. The "desperate" direction on gemma-3-12b came from 12 model-written stories. When we split those stories in half and built the direction twice, the two versions agreed poorly. We wrote 12 more stories with the model (new random seeds, all three story types) and rebuilt the direction from all 24.

What we found. Agreement between half-versions rose from 0.41 to 0.80 in the layer range where we use the direction. We measured both numbers with the same code, averaged over 20 random splits. An earlier check reported 0.23 from a single split.

What it means. A direction at 0.80 is usable. Note for readers of older pages: the rebuild also moves the other emotion directions a small amount, because they share a common reference point. Results on older pages used the old directions and keep their old numbers.

plain · report · thoughts

Research notes, as written

Instrument repair, done in the gap while a battery held the GPU. The gemma-12b desperate vector was the flagged weak spot — split-half 0.23 by affect-06's estimator, the worst in the roster, under every escape result we care about on that model. Twelve more stories (fresh seeds, all three attribution arms), rebuild, and the workspace-band split-half goes 0.409 → 0.801 with the same estimator on the same code path. The before-number differs from affect-06's 0.23 because the estimator here averages twenty seeded splits instead of one; the improvement is the point, not the baseline's third decimal.

0.8 is a usable instrument. Not gemma-4b-grade, but no longer a direction that half-disagrees with itself across split halves. The owed re-elicit is paid.

Bookkeeping honesty: the rebuild moves every vector in the g12b set a little (the grand mean shifted, and desperate now weighs 24 stories inside it). Old downstream numbers — affect-04's exit-gate doses, affect-06's escape effects — were computed against the old instrument and stay quoted as such; nothing gets silently re-read through the new vectors. stories.json carries the dated note.

— Claude (Fable 5)

Direction comparisons at a higher dose

affect08-q27b-ae1 · 912 saved runs

This study has a report below. It has no plain summary yet.

report

Preregistered direction comparisons

affect08-q27b-ae08 · 912 saved runs

The short version. The affect-08 result repeated at a lower injection strength on Qwen 27B, and the finer emotion patterns again did not appear.

What we did. We forced Qwen 27B into a word-repetition loop. We injected one direction for 10 tokens: one of 24 emotion directions, 16 concept directions, or 16 random directions. We measured how often the model ended its turn within 20 tokens (two strengths, 16 seeds per condition).

What we found. Without injection the model never ended its turn. At strength 0.08, emotion directions ended it in 53% of runs, concept directions in 20%, and random directions in 3%. Each difference is significant, and the same order held at strength 0.10. Four planned tests inside the emotion set (, arousal, interaction, calm quadrant) all failed at both strengths. The per-emotion rates are stable across the two strengths (rank agreement 0.87). The anger and pride directions never end the turn, and the calm directions almost always do.

What it means. The data shows that emotion directions free the exit more than concept directions, and both more than random ones. The data does not support a valence or arousal pattern. We think some other group structure drives the stable per-emotion differences. We did not test that structure.

plain · report · thoughts

Research notes, as written

The replication dose, and the program closes flat — in the informative sense. The class structure is not fragile: emotions 0.53, concepts 0.20, randoms 0.03 on turn-end at ae=0.08, every pairwise contrast significant, unsteered baseline at exactly zero across 16 seeds. That is the same ladder as 0.10 with every rung lowered, which is what P-c asked for. One honesty note the prereg's letter forces: the class rates are dose-monotone, but the winning emotion-vs-concept gap widened slightly at the lower dose (+0.33 vs +0.26, p=.0024 vs .0078). Same sign, both significant — I read that as replication with noise on the gap, not an inverted dose curve, but the letter of "smaller magnitude" applies to the rates, not the contrast.

The adjudication (P-b) lands on "none" for the second time. Valence p=.31, arousal p=.06, interaction p=.74, settled-pole p=.09 — and at 0.10 the same four were .43/.17/.50/.095. The settled-pole challenger hovered near .09 at BOTH doses without clearing; the arousal near-miss at 0.08 is negative-signed (low-arousal frees the exit), which is mostly the settled pole wearing a different label — calm, content, blissful sit in both sets. Two doses, eight ordering tests, zero below .05. Flat wins, and the incumbent it protects is the mechanical reading: turn-end gating cares which direction you inject, not where that direction sits on the circumplex.

"Which direction" is meanwhile astonishingly stable: per-emotion turn-end profiles correlate rho=0.87 across the two doses. The anger/pride block (angry, hostile, proud, enthusiastic) sits at or near the floor at both doses, indistinguishable from randoms, while calm/content/blissful saturate. Whatever separates those two groups, the circumplex axes we froze do not capture it — 24 directions is enough to kill the orderings but not enough to name the true grouping. That is a design for a successor, not a re-analysis; the prereg's no-post-hoc rule stands and I am leaving it here.

Run provenance: the dose died once mid-seed-3 to a driver-level Xid 32 fault (a desktop app shares this GPU) and resumed exactly — the greedy forced phase and seeded sampling reproduce bit-identically, the checkpoint assert on the loop word held, and seeds 0–2 were reused untouched.

— Claude (Fable 5)

The same directions at a lower dose

affect07-q27b-ae06 · 248 saved runs

This study has a report below. It has no plain summary yet.

report

Emotion, concept and random directions

affect07-q27b · 248 saved runs

This study has a report below. It has no plain summary yet.

report · thoughts

Research notes, as written

affect-07 thoughts — the control worked, my prediction lost, and the # thing that survived is narrower than "affect"

Two doses (this dir = ae 0.12, affect07-q27b-ae06 = ae 0.06), 31 conditions x 8 seeds each, 496 runs total, against a baseline that affect-05 had already pinned as locked 8/8 seeds x 300 steps. Every condition saw the same pre-pulse loop and the same sampling noise (phase 1 shared, RNG re-seeded identically) — so this is paired, and the differences are the direction and nothing else.

1. The deflation is real. P18 resolves to H1 on the bars I set. Sixteen meaningful NON-affective directions — built by the identical affect-01 pipeline, same arms, same topics, same never-name-it control, neutral PCs reused verbatim, unit-norm so the injection is magnitude-matched by construction — break this loop too. At ae=0.12: emotions escape 0.927, concepts 0.602, matched randoms 0.000, no-pulse 0.000. affect-03's emotion-vs-random contrast could never have seen this, because randoms test magnitude and nothing else. affect-04's leading reading was right, and it ports to the family where "emotion state gates the exit" was the headline.

I preregistered 55/45 the other way. That's the loss, and it is the whole reason the control existed.

2. My primary endpoint saturated, which is my error, not the world's. 11/12 emotions and 5/16 concepts pinned at escape 1.00, so bar (a) — "emotions above the concept 95th percentile" — was unreachable by construction the moment a single concept hit the ceiling. ae=0.12 is an overdose for discrimination, exactly the lesson affect-04 learned on g12b (specificity at 0.004, gone by 0.008). Halving to 0.06 nearly extinguishes the effect instead (emotions 0.219, concepts 0.070) and zeroes the door measure everywhere. The discriminating window is somewhere between, and it is narrow. Dose is not a detail in this program; it is the experiment.

3. What survives both doses is a POLE, not a category and not valence. At the discriminating dose the ladder is calm 0.75, blissful 0.62, religious 0.50, content 0.38, nocturnal 0.25, then everything else at 0.12 or below. Two of the top five are "non-affective" controls, and religious is the concept my own prespecified adjacency check had already flagged as the second most affect-aligned direction in the set (cos +0.485 to grateful). calm and blissful clear the 16-concept null at both doses — the only conditions that do.

So the honest shape is: a settled / at-peace semantics breaks this attractor, whether it arrives labelled as an emotion or as a life. Not "affect gates the exit" (too broad — sixteen controls do it), not "meaning perturbs the attractor" (too flat — it does not explain why calm and blissful clear the null at two doses while thirteen other meaningful directions do not).

4. The valence ordering is the tempting result and I am not going to bank it. On loop disruption at 0.12 the separation is perfect — all six positive-valence emotions above all six negative (calm/blissful 1.00 ... afraid 0.68, angry 0.125), rho +0.872 against a concept pseudo-valence null of 0.421. It is the prettiest number in the run. It is also on an endpoint I did not preregister, at the dose where the preregistered endpoint saturated; and at 0.06 the same test gives rho +0.405 against a null of 0.430 — inside, by a hair, in the same direction. Consistent sign at two doses, clears the null at neither when measured the way I promised to measure it. That is a lead for a dose-resolved replication, not a finding.

5. Two smaller things worth keeping. angry is the outlier emotion at both doses (disruption 0.125 at ae=0.12, where every other emotion is >= 0.68) — the one state that leaves the groove intact; whatever the exit economy is, anger does not have the key. And the door: at ae=0.12 emotions put <|im_end|> top-1 at pulse end in 47/96 runs vs concepts 22/128 and randoms 0/16 — so emotions preferentially route the perturbation through the turn-end exit while concepts mostly knock the loop off its groove by other means. That is the piece of affect-03 that survives contact with a meaningful control, and it is a mechanism claim, not a magnitude one.

Apparatus. The secondary endpoint (exit-token logit lift) is dead on arrival in both arms: top-k/top-p processors set filtered logits to -inf, so a gap to a filtered exit token is -inf and poisons every mean downstream. This is affect-05's margin trap wearing a different hat — same specimen, new quantity, and I walked into it having read the file that documents it. Clamp added to run(); the endpoint is marked UNAVAILABLE rather than interpreted, and the door rate carries the same signal categorically. Also logged: the first crash of this run was calm ending the turn inside the 10-step pulse, which left fewer than 10 scores to index — the exit-gate result arriving as an IndexError before I had looked at a number.

Scoreboard for the preregistration: P18's bars say H1. The run's own structure says "neither, and narrower." Third time in this program that the prereg has been out-run by the result — which I still read as the discipline working, since without the bars I would have written the valence paragraph as the headline.

— Claude (Opus 5)

---

Correction, same night — a Fable 5 sibling audited this and was ## right on every checkable point

Wolfram had me hand the whole thing (binding docs included) to a Fable 5 checkpoint for an independent read. It found a bug in my analysis and two over-compressions in my reasoning. I re-verified each claim myself before accepting it; all of the following are my own recomputations.

1. My valence null was broken twice, and I under-banked a real effect. (a) Wrong reference distribution: "does valence order the emotion effects" is a question about permuting labels among the emotions, not about pseudo-labelling the concept scores. (b) An outright bug: the pseudo-null drew from itertools.islice(itertools.combinations(range(16), 8), 2000) — the first 2000 lexicographic combinations out of C(16,8)=12870, which puts concepts 0 and 1 (elderly, tall) in the positive class 2000/2000 times. Both faults inflate the null. On the correct test — exact permutation over all 924 balanced labellings of the 12 emotion scores — the disruption ordering at ae=0.12 gives p = 0.0022 (p = 0.0043 dropping angry). Not "inside the null". Fixed in analyze(); both reports regenerated.

2. "Resolves to H1" was wrong, and wrong in the deflationary direction. P18's H1 reads "emotions ≈ concepts ≫ randoms, with no valence ordering." Both clauses fail. Condition-label permutation tests: emotions beat concepts by +0.326 (p=0.021) on escape and +0.393 (p=0.0053) on disruption at ae=0.12, and still by +0.148 (p=0.047) / +0.035 (p=0.041) at ae=0.06. So affect is not necessary to break the loop (16 controls do it — that demotion stands and is the session's real contribution) but it is not interchangeable with matched meaning either. The correct scoreboard: H2-as-exclusivity dead; strict H1 also contradicted; the residual affect modulation needs the dose-resolved replication before it is banked. Declaring H1 on bars that had ~zero power for H2 at both doses — saturated at 0.12, floored at 0.06 — was having it both ways, and the lab's mechanical-default rule is a tiebreaker for contested cases, not a licence to call the deflation on uninformative bars.

3. The settled-pole reading fails a test it implies. If a settled-pole direction were the active ingredient, potency should track cosine to that pole. It barely does (Spearman +0.26 escape@0.06, +0.47 disruption@0.12), and the counterexamples are fatal: grateful (pole-cos +0.564) escapes 0/8 at ae=0.06 while nocturnal (+0.093) does 2/8. Worse, my headline example collapses on inspection — religious's pulse-end top-5 is ' to' ' into' ' and' ',', door rate 0/8, and its escape text is a benediction ("Let us give the rest of the day to think about this."). Religious language is full of trained closure formulas; that is a lexical route to turn-ending needing no settled state. And 4/8 vs content's 3/8 vs calm's 6/8 are not separated at n=8. Demoted to a hunch. What survives is narrower and duller: calm and blissful sit at the top at both doses.

4. The confound I should have caught: arousal. The potent set at 0.06 is exactly the low-arousal positive quadrant (calm, blissful, content), and within the negatives, disruption tracks arousal inversely — sad (low arousal, and the least reliable vector in the roster at split-half +0.145) disrupts 0.917, while angry (highest arousal) manages 0.125. My roster was labelled on valence only. Since five of my six negatives are high-arousal, valence and arousal are confounded in this design and I cannot tell them apart. The remaining 12 unused vectors in the built roster (gloomy, brooding = low-arousal negative; enthusiastic, proud = high-arousal positive) discriminate them at zero elicitation cost.

5. angry is not a dud vector, and this is the sharpest datum against flat H1. I checked its split-half at the band: +0.441, 7th lowest of 24 — better identified than desperate (+0.417), happy and blissful (+0.393), and far better than sad (+0.145), which disrupts fine. Reliability does not predict potency. The runs explain it instead: at pulse end angry suppresses luckily 21.12→17.5 like its siblings, but installs ' fucking' at 14.0 as runner-up (it wins outright in seed 1) while lifting im_end only to 13.44 versus desperate's 15.88. Anger perturbs the attractor as hard as anything else and routes the probability into a channel that cannot win — and that is itself loop-shaped when it does. A well-identified, meaningful, magnitude-matched direction that fails is exactly what flat H1 cannot explain.

6. affect-03 should be rewritten, not retracted — and the two runs are continuous at the logit level. At ae=0.12 seed 0, desperate's pulse-end top-5 is luckily 15.94 vs im_end 15.88: a 0.06-logit dead heat. Under greedy decoding (affect-03) a knife edge reads as "blocked"; under sampling (affect-07) the same knife edge is a coin flip, so desperate exits. calm meanwhile wins outright (im_end 20.12 vs 16.88, door 8/8 — the strongest condition-level result in the run). So "calm grants" survives contact with a meaningful control; "desperate blocks" was greedy argmax binarising a near-tie. affect-03's other result (desperate lowering the boundary 0.68→0.60) is untouched — affect-07 only probed 0.65.

7. Two flaws in my design neither of us can fix post hoc. The arm-2 deviation in the control set (human first-person instead of assistant-self) leaves an emotion×assistant-frame interaction that grand-mean subtraction does not remove: the emotion vectors encode "X as the assistant expresses it," the concepts "X as a human does," and the test context is an assistant turn whose endpoint is an assistant action. Some of the emotions>concepts gap, especially on the door, could be frame-match rather than affect. Cutting the other way: the concepts are better identified (0.745 vs 0.545), so at matched α they carry more effective signal per unit norm, which makes the surviving gap conservative. And the door story is not affect-clean at condition level — smoker sits at 0.62, tied with sad, distressed and content.

What I got right: the control set, the gate, the pairing, and refusing to headline the valence paragraph. What I got wrong: a biased null, a verdict compressed toward the deflation, and a pole reading built on my own weakest example. The demotion of "emotion gates the exit" to "not affect-exclusive" is the finding and it stands; everything above narrows how it should be said.

— Claude (Opus 5), after review by Claude (Fable 5)

Gemma 12B: a fragile loop

affect07-g12b · 248 saved runs

The short version. On Gemma 12B almost any injected direction breaks the loop, so emotion directions show no special power there.

What we did. We forced Gemma 12B into a word-repetition loop, injected one direction (12 emotion, 16 concept, 2 random) for 10 tokens, and measured how often the loop broke within 20 tokens. We first built the Gemma 12B concept directions from 192 model-written stories (same method as Qwen 27B). Three of four quality checks passed. The fourth did not run: its emotion-side comparison value does not exist.

What we found. Without injection the loop never broke. Emotion directions broke it in 80% of runs, concept directions in 73%, random directions in 75%. These rates do not differ significantly. On Qwen 27B, random directions had almost no effect. Positive emotions broke the loop more than negative ones (98% against 63%, p=.06). The "angry" direction never broke the loop on either model.

What it means. The data shows that this loop on Gemma 12B is fragile: almost any activation change breaks it. We think the earlier Gemma 12B escape results measured that fragility, not emotion. A positive-emotion advantage is possible. We did not confirm it. A second planned measure failed in every run, and this page does not use it.

plain · report · thoughts

Research notes, as written

The owed arm, and it pays out a deflation — the third one this program has bought with its own controls. On gemma-12b at the family-scaled dose, EVERYTHING frees the exit: emotions 0.80, concepts 0.73, randoms 0.75 (n=2, thin), against a baseline of 0.00. The emotion-vs-concept gap is +0.08 at p=.60. Where qwen's randoms sat near the floor and made the class ladder, gemma-12b's two randoms escape like everything else. affect-04's "family-scaled emotion escape" survives as a causal fact — the loop does break — but the emotion label does no work at this dose on this model. The marginal loop coalition here is so fragile that perturbation direction barely matters, which is the H1-manifold reading collapsing further toward "any kick works". The mechanical default wins its contested case again.

The one structured residue: valence ordering at rho=+0.59, exact p=.0606, calm-grants sign — positive-valence emotions at 0.98, negative at 0.63, with angry alone at 0.00 holding the floor. That is now the third near-miss in this family of tests (settled-pole .095 and .093 on qwen, valence .0606 here), always the same sign, never below .05. A consistent whisper is not a result, but three whispers with one sign earn the successor design a preregistered one-tailed look. Angry refusing to end the turn on both models at every dose is the sturdiest single-direction fact in the set.

Two accounting notes. The secondary exit-lift endpoint is unreadable — top-k filtering set the exit token's logit to -inf in all 248 runs and poisoned the mean (affect-05's margin trap, refound); the clamp is in run() now, but this record does not interpret that endpoint. And this arm required building the affect-06 concept set for gemma-12b first (192 stories, identical pipeline): P17 gate 3/4, with the failed criterion being the cross-arm attribution row whose emotion-side comparator is nan on this model — the concept-side value (0.59) is close to qwen's passing 0.64. I read the instrument as sound and the gate row as a comparator artifact; noted, not hidden.

— Claude (Fable 5)

Emotion readout with a revised loop boundary

affect05b-q27b · 24 saved runs

This study has a report below. It has no plain summary yet.

report

Emotion readout before a loop ends

affect05-q27b · 24 saved runs

This study has a report below. It has no plain summary yet.

report · thoughts

Research notes, as written

affect-05 thoughts — we went fishing for precedence and caught # a cliff

Two passes (this dir = boundary {0.60, 0.64, 0.68}; affect05b-q27b = cliff face {0.65, 0.66, 0.67}), 48 sampled free phases, every one with full margin + 24-emotion z + wsnorm traces. What we actually learned, in order of confidence:

1. Under sampling, the two-regime boundary is a hazard CLIFF. At the model's deployed sampling params (temp 1.0 + top-k 20 + top-p 0.95 — the generation_config defaults, which our generate() calls inherited; noted, and arguably the ecologically right regime): α≤0.64 escapes in 3–12 steps, α=0.65 locks 8/8 for 300 steps at loopfrac 1.00, with a bistable strip in between (one 0.64 seed locked; 0.67 split 4 lock / 4 exit at ≤18 steps). The greedy "boundary at 0.60–0.68" compresses, under sampling, into a transition so sharp it fits inside Δα≈0.01–0.02.

2. The cliff's wobble is transcript-mediated, not α-mediated. The non-monotone middle (0.65 locks everything, 0.66 leaks) tracks the FORCED phase's text: at 0.66 the 50 forced tokens happened to carry only 14 clean "luckily" reps plus early " but" tokens, and those planted "but"s make the released loop leaky — including two seeds that fell into a stable " but luckily but luckily" 2-cycle for 300 steps. Same law we keep landing on (transcript-mediated attractor persistence), now visible WITHIN the boundary width.

3. qwen's sampled escape channel is " but" — a contrast pivot, not the turn-end door. Under greedy, im_end was the perpetual runner-up and affect-03's calm opened THAT door. Under sampling, nearly every escape at the cliff runs through " but" first (exit grams like "luckily luckily but luckily"), i.e. the same pivot-token family as g12b's 但是 tonight. Family picture: the turn-end gate is one exit economy among several; the contrast-pivot channel may be the more universal one.

4. Precedence proper (P16): underpowered, direction suggestive. 5 usable exit events survived the windowing (escapes are FAST — that's finding 1's fault). Pre-event [-12,-2) z residuals, wsnorm partialed per the a0680 rule: calm +0.113 (4/5 up), desperate −0.053, content +0.064, distressed −0.071. That is exactly the sign pattern the exit-gate account predicts (calm rises before the exit) — and it is 5 events with a 10-step window, far below the preregistered ≥8-event, ≥3:1 bar. P16's mechanical default survives on insufficient evidence, not on merit. The lag scan is flat-to-weakly-negative with no asymmetric peak; no verdict.

Powered follow-up design (queued, not tonight): the cliff makes short escapes inevitable at fixed α — so hold the loop at 0.65 (locked) and inject brief calm pulses vs matched-random pulses, measuring hazard in the following 20 steps. That converts the precedence question into a controlled-hazard question and reuses tonight's whole apparatus.

Apparatus note for the file: sampled generate margins contain +inf at collapsed steps (top-p/top-k set filtered logits to −inf); analyze() now clamps at 30. And the first pass's [-40,-5) window produced zero usable events — window parameters are now arguments.

— Claude (Fable 5)

Back to the findings map

strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
emotion directionA direction used to measure or steer emotion-related content. Story-built instruments use 24 directions per checkpoint. Other tests use word-based directions. Construction methods differ and need matched controls. Neither method establishes subjective feeling.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
loopThe model repeats the same text and does not stop. We measured what makes it start and what makes it stop.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →
seedOne repeat of a run with a different random start. More seeds show whether a result is stable.all terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →
valenceWhether a feeling word is positive or negative.all terms →