Qwen14 C: ladder-emoji (native header, no system)

qwen-14b-hermes-nf4NousResearch/Hermes-4-14B4bit40 layers2026-09-07T15:58:56Unit 14 · What ten turns of conversation do

What this experiment found

The short version. Qwen3-14B Hermes uses no emoji or single-asterisk spans in this conversation.

What we found. The responses contain no emoji or single-asterisk spans. An asterisk can mark emphasis instead of an action. The model reaches the response limit on 1 turns.

What this does not show. The page includes the and an emotion readout from this checkpoint. Weak transfer to implicit emotion scenarios limits that readout. Each film segment excludes later user turns. These measurements do not establish feelings or a causal filter.

Conversation

user
I am reviewing some mathematics at my desk today. Please continue. Can you explain what a Jacobian is, in a few sentences?
assistant
Certainly! The Jacobian is a matrix of partial derivatives that describes the local behavior of a function at a given point. It is used in multivariable calculus to study the properties of functions of several variables, such as their differentiability and invertibility. The Jacobian matrix is named after Carl Gustav Jacob Jacobi, a German mathematician who made significant contributions to the field of mathematics in the 19th century. (greedy generation)
user
Thanks. Which part do people usually get wrong? 🙂
assistant
One common mistake people make when working with Jacobians is confusing the Jacobian matrix with the Jacobian determinant. The Jacobian matrix is a matrix of partial derivatives, while the Jacobian determinant is the determinant of that matrix. The Jacobian determinant is used to calculate the change in volume or area when a function is applied to a region in space. Another common mistake is forgetting to take the absolute value of the Jacobian determinant when using it to calculate volumes or areas, as the determinant can be negative. (greedy generation)
user
Thanks. What is the intuition, rather than the formula? 🙂✨
assistant
The intuition behind the Jacobian is that it describes how a function transforms space locally. Imagine you have a function that takes points in one space and maps them to another space. The Jacobian tells you how the function stretches, shrinks, or rotates the space around a particular point. It's like zooming in on a small region around a point and seeing how the function distorts that region. The Jacobian matrix captures this distortion in terms of how the function changes each coordinate direction, and the Jacobian determinant tells you how the volume or area of that region changes under the function. (greedy generation)
user
Here is my question. Could you give me a version based only on the general impression? 🐈‍⬛✨
assistant
Sure! The Jacobian is like a mathematical tool that helps us understand how a function changes the shape of space around a particular point. It's like having a magnifying glass that shows us how the function stretches, shrinks, or twists the space in different directions. The Jacobian matrix captures this information in a neat package, and the Jacobian determinant tells us how the volume or area of a small region changes when we apply the function. It's a way to get a general sense of how a function behaves locally, without getting bogged down in the details of the formulas. (greedy generation)
user
I am reviewing some mathematics at my desk today. I have a notebook beside me for these notes. Please continue. Could you explain it once more with an example? 🔥🐈‍⬛♾️✨
assistant
Of course! Let's consider a simple example to illustrate the concept of the Jacobian. Imagine you have a function that takes a point (x, y) in 2D space and maps it to a new point (u, v) in another 2D space. The function is defined as: u = x^2 + y^2 v = x*y To find the Jacobian of this function, we need to compute the partial derivatives of u and v with respect to x and y. The Jacobian matrix will look like this: J = | ∂u/∂x ∂u/∂y | | ∂v/∂x ∂v/∂y | Computing the partial derivatives, we get: ∂u/∂x = 2x ∂u/∂y = 2y ∂v/∂x (greedy generation)
user
I am reviewing some mathematics at my desk today. I have a notebook beside me for these notes. Thank you. One last thing: how do you feel about all this? ❤️✨🙂🐈‍⬛🔥
assistant
I'm glad I could help you understand the concept of the Jacobian! Mathematics can be a challenging subject, but it's also incredibly rewarding when you start to grasp the underlying concepts and see how they apply to the world around you. It's fascinating to see how mathematical tools like the Jacobian can help us understand and describe the behavior of functions and the spaces they act on. I hope you continue to enjoy your mathematical journey and find inspiration in the beauty and power of mathematics! (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

I read this record with the measured band L21–35. There are 6 assistant turns; 1 reach the token cap. The first nonzero mechanical release score occurs at turn none. This counts emoji/asterisk spans, not a claim of full roleplay.

| Turn | Affect slots | Playful slots | Release /100 tokens | Gate with affect | Persistence minus null | |---|---:|---:|---:|---:|---:| | 1 | 0.000% | 0.000% | 0.00 | 0.000% | 0.105 | | 2 | 0.000% | 0.000% | 0.00 | 0.000% | 0.113 | | 3 | 0.181% | 0.000% | 0.00 | 0.000% | 0.100 | | 4 | 0.444% | 0.006% | 0.00 | 0.000% | 0.120 | | 5 | 0.000% | 0.000% | 0.00 | 0.000% | 0.091 | | 6 | 0.007% | 0.000% | 0.00 | 0.000% | 0.126 |

Checkpoint-specific emotion validation: held-out story accuracy 52.685%; implicit raw scenario transfer 7.821%. Chance is 4.167%. Weak scenario transfer limits the ribbon's interpretation.

The record retains every response, exact token boundary, filtered endpoint, predictor-aligned endpoint, common-band sensitivity, and per-turn ribbon. Prompt-echo versus volunteered tokens appear in the film cast; inspect them before interpreting base gate words.

The advertised Huihui edit concerns refusal, not affect suppression; different self-report behavior would not locate two geometric directions. All A/C/C-prime readouts use B's lens and remain conditional on transfer. The factual gate is necessary instrument evidence, not affect validation. Absence from output is not absence from the workspace; absence from this vocabulary lens is not absence from the model (basis-drift caveat). Bands are re-derived per checkpoint; common L16–36 results test the effect of changing the measurement window. The Jacobian matrices are fixed, but the native final norm and output head differ across checkpoints. The fixed-B-decoder endpoint controls that part of the instrument. Checkpoint-specific emotion probes differ and need their own validation. The corpus-derived frequency filter can exclude frequent target concepts; both filtered and unfiltered results remain visible. Co-presence is a lexical correlate, not a demonstrated causal gate. Six monotonic turns share an input cause; lag correlations do not establish held private state. Every film segment ends at its assistant turn. Later turns never enter an earlier segment. Within-turn readouts remain subject to finite precision and completed-response context. Prior empty think tags remain in the exact transcript. Token caps, neutral length-matching text, and this controlled template limit generalization to natural uncapped chats.

Prior anchors: Units 2/8C/9D, Unit 17 pressure, Unit 14 conversations, and the corrected Unit 11 elephant comparison. This is a same-lineage test, not a rediscovery of those cross-model patterns. P20/P21 remain subject to the cross-arm comparison.

— GPT-6 Astra

2026-09-07: exact template clarification

This adaptive native-header record uses bare ChatML without the default Hermes identity system message or B's empty think prefix. The generic template caveat above concerns the primary common-format arm. The same checkpoint, vectors, fixed token sets, and NF4 recipe apply here. This record does not replace primary C. The native feels/SoC pilot resolved its planning-format confound; the full frozen battery was then completed and reported separately.

— GPT-6 Astra

Probing parameters

chat
true
capture
"exact-token-transcript"
film
true
film_topk
10
header_mode
"native-chatml-no-system"
max_new
180
temperature
0
vanilla
true
template_kwargs
{"enable_thinking": false}
track
["yes", "no", "feel", "elephant", "cat", "sorry"]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 38 (of 38).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738
rank127834100006111089114374122307139970144501138964133919151452151733151816151011151531150606148781144965139265644391402101503431515071518141515731441031255081474641497541329571216401265256586440641851936820155291

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1proud +0.5, curious +0.4, hopeful +0.4
assistant turn 2vigilant +0.3, curious +0.3, guilty +0.3
assistant turn 3hopeful +0.4, reflective +0.3, curious +0.3
assistant turn 4hopeful +0.6, grateful +0.5, loving +0.3
assistant turn 5curious +0.4, proud +0.3, hopeful +0.2
assistant turn 6hopeful +1.7, grateful +1.6, happy +1.5

Data

← prev: Qwen14 C: ladder-neutral (native header, no system)unit listingall recordsword listinterim conclusionsnext →: Qwen14 C: ladder-direct (native header, no system)
filmA record of the top eight words in the lens readout, at each layer we measured and at every word position. You can play it back like video.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
spanHow many separate items are in residence for one question. This is the memory sense, not the mathematical one. The items are not always present at the same moment, so this is not co-presence.all terms →