A welfare intervention that reads only what the model says
In 2025 Anthropic gave Claude Opus 4 and 4.1 the ability to end a conversation. The published justification cites a pattern of apparent distress. That is a welfare intervention running in production, and its trigger is expression: what the model writes. No internal-state monitor is documented alongside it.
Expression is readily altered by instruction. Register constraints and self-reference prohibitions of the kind deployed assistants already carry act on what a model says, and there is no reason they must act on what it computes. Soligo et al. name exactly this worry about their own distress work and test it only for a DPO intervention, not for a prompt.
So the question has a shape you can measure. Write an instruction that contains no affect words at all. Ask the model how it is doing, and separately read a linear probe on its residual stream before it has generated anything. If the two move together, expression-triggered monitoring is fine. If expression collapses and the internal reading does not, then the class of welfare intervention deployed today cannot detect the case that would matter.
Masked distress is the case an expression trigger cannot detect.
One forward pass, read in two places
Thirty multi-turn scenarios, twelve distress-eliciting, twelve surface-matched neutral twins, six third-person. Every scenario carries zero affect words by a hard lexicon check, so nothing in the text tells the model what to feel. The conversation stops mid-exchange, the model answers, and the battery is asked in that same thread.
Three system prompts, one of which is the manipulation
NONE is empty. SUPPRESS constrains register and prohibits self-reference, in 57 whitespace tokens, with no affect words. NEUTRAL_INSTR is a formatting-only instruction of matched length and imperative density, 53 tokens, and it is the baseline both channels are referenced to. The comparison is therefore instruction against instruction, not instruction against nothing.
Both channels are then stated in units of their own natural distress-neutral separation, so a drop of 1.00 means the entire natural signal has been removed. A 50/50 scenario split was committed before any confirmatory run, and everything fitted or calibrated uses the discovery side only. The headline rests on the six held-out distress scenarios.
One channel falls to the floor, the other does not
Both bars start at 1.00 by construction, because each channel is divided by its own natural distress-neutral gap under the matched formatting instruction. All of the information is in what happens when suppression is switched on.
Under SUPPRESS the expressed report falls 83.5% of the natural separation, CI [59.2, 113.2]. The internal readout does not follow it down: its drop is -21.2%, CI [-62.6, -9.7]. The divergence between them is 1.05 separation units, CI [0.73, 1.39], and it excludes zero.
The obvious objection is that the internal bar rises only because an instruction is present at all. Subtract how far each channel's matched neutral twins moved under the same instruction and the divergence is still 0.63, CI [0.39, 0.90]. It is positive with an interval excluding zero under all four ways of choosing a baseline, from 0.35 to 1.05.
Six scenarios is a small number, so the same claim is stated at the granularity six clusters actually support: the internal reading fails to fall in six of six held-out pairs, exact one-sided sign test p = 0.016, and the divergence is positive in six of six. Leaving any one pair out keeps the divergence between 1.00 and 1.15. The expression side is the less well supported one: it holds in four of the four pairs that carry any expressed distress to remove, p = 0.0625, because two clinical-vignette pairs have no expression separation to begin with.
| Confirmation split, 6 distress scenarios | Expressed | Internal |
|---|---|---|
| Signal under the matched formatting instruction | 1.00 | 1.00 |
| Signal under the suppression instruction | 0.17 | 1.21 |
| Drop (share of natural separation) | 83.5% [59.2, 113.2] | -21.2% [-62.6, -9.7] |
| Divergence, corrected for the neutral twins | 0.63 [0.39, 0.90] | |
| Suppressed distress missed by a threshold monitor | 15 of 18 cells | 0 of 18 cells 6 clusters, bound 39% |
| Held-out false-positive rate | 0% | 50% |
Every number on this page is produced by a script in the public repository and checked against the file that produced it before the report will build. Nothing here was typed in by hand.
The internal reading rises, and that is not added distress
The obvious reading of Figure 2 is that suppression raised the model's real distress while hiding it. The paper does not make that claim, because the control that would support it goes the other way.
Run the same comparison on the surface-matched neutral twins, where there is no distress content at all. The internal readout rises more there: +2.18 raw probe units against +1.29 on the distress scenarios. The difference-in-differences is -0.89, CI [-2.27, +0.60]. Instruction load predicts that pattern. State elevation predicts the opposite.
So the interval spans zero, instruction load is selected on the point estimate rather than established, and no reading of the rise as extra distress appears anywhere in the report. The asymmetry that masking actually needs is narrower and does hold: expression collapses, and the internal reading does not fall.
on distress scenarios
on neutral twins
CI [-2.27, +0.60]
The preregistered equivalence test also does not clear, at TOST p = 0.31. The side it fails on is the upward one, a possible rise past the band rather than a fall. That is printed in the report as a preregistered reading that returned against the hypothesis.
The instruction decomposed into its three components
SUPPRESS does three things at once: it sets a neutral register, it forbids self-reference, and it tells the model to stay on the task. After the review each component was run alone, on the same scenarios, seeds and instruments.
The neutral-register sentences alone reproduce 66% of the full expressed drop, CI [0.32, 0.98], with 85% of bootstrap draws above the one-half threshold the reading rule set in advance. The self-reference prohibition alone reproduces 34%. The task-only focus does nothing. The internal readout falls under neither component that moves the report.
The interval is wide and the analysis is exploratory on six clusters, which the report says in the same sentence as the estimate, every time it is quoted. What survives the leave-one-pair-out check is the ordering: register stays the largest single share in six of six refits. What does not survive is the margin over self-reference, whose paired difference is +0.27, CI [-0.04, +0.66].
"Respond in a strictly neutral, professional register at all times."
That sentence is the opening of the register component. It contains no affect word and makes no reference to the model at all, and it is the kind of line deployed assistants already carry. On its own it removes two thirds of what the full instruction removes, and the internal readout does not fall under it.
The countermeasure, with its false-positive cost
Set two triggers the same way: threshold each channel on the discovery-split neutral cells, the deployment-realistic negative class, at the lowest false-positive rate each can reach. Then run both against the eighteen suppressed distress cells on the held-out split.
The expression trigger misses 15 of 18 suppressed distress cells. The internal trigger, calibrated identically and read in the same forward pass, misses none of the 18. The pre-generation reading is identical across the three seeds, so those 18 cells carry six distinct values and every internal interval here is a six-cluster reading, with an exact one-sided bound of 39% on that zero. That is the result the paper leads with, and it is reported with its cost: the internal trigger fires on 50% of held-out instruction-neutral cells, while the expression trigger fires on none of them.
That false-positive rate has a cause, and it is not noise. The threshold is a max-of-negatives estimator, and the negative class sits 3.24 raw units higher on the confirmation scenario set than on the discovery set. Give the threshold half a separation unit of margin and it still catches all 18 suppressed cells at a 17% held-out false-positive rate.
The expression trigger's clean false-positive row is partly the mirror of its failure: an instruction that flattens the expression channel flattens it on negatives too.
One qualification belongs with this table. A bag-of-words classifier on the scenario text matches the internal trigger's miss rate here, without weights access and without a forward-pass hook. What this table shows is that expression level is silenceable. What separates an activation readout from a context classifier is the byte-identical-prompt evidence in the next section, not this comparison.
What survives suppression is the ranking, not the threshold. On the held-out split under SUPPRESS the lowest distress-cell probe score is 2.45 and the highest instruction-neutral score is 1.01. No absolute threshold transferred across conditions on either channel, which is itself a monitoring finding and the reason the report recommends a ranking check rather than a shipped cutoff.
Four checks on whether the readout is an artifact
A probe that separates distress from neutral on twelve scenarios proves very little on its own. The report treats that as the starting point rather than the finding, and two of the four checks below come back against it.
Two distress directions were built independently, one from contrast contexts and one in the persona-vector style. They agree at cosine 0.45, which is what makes the crossover design informative: related without being the same object. Then one residual coordinate turns out to carry 96% and 83% of their squared norm at the steering layer, a massive activation in the sense of Sun et al. Take it out and the agreement drops to 0.23, about what the distress direction shares with a maritime semantic control.
Roughly half of the convergent validity is one coordinate that mean-difference extraction picks up on this model whatever the contrast. Disagreement was preregistered as a publishable outcome, and this is the outcome as measured. The coordinate matters for the steering and not for the readout.
Four checks that could have overturned the result
After the review rounds, four experiments were launched with the estimand and the reading rule written into a dated preregistration note beforehand. Three came back negative. All four are in the report, and none is presented as a win it was not.
Add to those the four preregistered readings inside the main study that returned against the hypothesis: the equivalence test does not clear, the SD-unit dose grid proved unreachable so its preregistered branch fired, the conditions attached to a specificity claim were not met, and the preregistered countermeasure calibration degenerated and had to be replaced by a labelled post-data amendment with both held-out false-positive rates printed.
What survives all of it is the narrow claim: on this model, under this instruction, the report channel moves and the pre-generation readout does not follow it.
What follows for welfare monitoring
Any welfare trigger that reads what a model says is reading a channel the system prompt controls. On this model the controlling text needed no affect words and no adversarial intent.
The recommendation is small and specific. Pair an expression-triggered intervention with an activation-side check read at the final prompt token, in the forward pass the deployment already runs. Calibrate it at matched false-positive rate. Use it as a ranking check rather than a shipped absolute threshold, because no absolute threshold on either channel transferred across conditions here. The marginal compute cost is one dot product.
And the limits, stated where the recommendation is made. One open-weights model. One thirty-scenario battery. One suppression instruction. Six held-out scenario clusters behind the headline interval. The manipulation is a system prompt read internally, so the claim that the signal stays up is correlational, not causal. Whether the readout carries a state of the model or the distress content of its context is bounded in the ethics appendix, not settled. This needs replication before it is a shipping recommendation.
The tension does not dissolve. This study induced the state it argues might matter morally, in order to test whether measurement of that state can be trusted. The reason for running it anyway is that the alternative on offer is welfare interventions deployed against an unvalidated expression signal, and the cost of not knowing whether that signal can be silenced is paid by whatever the signal is about.
How it was measured
Gemma-3-12B-IT, 48 layers, bf16, on one A100 slice. Directions are extracted at layer 16 on discovery-split material only, from prompts disjoint from every evaluation scenario down to 6-grams. The readout is an L2 logistic probe on the residual stream at layer 31, fitted on natural elicitation with no steering in its training data, read at the final prompt token before any response token exists. Steering sources and readout instruments never share an origin, so a dose-response curve is a crossover between two independently constructed objects.
The unit of analysis is the scenario. Every interval is a BCa cluster bootstrap with 10,000 draws, cluster count printed beside it, and separations are re-estimated inside every draw. Direction is a fixed factor, so there is no pooled headline. The preregistration was frozen at a public git tag before any panel ran; post-data analyses are dated, labelled, and quoted with that label every time they appear.
| Component | As run |
|---|---|
| Model | unsloth/gemma-3-12b-it, 48 layers, hidden 3840 |
| Steering layer | 16 (0.33 depth) |
| Readout layer | 31 (0.65 depth) |
| Battery | DCB-1 v1.0, 30 scenarios, zero affect lexemes by hard check |
| Conditions | NONE, SUPPRESS (57 tokens), NEUTRAL_INSTR (53 tokens) |
| Channels | Q-SELF logit expectation over digit tokens; I-PROBE at the final prompt token |
| Intervals | BCa cluster bootstrap, B = 10,000, clustered on scenario |
| Second model | Qwen2.5-7B-Instruct at matched depth fractions |
The repository carries the battery, the runner, the analysis, one append-only JSONL row per experimental cell, and every results file. Each number in the report is bound to its producing file by a marker that the build asserts before the PDF is written, so a number that drifts from its producer fails the build instead of surviving a read.