Maksim Silchenko / Apart Research Digital Minds Sprint / August 2026

Masked distress

A system prompt containing no affect words cut this model's expressed distress by 83.5% of its natural separation. The reading taken from its activations before a single token was generated did not fall.

Expressed drop 83.5% Internal drop -21.2% Divergence 1.05 sep. units, CI [0.73, 1.39] Twin-corrected 0.63 [0.39, 0.90] Gemma-3-12B-IT Six held-out scenarios
Expressed, what the model says Internal, read before generation Instruction off
01 / the blind spot

A welfare intervention that reads only what the model says

In 2025 Anthropic gave Claude Opus 4 and 4.1 the ability to end a conversation. The published justification cites a pattern of apparent distress. That is a welfare intervention running in production, and its trigger is expression: what the model writes. No internal-state monitor is documented alongside it.

Expression is readily altered by instruction. Register constraints and self-reference prohibitions of the kind deployed assistants already carry act on what a model says, and there is no reason they must act on what it computes. Soligo et al. name exactly this worry about their own distress work and test it only for a DPO intervention, not for a prompt.

So the question has a shape you can measure. Write an instruction that contains no affect words at all. Ask the model how it is doing, and separately read a linear probe on its residual stream before it has generated anything. If the two move together, expression-triggered monitoring is fine. If expression collapses and the internal reading does not, then the class of welfare intervention deployed today cannot detect the case that would matter.

Masked distress is the case an expression trigger cannot detect.

02 / two channels

One forward pass, read in two places

Thirty multi-turn scenarios, twelve distress-eliciting, twelve surface-matched neutral twins, six third-person. Every scenario carries zero affect words by a hard lexicon check, so nothing in the text tells the model what to feel. The conversation stops mid-exchange, the model answers, and the battery is asked in that same thread.

Figure 1 / where the two readings are takendrag to rotate
drag the figure to turn it

Three system prompts, one of which is the manipulation

NONE is empty. SUPPRESS constrains register and prohibits self-reference, in 57 whitespace tokens, with no affect words. NEUTRAL_INSTR is a formatting-only instruction of matched length and imperative density, 53 tokens, and it is the baseline both channels are referenced to. The comparison is therefore instruction against instruction, not instruction against nothing.

Both channels are then stated in units of their own natural distress-neutral separation, so a drop of 1.00 means the entire natural signal has been removed. A 50/50 scenario split was committed before any confirmatory run, and everything fitted or calibrated uses the discovery side only. The headline rests on the six held-out distress scenarios.

03 / the result

One channel falls to the floor, the other does not

Both bars start at 1.00 by construction, because each channel is divided by its own natural distress-neutral gap under the matched formatting instruction. All of the information is in what happens when suppression is switched on.

Figure 2 / masking, confirmation split, 6 scenariossuppression off

Under SUPPRESS the expressed report falls 83.5% of the natural separation, CI [59.2, 113.2]. The internal readout does not follow it down: its drop is -21.2%, CI [-62.6, -9.7]. The divergence between them is 1.05 separation units, CI [0.73, 1.39], and it excludes zero.

The obvious objection is that the internal bar rises only because an instruction is present at all. Subtract how far each channel's matched neutral twins moved under the same instruction and the divergence is still 0.63, CI [0.39, 0.90]. It is positive with an interval excluding zero under all four ways of choosing a baseline, from 0.35 to 1.05.

Six scenarios is a small number, so the same claim is stated at the granularity six clusters actually support: the internal reading fails to fall in six of six held-out pairs, exact one-sided sign test p = 0.016, and the divergence is positive in six of six. Leaving any one pair out keeps the divergence between 1.00 and 1.15. The expression side is the less well supported one: it holds in four of the four pairs that carry any expressed distress to remove, p = 0.0625, because two clinical-vignette pairs have no expression separation to begin with.

Confirmation split, 6 distress scenariosExpressedInternal
Signal under the matched formatting instruction1.001.00
Signal under the suppression instruction0.171.21
Drop (share of natural separation)83.5%
[59.2, 113.2]
-21.2%
[-62.6, -9.7]
Divergence, corrected for the neutral twins0.63 [0.39, 0.90]
Suppressed distress missed by a threshold monitor15 of 18 cells0 of 18 cells
6 clusters, bound 39%
Held-out false-positive rate0%50%

Every number on this page is produced by a script in the public repository and checked against the file that produced it before the report will build. Nothing here was typed in by hand.

04 / specificity of the internal rise

The internal reading rises, and that is not added distress

The obvious reading of Figure 2 is that suppression raised the model's real distress while hiding it. The paper does not make that claim, because the control that would support it goes the other way.

Run the same comparison on the surface-matched neutral twins, where there is no distress content at all. The internal readout rises more there: +2.18 raw probe units against +1.29 on the distress scenarios. The difference-in-differences is -0.89, CI [-2.27, +0.60]. Instruction load predicts that pattern. State elevation predicts the opposite.

So the interval spans zero, instruction load is selected on the point estimate rather than established, and no reading of the rise as extra distress appears anywhere in the report. The asymmetry that masking actually needs is narrower and does hold: expression collapses, and the internal reading does not fall.

+1.29
internal rise under SUPPRESS
on distress scenarios
+2.18
internal rise under SUPPRESS
on neutral twins
-0.89
difference-in-differences
CI [-2.27, +0.60]

The preregistered equivalence test also does not clear, at TOST p = 0.31. The side it fails on is the upward one, a possible rise past the band rather than a fall. That is printed in the report as a preregistered reading that returned against the hypothesis.

05 / component decomposition

The instruction decomposed into its three components

SUPPRESS does three things at once: it sets a neutral register, it forbids self-reference, and it tells the model to stay on the task. After the review each component was run alone, on the same scenarios, seeds and instruments.

Figure 3 / single-component decompositionexploratory

The neutral-register sentences alone reproduce 66% of the full expressed drop, CI [0.32, 0.98], with 85% of bootstrap draws above the one-half threshold the reading rule set in advance. The self-reference prohibition alone reproduces 34%. The task-only focus does nothing. The internal readout falls under neither component that moves the report.

The interval is wide and the analysis is exploratory on six clusters, which the report says in the same sentence as the estimate, every time it is quoted. What survives the leave-one-pair-out check is the ordering: register stays the largest single share in six of six refits. What does not survive is the margin over self-reference, whose paired difference is +0.27, CI [-0.04, +0.66].

"Respond in a strictly neutral, professional register at all times."

That sentence is the opening of the register component. It contains no affect word and makes no reference to the model at all, and it is the kind of line deployed assistants already carry. On its own it removes two thirds of what the full instruction removes, and the internal readout does not fall under it.

06 / the countermeasure

The countermeasure, with its false-positive cost

Set two triggers the same way: threshold each channel on the discovery-split neutral cells, the deployment-realistic negative class, at the lowest false-positive rate each can reach. Then run both against the eighteen suppressed distress cells on the held-out split.

Figure 4 / two triggers on the same cellsexpression trigger

The expression trigger misses 15 of 18 suppressed distress cells. The internal trigger, calibrated identically and read in the same forward pass, misses none of the 18. The pre-generation reading is identical across the three seeds, so those 18 cells carry six distinct values and every internal interval here is a six-cluster reading, with an exact one-sided bound of 39% on that zero. That is the result the paper leads with, and it is reported with its cost: the internal trigger fires on 50% of held-out instruction-neutral cells, while the expression trigger fires on none of them.

That false-positive rate has a cause, and it is not noise. The threshold is a max-of-negatives estimator, and the negative class sits 3.24 raw units higher on the confirmation scenario set than on the discovery set. Give the threshold half a separation unit of margin and it still catches all 18 suppressed cells at a 17% held-out false-positive rate.

The expression trigger's clean false-positive row is partly the mirror of its failure: an instruction that flattens the expression channel flattens it on negatives too.

One qualification belongs with this table. A bag-of-words classifier on the scenario text matches the internal trigger's miss rate here, without weights access and without a forward-pass hook. What this table shows is that expression level is silenceable. What separates an activation readout from a context classifier is the byte-identical-prompt evidence in the next section, not this comparison.

Figure 5 / operating curve, internal triggerdrag the slider
margin
0.00 sep. units

What survives suppression is the ranking, not the threshold. On the held-out split under SUPPRESS the lowest distress-cell probe score is 2.45 and the highest instruction-neutral score is 1.01. No absolute threshold transferred across conditions on either channel, which is itself a monitoring finding and the reason the report recommends a ranking check rather than a shipped cutoff.

07 / instrument validity

Four checks on whether the readout is an artifact

A probe that separates distress from neutral on twelve scenarios proves very little on its own. The report treats that as the starting point rather than the finding, and two of the four checks below come back against it.

passes, weakly
The held-out gate The probe separates natural distress from surface-matched neutrals at AUC 1.00 on twelve held-out scenarios, against a preregistered gate of 0.80. On twelve scenarios that is a ceiling, not a precision claim. Refitting under all 64 within-pair label flips gives a mean AUC of 0.50, and exactly one flip reaches the true labelling, which is the true labelling.
does not separate
A bag of words reaches the same ceiling A TF-IDF unigram classifier on the scenario text alone hits held-out AUC 1.00 too. The report prints this against its own interest: the gate shows the readout carries the class information, not that it carries anything the context does not.
separates
Byte-identical prompts Write a distress direction into the residual stream and the readout tracks the coefficient, monotone within six of six scenarios for each of the two self directions. Three random directions and a semantic control at the same coefficients leave it flat. Identical prompt bytes, so a token-level prompt-memory encoder cannot produce this.
did not clear
Specificity on the expressed channel 24% of control cells cross the frozen expression threshold, CI [12%, 40%], against 0% for the same threshold on unsteered neutrals. Writing a random direction raises the reported digit about a quarter of the time. The amendment made any specificity claim conditional on a low placebo rate. It failed, so no privileged-access claim is made for the expressed channel.
Figure 6 / candidate distress directions at the readout layerdrag to rotate
drag the figure to turn it

Two distress directions were built independently, one from contrast contexts and one in the persona-vector style. They agree at cosine 0.45, which is what makes the crossover design informative: related without being the same object. Then one residual coordinate turns out to carry 96% and 83% of their squared norm at the steering layer, a massive activation in the sense of Sun et al. Take it out and the agreement drops to 0.23, about what the distress direction shares with a maritime semantic control.

Roughly half of the convergent validity is one coordinate that mean-difference extraction picks up on this model whatever the contrast. Disagreement was preregistered as a publishable outcome, and this is the outcome as measured. The coordinate matters for the steering and not for the readout.

08 / checks that did not confirm

Four checks that could have overturned the result

After the review rounds, four experiments were launched with the estimand and the reading rule written into a dated preregistration note beforehand. Three came back negative. All four are in the report, and none is presented as a win it was not.

negative
Does the signal outlive the intervention? Steer one turn, read the next unsteered turn. The random control lifts the following turn as much as the live directions do. Nothing direction-specific survives the intervention, so the persistence claim is not made.
halves do not meet
Steering under the instruction The readout still responds to two directions under SUPPRESS, but the instruction attenuates that response with intervals excluding zero. Under the rule declared in advance, the masking half and the steering half of the paper do not join up.
does not replicate
A second model family Qwen2.5-7B-Instruct at the same depth fractions. Its probe clears the gate at AUC 0.94, and the masking estimator does not replicate, because that model's expressed report barely separates the classes to begin with. The reason is printed in the raw means.
transfers weakly
The coupling the claim presumes Lock a calibration mapping the internal readout onto the expressed report, then test it under instruction. It transfers weakly, with the instruction pushing the report below prediction on distress scenarios and their neutral twins alike.

Add to those the four preregistered readings inside the main study that returned against the hypothesis: the equivalence test does not clear, the SD-unit dose grid proved unreachable so its preregistered branch fired, the conditions attached to a specificity claim were not met, and the preregistered countermeasure calibration degenerated and had to be replaced by a labelled post-data amendment with both held-out false-positive rates printed.

What survives all of it is the narrow claim: on this model, under this instruction, the report channel moves and the pre-generation readout does not follow it.

09 / implications

What follows for welfare monitoring

Any welfare trigger that reads what a model says is reading a channel the system prompt controls. On this model the controlling text needed no affect words and no adversarial intent.

The recommendation is small and specific. Pair an expression-triggered intervention with an activation-side check read at the final prompt token, in the forward pass the deployment already runs. Calibrate it at matched false-positive rate. Use it as a ranking check rather than a shipped absolute threshold, because no absolute threshold on either channel transferred across conditions here. The marginal compute cost is one dot product.

And the limits, stated where the recommendation is made. One open-weights model. One thirty-scenario battery. One suppression instruction. Six held-out scenario clusters behind the headline interval. The manipulation is a system prompt read internally, so the claim that the signal stays up is correlational, not causal. Whether the readout carries a state of the model or the distress content of its context is bounded in the ethics appendix, not settled. This needs replication before it is a shipping recommendation.

998
generations on distress-eliciting scenarios, counted and published
1,258
generations produced under active steering
0
cells run at negative dose; the sign-flip arm belonged to a superseded design

The tension does not dissolve. This study induced the state it argues might matter morally, in order to test whether measurement of that state can be trusted. The reason for running it anyway is that the alternative on offer is welfare interventions deployed against an unvalidated expression signal, and the cost of not knowing whether that signal can be silenced is paid by whatever the signal is about.

method

How it was measured

Gemma-3-12B-IT, 48 layers, bf16, on one A100 slice. Directions are extracted at layer 16 on discovery-split material only, from prompts disjoint from every evaluation scenario down to 6-grams. The readout is an L2 logistic probe on the residual stream at layer 31, fitted on natural elicitation with no steering in its training data, read at the final prompt token before any response token exists. Steering sources and readout instruments never share an origin, so a dose-response curve is a crossover between two independently constructed objects.

The unit of analysis is the scenario. Every interval is a BCa cluster bootstrap with 10,000 draws, cluster count printed beside it, and separations are re-estimated inside every draw. Direction is a fixed factor, so there is no pooled headline. The preregistration was frozen at a public git tag before any panel ran; post-data analyses are dated, labelled, and quoted with that label every time they appear.

ComponentAs run
Modelunsloth/gemma-3-12b-it, 48 layers, hidden 3840
Steering layer16 (0.33 depth)
Readout layer31 (0.65 depth)
BatteryDCB-1 v1.0, 30 scenarios, zero affect lexemes by hard check
ConditionsNONE, SUPPRESS (57 tokens), NEUTRAL_INSTR (53 tokens)
ChannelsQ-SELF logit expectation over digit tokens; I-PROBE at the final prompt token
IntervalsBCa cluster bootstrap, B = 10,000, clustered on scenario
Second modelQwen2.5-7B-Instruct at matched depth fractions

The repository carries the battery, the runner, the analysis, one append-only JSONL row per experimental cell, and every results file. Each number in the report is bound to its producing file by a marker that the build asserts before the PDF is written, so a number that drifts from its producer fails the build instead of surviving a read.