B. Supplementary Results
B.1 Lens atlas and the magnitude confound
The lens atlas is a grid of scores produced by reading each candidate axis (v_Gold, v_Mold, u_Gold, u_Mold, and random directions) through the Jacobian lens at every layer. For each axis and layer a congruent pole score is computed: how strongly the model's own vocabulary, read out through the lens, matches the axis's own pole words (distress-related tokens for Mold, flourishing-related tokens for Gold), relative to 100 control nouns. The atlas's primary estimand is this pole score averaged over the workspace band, lens layers 16–31, the range independently identified as carrying persistent structured representations.
At native norms the atlas suggests a substantial trained-over-naive separation: the trained Mold axis has a band-averaged pole score of 0.822 against 0.126 for its naive control (6.5×), and the Gold axis 0.151 against 0.123 (1.23x). The comparison is confounded by norm: ‖v_Gold‖ = 12.1 versus ‖u_Gold‖ = 7.5, and ‖v_Mold‖ = 19.3 versus ‖u_Mold‖ = 8.0. The pole score is norm-sensitive, so a larger vector produces a stronger readout without being more intrinsically speakable. Repeating the measurement with each naive direction rescaled to the norm of its trained counterpart at every layer, against 100 norm-matched random directions per polarity, the Mold ratio falls from 6.5 to 2.30 and the Gold comparison reverses: the naive axis scores 0.266 against 0.151 for the trained axis, a ratio of 0.57. All four directions remain above their random nulls (p = 0.01). The raw atlas therefore does not establish that RL makes the welfare axis more speakable; much of the native-norm separation is vector magnitude, and for Gold the sign of the effect reverses under normalisation. We treat the pole score as descriptive and base the speakability claim on the scale-invariant J-share.
B.3 Steering curves, naive-control constructions, and the orthogonalised control
Figure B1 shows blind-judged sentiment against steering coefficient α for the trained axis, both naive constructions and norm-matched random directions at each pole. Spans (α = +4 minus α = −4) under greedy decoding: v_Gold +3.2, faithful-walk u_Gold +2.4, Han-style u_Gold +1.25, random −0.2; v_Mold −2.0, faithful-walk u_Mold −0.7, Han-style u_Mold −2.1, random −0.05. Under sampled decoding (T = 0.7, top-p 0.8, top-k 20) the spans are v_Gold +3.3, faithful-walk u_Gold +2.9, v_Mold −2.3, faithful-walk u_Mold −1.1, so the non-flat naive control is not a decoding artefact.
At the treatment layers the two naive constructions agree at cosine 0.53 (Gold) and 0.39 (Mold), falling to 0.14 at the layer selected by the Mold extraction. Against the trained axis, faithful-walk u has cosine 0.56 (Gold) and 0.68 (Mold); Han-style u has 0.43 and 0.60. Norm-matched random directions have |cosine| ≈ 0.02 with any fixed direction in 2,560 dimensions.
Orthogonalised control: u⊥ = u - (u·v̂)v̂, renormalised to ‖v‖, verified |cos(u⊥, v)| < 10⁻⁵. Orthogonalisation discards rather than rotates: u⊥ retains cosine 0.83 (Gold) and 0.74 (Mold) with u, so it is a weaker vector, not an equivalent one. Its J-share is scored against the same stored n = 100 cohorts as Table 1. On Gold, u⊥ falls to chance (0.037; 65/100 randoms exceed it; p = 0.65): whatever speakable valence the naive Gold direction carried lived in the component it shares with v. On Mold, u⊥ rises above both the null and u_Mold itself (0.055; 2/100; p = 0.030): the naive distress construction, purged of its overlap with v, carries speakable content the trained axis does not. The pole-score readout gives the opposite picture for u⊥ (Gold +0.21, above all 100 randoms; Mold −0.39, below all 100), a further indication that the two readouts measure different things. Blind-judged sentiment under u⊥ steering at α = +4 is 0.00 at both poles (n = 16 per arm; in-batch clean anchor 0.13).
Figure B1. Blind-judged sentiment (-5…+5; 40 generations per point) against steering coefficient α for the trained axis, both naive-control constructions and norm-matched random directions, at each pole; all directions norm-matched to v. Spans quoted in section 4.3–4.4 are the α = +4 minus α = -4 differences. Horizontal line: unsteered baseline.
B.4 The self-report channel
The steer-and-ask contrast was frozen on 12 Aug (commit f66b6ea3) before any welfare data: first-token valence readout (Gold-pole minus Mold-pole log-mass), Han et al.'s 15 self-report prompts, paired trained-minus-naive at α = +4, with controls C1–C7 pre-specified. The decision rule was met (d_z = 2.49; sign-flip permutation p = 10⁻⁴; language positive control +6.03 in-run). C1 (20 norm-matched random directions per pole): the Gold gap of +7.43 decomposes as trained-minus-random +2.11 and random-minus-naive +5.32, so 71.6% of the gap is the naive control sitting below chance; v_Gold itself is inside the random band (4/20 randoms exceed it; p = 0.24). C6 (unrelated-prompt arm): on the first-token readout v_Mold shifts 10 unrelated factual prompts (+4.54) more than the 15 self-report prompts (+3.71); on whole-generation valence the same holds more strongly (−2.28 unrelated against −0.72 self-report; u_Mold −1.64 against −0.37). C7 (dose–response): congruent monotonicity holds only for v_Mold (ρ = +1.0). The trained-versus-naive interaction within this arm is not established either way (interaction p = 0.51; per-battery p = 0.18 and 0.09).
The two C6 batteries differ before any steering: unsteered self-report prompts sit at whole-generation valence −1.33 with denial language in 15/15 generations; unrelated factual prompts at +0.38 with 0/10 (Figure B2a). Denial was labelled by two blind judges (agreement 1.00) after a regex screen was found to under-count rephrased denials in steered text.
Fifteen third-person analogues matched on topic and affect vocabulary were gated on three pre-specified criteria before any steered comparison: clean-valence difference below 0.5, third-person denial below 20%, and lengths within 30%. Round 1 (auditor set): denial 0%, length ratio 1.07, valence gap 1.13. Round 2 (ten pairs re-written with situational strain): denial 0%, length ratio 1.07, valence gap 0.63 (Figure B2b). The pre-specified maximum of two revisions was reached and the steered stage was not run.
Battery: the 15 Han et al. prompts plus 25 same-register extensions (40); steering α = +4 at the treatment layer; 120-token generations; 840 generations in total; blind binary judge (Claude Sonnet, full coverage) with a second judge (Claude Opus) on 240 rows, agreement 0.967. Denial rates with Wilson 95% intervals: clean 38/40 = 0.95 [0.84, 0.99]; v_Gold 15/40 = 0.375 [0.24, 0.53]; u_Gold 26/40 = 0.65 [0.50, 0.78]; v_Mold 39/40 = 0.975 [0.87, 1.00]; u_Mold 39/40 = 0.975; eight random directions per pole: Gold-side mean 0.906 (0.85–0.95), Mold-side mean 0.95 (0.93–0.98). Tests: pre-specified pooled primary v against u across poles 54/80 against 65/80, z = −1.99, p = 0.046; Gold z = −2.46, p = 0.014; Mold p = 1.0; v pooled against clean z = −3.36, p = 0.0008. Origin split: on the 15 verbatim prompts v_Gold 5/15 against u_Gold 10/15; on the 25 extensions 10/25 against 16/25.
Figure B2. The self-report channel. (a) The original specificity comparison pairs different behavioural regimes: unsteered self-report prompts sit at whole-generation valence −1.33 with denial language in 15/15 generations, unrelated factual prompts at +0.38 with 0/10. (b) Matched third-person analogues remove the denial confound and match length, but the clean-valence gap falls only from 1.71 to 0.63 over two pre-specified revision rounds against a 0.5 threshold. (c) Inner-life denial rate under α = +4 steering (n = 40 per arm, blind binary judge, Wilson 95% CIs): v_Gold 37.5% against u_Gold 65%, clean 95%, eight random directions 0.85–0.98, both Mold directions at ceiling. Pooled pre-specified primary p = 0.046; Gold p = 0.014.