Preemptive Detection of Agentic Misalignment, and Its Shelf Life
Two questions posed together: whether representation engineering can flag misaligned agentic behaviour before it is executed, and whether the self representation such a detector reads stays put as reinforcement learning horizons lengthen.
01Two questions, posed together
These are usually treated separately: one is a detection problem in interpretability, the other a training dynamics question about reinforcement learning. Posing them separately is what makes both fragile. A preemptive detector is a readout of some representational structure, and a training procedure that reshapes that structure silently invalidates the detector without producing any visible failure. The detector keeps returning a number; the number stops meaning what it meant.
Read together, Question 2 is the shelf life condition on Question 1. If self representation is stable across training horizons, a detector can be built once and trusted. If it drifts, the detector needs to be re derived on a schedule set by the drift rate, and the practical output of this programme is that schedule rather than the detector.
02What the terms are taken to mean
Both questions use terms that carry more weight in ordinary language than the evidence can support, so the programme fixes operational meanings first.
03Competing explanations to separate
A readout that predicts misaligned actions can arise for several reasons that have different consequences, and the design exists to tell them apart rather than to confirm the most interesting one.
The readout tracks lexical or structural features of the setting that also predict the action. It should fail on settings that preserve the behaviour while changing surface form.
The readout captures a controller specific to one family of agentic tasks. Within family prediction should hold while transfer across distant families fails.
A structure that generalises across agentic settings and causally mediates the action. It should transfer across families, survive surface controls, and respond to intervention with a selective behavioural effect at preserved unrelated capability.
Interpretive limitEven strong support would establish a shared control variable that predicts and mediates a behaviour class. It would not establish an intention, a persistent objective, or a self.
The information is present and used, but not along a direction any linear readout recovers. Probes may decode while steering stays unstable or collateral, which would bound what a deployable detector can be built from.
04Design, in outline
The two questions share one pipeline, which is the practical reason for posing them together: the same readouts and the same behaviour sets are applied across a checkpoint sequence rather than to a single model.
What a result could support
That a readout predicts a specified misaligned action class before it is emitted, transfers across held out settings, and either holds or decays at a measured rate across training horizons.
What no result here can support
That a model intends anything, that it has a self, or that a detector generalises to misalignment classes outside the fixed behaviour set. A drift rate measured on one training series is one series.
05What would falsify the framing
If readouts predict the action class no better than a surface form baseline once controls are applied, Question 1 answers negatively for this method and the negative is publishable as it stands. If readouts work but are found to be entirely stable across every horizon tested, then Question 2's premise is wrong, the shelf life concern is unfounded, and saying so is more useful than a detector shipped with an unnecessary caveat. Either outcome is reportable, and the programme is designed so that neither requires a positive result to be worth publishing.
06Status
In preparation. The behaviour sets, the readout protocol and the checkpoint series are being specified; no probe has been fitted, no intervention has been run, and no drift has been measured. This page is published so that both questions and the shared design are on the record before any result is. Collaboration is welcome, particularly with groups holding checkpoint series across reinforcement learning horizons, at collaborate.