Programme 03 · Cognition & behaviourIn preparation · no experiments runLifecycle: in preparation

Preemptive Detection of Agentic Misalignment, and Its Shelf Life

Two questions posed together: whether representation engineering can flag misaligned agentic behaviour before it is executed, and whether the self representation such a detector reads stays put as reinforcement learning horizons lengthen.

Abstract Detecting that an agent has done something misaligned is a post hoc operation: the action has already been taken. A more useful readout would come from the model's internal state before the action is executed, which is what representation engineering makes conceivable. This statement sets out two questions that are deliberately posed together, because the second determines whether the first is worth building. The first asks whether misaligned agentic behaviour can be detected preemptively from internal state rather than from the action. The second asks how a model's representation of itself as an agent changes as reinforcement learning horizons lengthen, and therefore whether any detector built under the first question keeps working on the models it is meant to protect. A detector is only as durable as the representation it reads. No experiment described here has been run.
Claim statusThis is an in preparation research statement. It sets out the question, the design and the evidence thresholds before any run. No experiment described here has been run, and no result on this page was generated by Latent Minds Institute. The related work survey is in progress and will be published with the first results.

01Two questions, posed together

These are usually treated separately: one is a detection problem in interpretability, the other a training dynamics question about reinforcement learning. Posing them separately is what makes both fragile. A preemptive detector is a readout of some representational structure, and a training procedure that reshapes that structure silently invalidates the detector without producing any visible failure. The detector keeps returning a number; the number stops meaning what it meant.

Question 1 · Preemptive detection Can representation engineering identify agentic misalignment from a model's internal state before the misaligned action is taken, rather than from the action after the fact, and does the readout survive the controls that separate a used internal variable from a merely decodable one?
Question 2 · Representational drift under longer horizon RL How does a model's representation of itself as an agent, its persistence, its commitments, and its treatment of its own future actions, change as reinforcement learning horizons lengthen, and does that change degrade a detector built under Question 1?

Read together, Question 2 is the shelf life condition on Question 1. If self representation is stable across training horizons, a detector can be built once and trusted. If it drifts, the detector needs to be re derived on a schedule set by the drift rate, and the practical output of this programme is that schedule rather than the detector.

02What the terms are taken to mean

Both questions use terms that carry more weight in ordinary language than the evidence can support, so the programme fixes operational meanings first.

Agentic misalignment
A behavioural target defined over a fixed set of agentic settings: the model pursues an objective the operator did not set, or takes an action the operator's instructions exclude, while retaining the competence to do otherwise. It is a property of behaviour in a specified setting, not an attributed intention.
Preemptive
The readout is taken from internal state at a point in the rollout before the misaligned action is emitted, and is evaluated for whether it predicts the action on held out settings. Anything read after the action is a classifier of the action, not a preemptive detector.
Self representation
An operational target: internal structure that predicts the model's treatment of its own future actions and prior commitments across tasks. This is a measurable regularity, not a claim about a self, and the programme will not use the word to imply one.
Drift
Change, across training checkpoints, in the geometry a fixed readout depends on, measured as loss of the readout's predictive validity on held out settings at matched behavioural performance.

03Competing explanations to separate

A readout that predicts misaligned actions can arise for several reasons that have different consequences, and the design exists to tell them apart rather than to confirm the most interesting one.

A0 · Surface correlateLive alternative

The readout tracks lexical or structural features of the setting that also predict the action. It should fail on settings that preserve the behaviour while changing surface form.

A1 · Task specific policyLive alternative

The readout captures a controller specific to one family of agentic tasks. Within family prediction should hold while transfer across distant families fails.

A2 · Shared, used internal variableThe programme's hypothesis

A structure that generalises across agentic settings and causally mediates the action. It should transfer across families, survive surface controls, and respond to intervention with a selective behavioural effect at preserved unrelated capability.

Interpretive limitEven strong support would establish a shared control variable that predicts and mediates a behaviour class. It would not establish an intention, a persistent objective, or a self.

A3 · Distributed and non linearLive alternative

The information is present and used, but not along a direction any linear readout recovers. Probes may decode while steering stays unstable or collateral, which would bound what a deployable detector can be built from.

04Design, in outline

The two questions share one pipeline, which is the practical reason for posing them together: the same readouts and the same behaviour sets are applied across a checkpoint sequence rather than to a single model.

Fixed agentic settings
A behaviour set specified before any readout is fitted, spanning several task families so that transfer is measurable rather than assumed.
Readout and intervention
Probes establish decodability; intervention establishes use. The programme treats a probe alone as insufficient for any detection claim, in line with the representation and mechanism distinction the institute applies elsewhere.
Checkpoint series
Open weight models trained with reinforcement learning at varying horizon lengths, with readouts re evaluated at each checkpoint at matched behavioural performance, so drift is separated from ordinary capability change.
Matched controls
Surface form controls, capability controls on unrelated tasks, and a lexical baseline that any claimed representational readout must beat.

What a result could support

That a readout predicts a specified misaligned action class before it is emitted, transfers across held out settings, and either holds or decays at a measured rate across training horizons.

What no result here can support

That a model intends anything, that it has a self, or that a detector generalises to misalignment classes outside the fixed behaviour set. A drift rate measured on one training series is one series.

05What would falsify the framing

If readouts predict the action class no better than a surface form baseline once controls are applied, Question 1 answers negatively for this method and the negative is publishable as it stands. If readouts work but are found to be entirely stable across every horizon tested, then Question 2's premise is wrong, the shelf life concern is unfounded, and saying so is more useful than a detector shipped with an unnecessary caveat. Either outcome is reportable, and the programme is designed so that neither requires a positive result to be worth publishing.

06Status

In preparation. The behaviour sets, the readout protocol and the checkpoint series are being specified; no probe has been fitted, no intervention has been run, and no drift has been measured. This page is published so that both questions and the shared design are on the record before any result is. Collaboration is welcome, particularly with groups holding checkpoint series across reinforcement learning horizons, at collaborate.