Latent Minds

Research

Research objects

Published


Abstract editorial illustration of a latent direction separating into gold and teal axes beside a stream of small squares. It is not an experimental result.
Interpretability2026-08-16

Is functional welfare speakable?

A functional welfare axis is speakable before any RL, and training amplifies the speakable share of the axis, distress first.

Sprint paperRead →

In preparation


Safety2026-07-28

Safeguard Durability in Open Weight Models

Incoming.

In preparation
Model Cognition2026-07-28

Preemptive Detection of Agentic Misalignment, and Its Shelf Life

Incoming.

In preparation
Model Cognition2026-08-06

Distress, Flourishing, and Valence in Model Output

When models report distress, satisfaction, or engagement, whether those signals stay stable across framings, and whether they track anything behavioural or only the prompt.

In preparation
Model Cognition2026-08-06

Introspection and the Reliability of Model Self Report

Whether models can report their own internal states accurately, and whether structured prompting or interpretability methods make those reports more reliable.

In preparation