# Latent Minds Institute, full public research context This file is generated from data/research-objects.json. Preserve every object's status and evidence limits when quoting or summarizing it. ## Safeguard Durability in Open Weight Models - Slug: open-weight-safeguard-durability - Canonical URL: https://latentmindsinstitute.com/research/open-weight-safeguard-durability/ - API URL: https://latentmindsinstitute.com/api/research-objects/open-weight-safeguard-durability.json - Type: Research statement (statement) - Lifecycle: in-preparation - Epistemic status: design-only - Status statement: In preparation; design, cost measure and controls specified; no attack run, no model evaluated, no result - Programme: deployment - Date: 2026-07-28 - Version: v0.1 - Authors: Latent Minds Institute - Research question: When a model's weights are public, how much adversarial effort does each published safeguard actually impose, and how much of that cost survives a modest fine tuning budget? - Evidence: none; this is a design statement published before any run - Provenance: Published in advance of results so the question, the cost measure and the controls are on the record before any finding is - Methods: adversarial red teaming across input, decoding, weight modification and composed surfaces (planned); attack cost curves against a fixed behaviour set (planned); capability controls separating a defeated safeguard from a degraded model (planned); attacker effort controls so null results read as bounded search (planned) - Models: published open weight models (planned) - Code: not publicly linked - Data: not publicly linked - Related slugs: model-entrenchment ## Preemptive Detection of Agentic Misalignment, and Its Shelf Life - Slug: agentic-misalignment-detection - Canonical URL: https://latentmindsinstitute.com/research/agentic-misalignment-detection/ - API URL: https://latentmindsinstitute.com/api/research-objects/agentic-misalignment-detection.json - Type: Research statement (statement) - Lifecycle: in-preparation - Epistemic status: design-only - Status statement: In preparation; two adjacent questions and a shared design specified; no probe fitted, no intervention run, no drift measured - Programme: cognition - Date: 2026-07-28 - Version: v0.1 - Authors: Latent Minds Institute - Research question: Can representation engineering detect agentic misalignment from internal state before the action is taken, and does the self representation such a detector reads remain stable as reinforcement learning horizons lengthen? - Evidence: none; this is a design statement published before any run - Provenance: The two questions are posed together because the second sets the shelf life of any detector built under the first - Methods: readouts over a fixed set of agentic settings, taken before the action is emitted (planned); intervention to separate a used internal variable from a merely decodable one (planned); re-evaluation across a reinforcement learning checkpoint series at matched behavioural performance (planned); surface form, capability and lexical baseline controls (planned) - Models: open weight models with reinforcement learning checkpoint series (planned) - Code: not publicly linked - Data: not publicly linked - Related slugs: evaluation-state, latent-signatures-strategic-games ## Construct map - Slug: construct-map - Canonical URL: https://latentmindsinstitute.com/research/construct-map/ - API URL: https://latentmindsinstitute.com/api/research-objects/construct-map.json - Type: Methods note (note) - Lifecycle: n/a - Epistemic status: methods-note - Status statement: Methods documentation; no results, no numbers - Programme: cognition - Date: 2026-07-18 - Version: v1 - Authors: Latent Minds Institute - Research question: For each strategic game family, what construct is measured, what control does a reading require, and which claim does the evidence never license? - Evidence: documentation of method only; contains no measurements and states none - Provenance: Written for the latent arena game families behind Latent Signatures in Strategic Games - Methods: construct specification per game family; required controls stated before code; explicit never claim boundaries - Models: none stated - Code: not publicly linked - Data: not publicly linked - Related slugs: latent-signatures-strategic-games ## Latent Signatures in Strategic Games - Slug: latent-signatures-strategic-games - Canonical URL: https://latentmindsinstitute.com/papers/latent-signatures-strategic-games/ - API URL: https://latentmindsinstitute.com/api/research-objects/latent-signatures-strategic-games.json - Type: Working paper (working-paper) - Lifecycle: pilot - Epistemic status: research-programme - Status statement: Research programme; behavioural fleet measured across six models in negotiation and four in social deduction; latent analysis proposed, none run - Programme: cognition - Date: 2026-07-11 - Version: v0.2 - Authors: Latent Minds Institute - Research question: Do deception, opponent-modelling, and evaluation-awareness signatures in an open model's activations predict and cause its moves in strategic games? - Evidence: research programme with three signature hypotheses and a three-stage proposed pipeline; the latent-arena harness is built and tested, the behavioural fleet is measured (negotiation self-play across six models, social deduction with four model tables), no activations captured - Provenance: Built for Latent Minds Institute; the harness is the latent-arena repository, being prepared for public release and available on request until then - Methods: self-hosted game environments; behavioural self-play (measured); instrumented self-play (proposed); signature probes (proposed); steering during play (proposed) - Models: open-weight models (proposed) - Code: not publicly linked - Data: not publicly linked - Related slugs: evaluation-state, model-entrenchment ## Circuit Traces: Attribution Graphs for Open Models - Slug: circuit-traces - Canonical URL: https://latentmindsinstitute.com/instruments/circuit-traces/ - API URL: https://latentmindsinstitute.com/api/research-objects/circuit-traces.json - Type: Instrument (instrument) - Lifecycle: instrument - Epistemic status: instrument - Status statement: Live attribution graphs on Gemma 2 2B; hypotheses to confirm with interventions, stated as such - Programme: computation - Date: 2026-07-11 - Version: v1 - Authors: Latent Minds Institute - Research question: Which features caused which, from prompt to prediction, and does the causal story survive comparison across tasks, models, and methods? - Evidence: real causal graphs served live; replacement-model caveat stated in interface - Provenance: Method: Ameisen et al. / Lindsey et al. (2025), Transformer Circuits; implementation: safety-research/circuit-tracer; hosting: Neuronpedia - Methods: attribution graphs on transcoder replacement models; curated case-study reading guides; validated graph loading - Models: Gemma 2 2B (Gemma Scope transcoders) - Code: https://github.com/safety-research/circuit-tracer - Data: not publicly linked - Related slugs: latent-observatory, transformer-explainer ## Latent Observatory: SAE Feature Explorer - Slug: latent-observatory - Canonical URL: https://latentmindsinstitute.com/instruments/latent-observatory/ - API URL: https://latentmindsinstitute.com/api/research-objects/latent-observatory.json - Type: Instrument (instrument) - Lifecycle: instrument - Epistemic status: instrument - Status statement: Live model data via the open Neuronpedia API; explanations labelled as hypotheses - Programme: representations - Date: 2026-07-11 - Version: v1 - Authors: Latent Minds Institute - Research question: What has a sparse autoencoder actually learned about a model, and does the auto-interp story survive contact with the activation evidence? - Evidence: model data from the open Neuronpedia API - Provenance: Built by Latent Minds Institute on the open Neuronpedia API (MIT) - Methods: semantic search over auto-interp explanations; activation-record inspection; logit-effect readout; decoder-space nearest neighbours - Models: GPT-2 small (RES-JB); Gemma 2 2B (Gemma Scope res 16k) - Code: not publicly linked - Data: https://www.neuronpedia.org/api-doc - Related slugs: interpretability-map, evaluation-state ## Evaluation State in Language Models - Slug: evaluation-state - Canonical URL: https://latentmindsinstitute.com/papers/evaluation-state/ - API URL: https://latentmindsinstitute.com/api/research-objects/evaluation-state.json - Type: Working paper (working-paper) - Lifecycle: proposal - Epistemic status: conceptual-framework - Status statement: Research programme · five hypotheses, experiments proposed, none run - Programme: cognition - Date: 2026-07-10 - Version: v0.1 - Authors: Latent Minds Institute - Research question: What evidence would show that a model represents evaluation as a shared internal state rather than reacting to surface cues? - Evidence: literature synthesis and experimental framework; no original empirical results - Provenance: not stated in registry - Methods: layer-wise probes vs lexical baselines (proposed); cross-family and cross-task transfer (proposed); activation steering with capability controls (proposed); Jacobian-lens readouts (proposed) - Models: open-weight instruction-tuned models (proposed) - Code: not publicly linked - Data: not publicly linked - Related slugs: model-entrenchment, interpretability-map ## Model Entrenchment: Why Useful AI Systems Become Difficult to Remove - Slug: model-entrenchment - Canonical URL: https://latentmindsinstitute.com/papers/model-entrenchment/ - API URL: https://latentmindsinstitute.com/api/research-objects/model-entrenchment.json - Type: Working paper (working-paper) - Lifecycle: proposal - Epistemic status: conceptual-framework - Status statement: Conceptual framework · experiments proposed, none run - Programme: deployment - Date: 2026-07-01 - Version: v1.1 - Authors: Latent Minds Institute - Research question: When an AI system resists removal, where does the resistance live: in the model's internal representations and behaviour, or in the web of dependence around it? - Evidence: conceptual framework with falsification criteria; every empirical claim is a prediction - Provenance: not stated in registry - Methods: linear probes (proposed); activation steering (proposed); counterfactual environments (proposed); behavioural evaluations (proposed); case coding (proposed) - Models: open-weight models (proposed) - Code: not publicly linked - Data: not publicly linked - Related slugs: interpretability-map, gpu-credit-underwriting, gpu-forward-market ## The Interpretability Map - Slug: interpretability-map - Canonical URL: https://latentmindsinstitute.com/instruments/interpretability-map/ - API URL: https://latentmindsinstitute.com/api/research-objects/interpretability-map.json - Type: Research map · instrument (research-map) - Lifecycle: instrument - Epistemic status: literature-synthesis - Status statement: Literature synthesis · sources cited per node - Programme: representations - Date: 2026-07-01 - Version: v1 - Authors: Latent Minds Institute - Research question: What has mechanistic interpretability actually established, in what order, with what dependencies, and what should a new researcher do this week? - Evidence: synthesis of primary sources; contains no original experimental claims - Provenance: not stated in registry - Methods: literature dependency graph; chronology; tiered reading pathway; runnable TransformerLens/SAELens protocols; model-access matrix; open problems as experiments - Models: GPT-2 small; Gemma 2 2B (protocols) - Code: not publicly linked - Data: /instruments/interpretability-map/atlas.json - Related slugs: model-entrenchment, transformer-explainer ## Transformer Explainer: GPT-2 Live in the Browser - Slug: transformer-explainer - Canonical URL: https://latentmindsinstitute.com/instruments/transformer-explainer/ - API URL: https://latentmindsinstitute.com/api/research-objects/transformer-explainer.json - Type: Instrument · adaptation (instrument) - Lifecycle: instrument - Epistemic status: adaptation - Status statement: Adapted instrument · original by Cho et al., Georgia Tech Polo Club - Programme: computation - Date: 2026-07-10 - Version: v1 - Authors: Latent Minds Institute - Research question: What does a transformer's forward pass actually compute, step by step, on real input? - Evidence: pedagogical instrument running a real model; the 2026 commentary is editorial, not original research - Provenance: Adaptation of poloclub/transformer-explainer (MIT); previously hosted as Transformer Visualiser at mo3.ca - Methods: ONNX GPT-2 (small) inference in-browser; attention map inspection; temperature and sampling controls - Models: GPT-2 small (124M) - Code: https://github.com/poloclub/transformer-explainer - Data: not publicly linked - Related slugs: interpretability-map ## Underwriting the Machine: A Field Guide to GPU Credit Risk - Slug: gpu-credit-underwriting - Canonical URL: https://latentmindsinstitute.com/papers/gpu-credit-underwriting/ - API URL: https://latentmindsinstitute.com/api/research-objects/gpu-credit-underwriting.json - Type: Working paper · economics strand (working-paper) - Lifecycle: proposal - Epistemic status: conceptual-framework - Status statement: Analytical framework with illustrative models · not investment advice - Programme: deployment - Date: 2026-07-11 - Version: v1 - Authors: Latent Minds Institute - Research question: How should lenders price credit risk on GPU-collateralised debt, and what does that market structure imply for compute dependence? - Evidence: analytical framework with worked illustrative numbers; no proprietary deal data - Provenance: not stated in registry - Methods: credit risk decomposition; DSCR/LTV sizing models; residual-value analysis; scenario stress - Models: none stated - Code: not publicly linked - Data: not publicly linked - Related slugs: gpu-forward-market, model-entrenchment ## A Forward Market for GPU Compute - Slug: gpu-forward-market - Canonical URL: https://latentmindsinstitute.com/papers/gpu-forward-market/ - API URL: https://latentmindsinstitute.com/api/research-objects/gpu-forward-market.json - Type: Working paper · economics strand (working-paper) - Lifecycle: proposal - Epistemic status: conceptual-framework - Status statement: Market-design proposal · conceptual - Programme: deployment - Date: 2026-07-01 - Version: v1 - Authors: Latent Minds Institute - Research question: What would a functioning forward market for GPU compute look like, and what would it change about the economics of AI deployment? - Evidence: market-design proposal; no live market data beyond cited public sources - Provenance: not stated in registry - Methods: market design; contract specification; term-structure analysis - Models: none stated - Code: not publicly linked - Data: not publicly linked - Related slugs: gpu-credit-underwriting, model-entrenchment