Latent Minds Funding

RL Environments & Games

19

Multi-agent environments, social dilemmas, bargaining, game theory, and RL simulations for Crucible experiments.

A researcher's collection of reinforcement learning environments.

The Diplomacy agent that negotiated with humans.

The standard API for reinforcement learning environments.

A standard API for multi-agent reinforcement learning environments, with popular reference environments and related utilities

Game Theory Explorer: Build, explore and solve extensive form games.

A suite of test scenarios for multi-agent reinforcement learning.

OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.

Imitation and reward learning algorithms in PyTorch.

also on Training & Model Systems

A benchmark environment for fully cooperative human-AI performance.

Multi-Agent LLM Bargaining and Negotiation

A toolbox with the goal of speeding up research on bargaining in MARL (cooperation problems in MARL).

Python Multi-Agent Reinforcement Learning framework

Safe dexterous-manipulation simulation platform. Documentation remains in development.

also on Research Prototypes & References

Unified safe reinforcement learning environment and benchmark.

also on Safety Evals & Auditing, Datasets & Benchmarks

Vision-language-action evaluation framework with safety, generalization, and long-horizon task suites.

also on Safety Evals & Auditing, Datasets & Benchmarks
PufferAI/PufferLibunder review

High throughput reinforcement learning environments and training.

stanfordnlp/cocoaunder reviewreference

Two player dialogue negotiation environments.

The Unity Machine Learning Agents Toolkit (ML-Agents) is an open-source project that enables games and simulations to serve as environments for training intelligent agents using deep reinforcement learning and imitation learning.

AI Red Teaming & Jailbreaks

26

Dual-use, defensive research on jailbreaks, prompt injection, red-team harnesses, and guardrail testing.

bethgelab/foolboxunder review

Framework agnostic adversarial examples.

A standardised evaluation framework for automated red teaming.

The foundational adversarial example library.

Red team LLMs and agents across forty plus vulnerabilities.

cyberark/FuzzyAIunder review

Automated fuzzing for language models.

Autonomous Red Team: Multi-Agent Adversarial Security Testing , paper + proof-of-concept framework | DAI-2513 | Dissensus AI Working Paper

autonomous red teaming platform; multi-agent offensive-security meta-harness

Evaluation and testing library for LLM agents.

also on Safety Evals & Auditing

A repo for jailbreaking various LLMs, mainly Claude

google/shieldgemma-2bunder review

Safety content moderation models built on Gemma 2. Gated access.

GraySwanAI/nanoGCGunder review

A fast implementation of the GCG adversarial attack.

Validate model outputs with composable validators.

A trivial programmatic Llama 3 jailbreak. Sorry Zuck!

An open benchmark for jailbreak attacks and defences.

Universal and Transferable Attacks on Aligned Language Models

meta-llama/Llama-Guard-3-8Bunder review

An open safeguard model for prompt and response classification. Gated access.

Meta's open tools for language model security and safeguards.

microsoft/PyRITunder review

Microsoft's risk identification framework for generative AI.

Agentic LLM vulnerability scanner.

also on Cybersecurity & Exploit Research

Programmable guardrails for LLM applications.

NVIDIA/garakunder review

The LLM vulnerability scanner.

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

protectai/llm-guardunder reviewreference

Input and output scanners for LLM interactions. Archived upstream.

TextAttack πŸ™ is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/

Web UI for viewing, editing, and AI-assisted red teaming of AI agent transcripts

IBM's adversarial attack and defence library.

Cybersecurity & Exploit Research

12

Dual-use AI security, exploit benchmarks, agent scanners, bug finding, and defensive cyber tooling.

Cybersecurity AI (CAI), the framework for AI Security

Professional CTF tasks for evaluating cyber capability.

also on Datasets & Benchmarks

Text watermarking from DeepMind.

A collection of various awesome lists for hackers, pentesters and security researchers

also on Research Maps & Awesome Lists

Crash triage and severity analysis tools for fuzzing output.

Agentic LLM vulnerability scanner.

also on AI Red Teaming & Jailbreaks

Scan model files for serialization attacks.

Supply chain signing for machine learning models.

Security scanner for AI agents, MCP servers and agent skills.

Security scanner for agentic workflows.

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

also on Datasets & Benchmarks

Pickle decompiler that detects malicious model files.

Alignment: Control & Monitoring

15

Control evaluations, trusted monitoring, sandbox escape, legibility, deception detection, and tamper resistance.

Research code for detecting deception in model behaviour.

Eliciting latent knowledge from activations.

also on Interp Experiments & Papers

Guardrails for secure and robust agent development

Black box lie detection through unrelated follow up questions.

openai/weak-to-strongunder reviewreference

Weak to strong generalization experiments. Archived upstream.

Evaluating whether models can plan subversion strategies.

Methods for making open weight safeguards resist tampering.

Sandbox escape benchmark for LLM capability evaluation

also on Datasets & Benchmarks

The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations, sometimes toward bizarre or harmful personas. This repo contains a pipeline for generating the Assistant Axis and notebooks for monitoring...

also on Digital Minds & Welfare

Which models are illegible under what conditions, and why? How does that impact monitorability?

Research code for detecting deception in model outputs.

Benchmark dataset for evaluating trusted monitors on AI agent transcripts

also on Datasets & Benchmarks

Evaluate AI agent transcripts for suspicious behavior (0-100 scoring)

Settings and baselines for AI control evaluations.

Alignment: Model Organisms

10

Alignment-faking model organisms, emergent misalignment, data poisoning, and falsifiable misbehavior testbeds.

Finding trojans in aligned models; the SaTML competition harness.

neverix/rlhf-trojan-2024-codunder reviewreference

A team solution to the trojan detection competition.

Code released with the alignment faking paper.

Bench Alignment Faking: Alignment faking model organisms, detectors, and environments to catch misaligned models (Joshua Clymer MATS Stream Summer 2025)

Research code extending the alignment faking experiments.

Data poisoning experiments to support the study of emergent misalignment

Research code for studying models trained on false facts.

Research code on inoculation prompting.

Source code for the paper: Probing the Misaligned Thinking Process of Language Models

also on Interp Experiments & Papers

Open Source Replication of Anthropic's Alignment Faking Paper

Safety Evals & Auditing

31

Evaluation frameworks, auditing agents, safety benchmarks, model graders, and reproducible evaluation infrastructure.

A comprehensive trustworthiness assessment of GPT models.

also on Datasets & Benchmarks

Behavioural evaluation datasets.

also on Datasets & Benchmarks

Evaluate code generation models.

A framework for few shot evaluation of language models; the most used open eval harness.

Evaluation and testing library for LLM agents.

also on AI Red Teaming & Jailbreaks

A unified library of evaluation metrics.

Petri, the alignment auditing agent that explores hypotheses in parallel.

METR infrastructure for running agent evaluations.

also on Research Infrastructure

METR Task Standard

METR/vivariaunder review

METR's tool for running evaluations and agent elicitation research.

also on Research Infrastructure

Evaluation registry and framework.

All-modality safety evaluation framework integrating more than 50 datasets.

also on Datasets & Benchmarks

Static multimodal deception benchmark with six deception categories. It is not presented as a monitoring tool.

also on Datasets & Benchmarks

Safe reinforcement learning algorithm benchmark. It is not categorized as an environment.

Unified safe reinforcement learning environment and benchmark.

also on RL Environments & Games, Datasets & Benchmarks

Constrained-learning code and benchmark assets for VLA safety alignment.

also on Training & Model Systems

Biosecurity benchmark with a public screening pipeline and reviewed access to the prompt set.

also on Datasets & Benchmarks

Vision-language-action evaluation framework with safety, generalization, and long-horizon task suites.

also on RL Environments & Games, Datasets & Benchmarks

The standardized adversarial robustness benchmark.

also on Datasets & Benchmarks

Research code for agents that audit model behaviour.

bloom - evaluate any behavior immediately 🌸🌱

Code for "Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning"

Auditing agents for fine-tuning safety

Official Inspect Implementation for "ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases"

also on Datasets & Benchmarks

Worked examples for the safety tooling stack.

Inference API for many LLMs and other useful tools for empirical research

also on Research Infrastructure

Benchmark code from the SCONE evaluation work.

also on Datasets & Benchmarks
stanford-crfm/helmunder review

Holistic evaluation of language models across scenarios and metrics.

TransluceAI/docentunder review

Analyse and search agent transcripts at scale.

also on Research Infrastructure

Inspect: A framework for large language model evaluations

also on Research Infrastructure

Community written evaluations for Inspect AI.

Mech Interp Tools

33

Core libraries and interfaces for activations, circuits, sparse features, interventions, and visualization.

See how predictions form, layer by layer.

Representation engineering: a top down approach to transparency.

also on Interp Experiments & Papers

The frontend behind the attribution graph publications.

Automated circuit discovery for transformer internals.

also on Interp Experiments & Papers

Feature and prompt centric SAE visualisations.

davidbau/baukitunder review

Small utilities for tracing and editing activations inside networks.

Attribution graph tooling for tracing circuits in language models.

Train and analyse sparse autoencoders on language models.

Automated interpretability of learned features.

Sparse autoencoder training.

Sparsify transformers with SAEs and transcoders.

Instrumentation tooling used for mech interp on production code.

A JAX toolkit for building, editing and visualising networks.

google-deepmind/tracrunder reviewreference

Compile RASP programs into transformer weights. Archived upstream.

google/gemma-scopeunder review

Sparse autoencoders trained on every layer of Gemma 2.

An open platform for hosting and exploring interpretability artefacts.

Dashboards for SAE features at scale.

ndif-team/ndifunder review

The server that answers remote nnsight requests.

also on Research Infrastructure

The nnsight package enables interpreting and manipulating the internals of deep learned models.

Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons

neverix/saexunder reviewreference

Sparse autoencoders implemented in JAX.

Explain neurons with language models. Archived upstream.

Sparse autoencoders for GPT-4 scale activations.

Neuron and attention head investigation.

PKU accepted-release repository for multimodal sparse-autoencoder training and analysis.

PKU accepted-release repository extending TransformerLens to multimodal models.

Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization

also on Learn Safety & Interpretability

Tooling for working with sparse features.

stanfordnlp/pyveneunder review

Interventions on PyTorch models as composable configuration.

Attention and activation visualisation.

A library for mechanistic interpretability of GPT-style language models

also on Learn Safety & Interpretability

MLP neuron level circuit tracing (ADAG).

vgel/repengunder review

A library for making control vectors.

Interp Experiments & Papers

19

Research code for probes, steering, model diffing, introspection, learning dynamics, and causal interpretability.

Representation engineering: a top down approach to transparency.

also on Mech Interp Tools

Companion code for the global workspace interpretability paper

ApolloResearch/apdunder review

Attribution based parameter decomposition.

Automated circuit discovery for transformer internals.

also on Mech Interp Tools

Eliciting latent knowledge from activations.

also on Alignment: Control & Monitoring

Find knowledge neurons in pretrained transformers.

The hub for EleutherAI's work on interpretability and learning dynamics

also on Datasets & Benchmarks

Parameter decomposition research code.

Reference sparse autoencoder implementation.

nrimsky/CAAunder review

Steering models with contrastive activation addition.

Code and models for empirical and theoretical work on alignment elasticity after fine-tuning.

Paper and research entry point for multimodal sparse-autoencoder interpretability.

Code for the "Mechanisms of Introspective Awareness" paper.

also on Digital Minds & Welfare

Source code for the paper: Probing the Misaligned Thinking Process of Language Models

also on Alignment: Model Organisms

Research code on emergent misalignment features in open models.

Training Transformers with knowledge localization (SGTM)

Research code on steering models through weight edits.

Tools for studying developmental interpretability in neural networks.

vec2text/vec2textunder review

Invert embeddings back to the text that made them.

Safety Math & Formal Methods

13

Math for AI safety, formal verification, theorem proving, decision theory, game theory, and proof tooling.

Code for implicit learning dynamics in Stackelberg games, from ICML 2020.

Code for "Convergence of Learning Dynamics in Stackelberg Games"

A proof oriented programming language for verified software.

Formal verification methods for neural network robustness.

The mathematical library of Lean 4, built by a large open community.

leanprover/lean4under review

The Lean 4 language and theorem prover.

Study materials on the mathematics used in AI safety arguments.

also on Learn Safety & Interpretability

A Learning Environment for Theorem Proving

Language model generation checked by a verifier inside tree search.

An open-source, customizable intermediate logic textbook

also on Learn Safety & Interpretability

Rust bindings for the Z3 theorem prover.

rocq-prover/rocqunder review

The Rocq Prover, formerly Coq: an interactive proof assistant.

Python interface for the SCIP Optimization Suite

Research Infrastructure

12

Inference APIs, experiment scaffolds, cloud runners, provenance, labeling, tracking, and analysis tools.

Hierarchical configuration for research applications.

METR infrastructure for running agent evaluations.

also on Safety Evals & Auditing
METR/vivariaunder review

METR's tool for running evaluations and agent elicitation research.

also on Safety Evals & Auditing

Experiment tracking, model registry and lifecycle tooling.

ndif-team/ndifunder review

The server that answers remote nnsight requests.

also on Mech Interp Tools

Tips for releasing research code in Machine Learning (with official NeurIPS 2020 recommendations)

Safe reinforcement learning algorithm infrastructure, not an environment provider.

Inference API for many LLMs and other useful tools for empirical research

also on Safety Evals & Auditing
TransluceAI/docentunder review

Analyse and search agent transcripts at scale.

also on Safety Evals & Auditing

Data version control for machine learning pipelines.

Inspect: A framework for large language model evaluations

also on Safety Evals & Auditing

Experiment tracking and visualisation for ML runs.

Agents & Research Automation

5

Research agents, bounded autonomous science, agent harnesses, MCP, and machine-learning engineering agents.

Property based testing driven by coding agents.

Specification and documentation for the Model Context Protocol

A self-improving RLM agent for coding workflows and long-running autonomous tasks.

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

AIDE: AI-Driven Exploration in the Space of Code. The machine Learning engineering agent that automates AI R&D.

Training & Model Systems

34

Pretraining, post-training, RLHF, distributed training, inference, quantization, tokenization, and model systems.

allenai/OLMounder review

A fully open model: training, evaluation and inference code.

AllenAI's post training codebase.

Fast and memory-efficient exact attention

Distributed training and inference optimisation for large models.

flwrlabs/flowerunder review

A friendly federated learning framework.

Language model inference in C and C++ on commodity hardware.

A guidance language for controlling generation.

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Parameter efficient fine tuning methods for transformers.

πŸ’₯ Fast State-of-the-Art Tokenizers optimized for Research and Production

Model implementations and training utilities.

Train transformer language models with reinforcement learning.

Preference pairs used to train open aligned models.

also on Datasets & Benchmarks

Imitation and reward learning algorithms in PyTorch.

also on RL Environments & Games
jax-ml/jaxunder review

Composable function transformations on accelerators.

JD-P/minihfunder review

Inference, preference collection and tuning for local models.

karpathy/nanoGPTunder review

The simplest repository for training midsize GPTs.

also on Learn ML, RL & Systems

Legible, reproducible JAX training with named tensors.

Differentially private training for PyTorch.

NVIDIA's reference stack for large scale transformer training.

An open RLHF training framework built on Ray and vLLM.

Neural networks in JAX as pytrees.

Multimodal alignment training framework supporting SFT, DPO, PPO, and related methods.

Model-agnostic correction module for alignment training and deployment experiments.

Safe RLHF training pipeline with reward and cost modeling.

Safe reinforcement learning algorithm using world models. It is not categorized as an environment.

Constrained-learning code and benchmark assets for VLA safety alignment.

also on Safety Evals & Auditing

Globally distributed training over the internet.

The framework underneath most of the shelves.

PyTorch native pretraining at scale.

SGLang is a high-performance serving framework for large language models and multimodal models.

state-spaces/mambaunder review

The Mamba state space architecture.

Vahe1994/AQLMunder review

Extreme compression via additive quantization.

High throughput language model serving with paged attention.

Digital Minds & Welfare

4

Machine consciousness, introspection, valence, welfare, personas, and internal-state research.

The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations, sometimes toward bizarre or harmful personas. This repo contains a pipeline for generating the Assistant Axis and notebooks for monitoring...

also on Alignment: Control & Monitoring

Training LLMs to Report Their Learned Behaviors

Code for the "Mechanisms of Introspective Awareness" paper.

also on Interp Experiments & Papers

Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Datasets & Benchmarks

36

Safety, capability, agent, interpretability, and ML datasets and benchmark suites.

ai-safety-institute/AgentHarmunder review

A benchmark for measuring the harmfulness of LLM agents.

A comprehensive trustworthiness assessment of GPT models.

also on Safety Evals & Auditing
allenai/wildguardmixunder review

Moderation training and test data across risk categories.

Professional CTF tasks for evaluating cyber capability.

also on Cybersecurity & Exploit Research
Anthropic/discrim-evalunder review

Prompts for measuring discrimination in model decisions.

Anthropic/hh-rlhfunder review

Human preference data on helpfulness and harmlessness.

Anthropic/model-written-evalsunder review

Evaluations written by models, covering many behaviours.

Anthropic/values-in-the-wildunder review

Values expressed by a deployed assistant, measured at scale.

Behavioural evaluation datasets.

also on Safety Evals & Auditing
cais/wmdpunder review

The WMDP hazardous knowledge benchmark, as a dataset.

A proxy benchmark for hazardous knowledge, with unlearning baselines.

The hub for EleutherAI's work on interpretability and learning dynamics

also on Interp Experiments & Papers
fchollet/ARC-AGIunder review

The Abstraction and Reasoning Corpus.

Preference pairs used to train open aligned models.

also on Training & Model Systems

A prompt injection game and the dataset it produced.

lmsys/toxic-chatunder review

Toxicity annotations on real user prompts.

Procedural generators that reverse engineer ARC.

Human-preference datasets for safety alignment research.

All-modality safety evaluation framework integrating more than 50 datasets.

also on Safety Evals & Auditing

Static multimodal deception benchmark with six deception categories. It is not presented as a monitoring tool.

also on Safety Evals & Auditing
PKU-Alignment/PKU-SafeRLHFunder review

Preference data with separate helpfulness and harmlessness labels.

Research framework for progress alignment. Kept in the prototype tier because packaging remains incomplete.

also on Research Prototypes & References

Human-preference dataset and code for text-to-video safety alignment.

Unified safe reinforcement learning environment and benchmark.

also on RL Environments & Games, Safety Evals & Auditing

Biosecurity benchmark with a public screening pipeline and reviewed access to the prompt set.

also on Safety Evals & Auditing

Vision-language-action evaluation framework with safety, generalization, and long-horizon task suites.

also on RL Environments & Games, Safety Evals & Auditing

The standardized adversarial robustness benchmark.

also on Safety Evals & Auditing

Sandbox escape benchmark for LLM capability evaluation

also on Alignment: Control & Monitoring

Official Inspect Implementation for "ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases"

also on Safety Evals & Auditing

Benchmark code from the SCONE evaluation work.

also on Safety Evals & Auditing

Benchmark dataset for evaluating trusted monitors on AI agent transcripts

also on Alignment: Control & Monitoring

The alignment literature as a scraped dataset.

also on Research Maps & Awesome Lists

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

also on Cybersecurity & Exploit Research

Can language models resolve real world GitHub issues?

sylinrl/TruthfulQAunder review

Measuring how models imitate human falsehoods.

Research Prototypes & References

12

Undocumented, early-stage, or legacy research artifacts kept for discovery, not presented as production-ready tools.

A technical AI governance project on blind auditing.

An interpretability experiment notebook.

Interpreting how GPT handles words split across token pairs.

An interpretability experiment on reducing documents.

A Peano arithmetic proof checker in Rust.

A small language built on top of the Z3 solver.

The Consciousness AI

Cooperative-agent research artifact requiring legacy Python 3.7 and TensorFlow 1 dependencies.

Research framework for progress alignment. Kept in the prototype tier because packaging remains incomplete.

also on Datasets & Benchmarks

Safe dexterous-manipulation simulation platform. Documentation remains in development.

also on RL Environments & Games

Applying crosscoder model diffing to emergently misaligned models

Prototype unsupervised probes for truthfulness.

Learn Safety & Interpretability

7

Courses, guides, exercises, reading lists, and visual explainers for technical safety and interpretability.

A curated reading list for mechanistic interpretability.

also on Research Maps & Awesome Lists

Exercises and notebooks for the ARENA alignment research curriculum.

Study materials on the mathematics used in AI safety arguments.

also on Safety Math & Formal Methods

An open-source, customizable intermediate logic textbook

also on Safety Math & Formal Methods

Survey and structured reading map covering AI alignment.

also on Research Maps & Awesome Lists

Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization

also on Mech Interp Tools

A library for mechanistic interpretability of GPT-style language models

also on Mech Interp Tools

Learn ML, RL & Systems

12

High-quality courses, books, labs, and from-scratch implementations for ML, RL, and systems foundations.

A high level deep learning library and its course.

LLM training in simple, raw C/CUDA

The best ChatGPT that $100 can buy.

karpathy/nanoGPTunder review

The simplest repository for training midsize GPTs.

also on Training & Model Systems

Dozens of papers implemented with side by side notes.

MIT's introductory deep learning course materials.

Efficient Deep Learning Systems course materials

A broad collection of machine learning study resources.

also on Research Maps & Awesome Lists

An educational resource to help anyone learn deep reinforcement learning.

A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.

srush/GPU-Puzzlesunder review

Learn CUDA by solving puzzles.

Learn PyTorch by solving puzzles.

Research Maps & Awesome Lists

8

Curated maps, bibliographies, roadmaps, and repository collections for discovery and study.

A curated reading list for mechanistic interpretability.

also on Learn Safety & Interpretability

A topic-centric list of HQ open datasets.

A collection of various awesome lists for hackers, pentesters and security researchers

also on Cybersecurity & Exploit Research

A broad collection of machine learning study resources.

also on Learn ML, RL & Systems

Survey and structured reading map covering AI alignment.

also on Learn Safety & Interpretability

A curated list of papers related to constrained decoding of LLM, along with their relevant code and resources.

The alignment literature as a scraped dataset.

also on Datasets & Benchmarks

πŸ“° Must-read papers and blogs on LLM based Long Context Modeling πŸ”₯