Nothing on the shelves matches. Try fewer words, or suggest what should be here.
RL Environments & Games
19
Multi-agent environments, social dilemmas, bargaining, game theory, and RL simulations for Crucible experiments.
A researcher's collection of reinforcement learning environments.
The Diplomacy agent that negotiated with humans.
The standard API for reinforcement learning environments.
A standard API for multi-agent reinforcement learning environments, with popular reference environments and related utilities
Game Theory Explorer: Build, explore and solve extensive form games.
A suite of test scenarios for multi-agent reinforcement learning.
OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.
Imitation and reward learning algorithms in PyTorch.
also on Training & Model SystemsA benchmark environment for fully cooperative human-AI performance.
Multi-Agent LLM Bargaining and Negotiation
A toolbox with the goal of speeding up research on bargaining in MARL (cooperation problems in MARL).
Python Multi-Agent Reinforcement Learning framework
Safe dexterous-manipulation simulation platform. Documentation remains in development.
also on Research Prototypes & ReferencesUnified safe reinforcement learning environment and benchmark.
also on Safety Evals & Auditing, Datasets & BenchmarksVision-language-action evaluation framework with safety, generalization, and long-horizon task suites.
also on Safety Evals & Auditing, Datasets & BenchmarksHigh throughput reinforcement learning environments and training.
llms playing coup
Two player dialogue negotiation environments.
The Unity Machine Learning Agents Toolkit (ML-Agents) is an open-source project that enables games and simulations to serve as environments for training intelligent agents using deep reinforcement learning and imitation learning.
AI Red Teaming & Jailbreaks
26
Dual-use, defensive research on jailbreaks, prompt injection, red-team harnesses, and guardrail testing.
Framework agnostic adversarial examples.
A standardised evaluation framework for automated red teaming.
The foundational adversarial example library.
Red team LLMs and agents across forty plus vulnerabilities.
Automated fuzzing for language models.
Autonomous Red Team: Multi-Agent Adversarial Security Testing , paper + proof-of-concept framework | DAI-2513 | Dissensus AI Working Paper
autonomous red teaming platform; multi-agent offensive-security meta-harness
Evaluation and testing library for LLM agents.
also on Safety Evals & AuditingA repo for jailbreaking various LLMs, mainly Claude
Safety content moderation models built on Gemma 2. Gated access.
A fast implementation of the GCG adversarial attack.
Validate model outputs with composable validators.
A trivial programmatic Llama 3 jailbreak. Sorry Zuck!
An open benchmark for jailbreak attacks and defences.
Universal and Transferable Attacks on Aligned Language Models
An open safeguard model for prompt and response classification. Gated access.
Meta's open tools for language model security and safeguards.
Microsoft's risk identification framework for generative AI.
Agentic LLM vulnerability scanner.
also on Cybersecurity & Exploit ResearchProgrammable guardrails for LLM applications.
The LLM vulnerability scanner.
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Input and output scanners for LLM interactions. Archived upstream.
TextAttack π is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/
Web UI for viewing, editing, and AI-assisted red teaming of AI agent transcripts
IBM's adversarial attack and defence library.
Cybersecurity & Exploit Research
12
Dual-use AI security, exploit benchmarks, agent scanners, bug finding, and defensive cyber tooling.
Cybersecurity AI (CAI), the framework for AI Security
Professional CTF tasks for evaluating cyber capability.
also on Datasets & BenchmarksText watermarking from DeepMind.
A collection of various awesome lists for hackers, pentesters and security researchers
also on Research Maps & Awesome ListsCrash triage and severity analysis tools for fuzzing output.
Agentic LLM vulnerability scanner.
also on AI Red Teaming & JailbreaksScan model files for serialization attacks.
Supply chain signing for machine learning models.
Security scanner for AI agents, MCP servers and agent skills.
Security scanner for agentic workflows.
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
also on Datasets & BenchmarksPickle decompiler that detects malicious model files.
Alignment: Control & Monitoring
15
Control evaluations, trusted monitoring, sandbox escape, legibility, deception detection, and tamper resistance.
Research code for detecting deception in model behaviour.
Guardrails for secure and robust agent development
Black box lie detection through unrelated follow up questions.
Weak to strong generalization experiments. Archived upstream.
Evaluating whether models can plan subversion strategies.
Hidden reasoning in model outputs.
also on Datasets & BenchmarksMethods for making open weight safeguards resist tampering.
Sandbox escape benchmark for LLM capability evaluation
also on Datasets & BenchmarksThe Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations, sometimes toward bizarre or harmful personas. This repo contains a pipeline for generating the Assistant Axis and notebooks for monitoring...
also on Digital Minds & WelfareWhich models are illegible under what conditions, and why? How does that impact monitorability?
Research code for detecting deception in model outputs.
Benchmark dataset for evaluating trusted monitors on AI agent transcripts
also on Datasets & BenchmarksEvaluate AI agent transcripts for suspicious behavior (0-100 scoring)
Settings and baselines for AI control evaluations.
Alignment: Model Organisms
10
Alignment-faking model organisms, emergent misalignment, data poisoning, and falsifiable misbehavior testbeds.
Finding trojans in aligned models; the SaTML competition harness.
A team solution to the trojan detection competition.
Code released with the alignment faking paper.
Bench Alignment Faking: Alignment faking model organisms, detectors, and environments to catch misaligned models (Joshua Clymer MATS Stream Summer 2025)
Research code extending the alignment faking experiments.
Data poisoning experiments to support the study of emergent misalignment
Research code for studying models trained on false facts.
Research code on inoculation prompting.
Source code for the paper: Probing the Misaligned Thinking Process of Language Models
also on Interp Experiments & PapersOpen Source Replication of Anthropic's Alignment Faking Paper
Safety Evals & Auditing
31
Evaluation frameworks, auditing agents, safety benchmarks, model graders, and reproducible evaluation infrastructure.
A comprehensive trustworthiness assessment of GPT models.
also on Datasets & BenchmarksEvaluate code generation models.
A framework for few shot evaluation of language models; the most used open eval harness.
Evaluation and testing library for LLM agents.
also on AI Red Teaming & JailbreaksA unified library of evaluation metrics.
Petri, the alignment auditing agent that explores hypotheses in parallel.
METR Task Standard
METR's tool for running evaluations and agent elicitation research.
also on Research InfrastructureEvaluation registry and framework.
All-modality safety evaluation framework integrating more than 50 datasets.
also on Datasets & BenchmarksStatic multimodal deception benchmark with six deception categories. It is not presented as a monitoring tool.
also on Datasets & BenchmarksSafe reinforcement learning algorithm benchmark. It is not categorized as an environment.
Unified safe reinforcement learning environment and benchmark.
also on RL Environments & Games, Datasets & BenchmarksConstrained-learning code and benchmark assets for VLA safety alignment.
also on Training & Model SystemsBiosecurity benchmark with a public screening pipeline and reviewed access to the prompt set.
also on Datasets & BenchmarksVision-language-action evaluation framework with safety, generalization, and long-horizon task suites.
also on RL Environments & Games, Datasets & BenchmarksThe standardized adversarial robustness benchmark.
also on Datasets & BenchmarksResearch code for agents that audit model behaviour.
bloom - evaluate any behavior immediately πΈπ±
Code for "Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning"
Auditing agents for fine-tuning safety
Official Inspect Implementation for "ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases"
also on Datasets & BenchmarksWorked examples for the safety tooling stack.
Inference API for many LLMs and other useful tools for empirical research
also on Research InfrastructureBenchmark code from the SCONE evaluation work.
also on Datasets & BenchmarksHolistic evaluation of language models across scenarios and metrics.
Analyse and search agent transcripts at scale.
also on Research InfrastructureInspect: A framework for large language model evaluations
also on Research InfrastructureCommunity written evaluations for Inspect AI.
Mech Interp Tools
33
Core libraries and interfaces for activations, circuits, sparse features, interventions, and visualization.
See how predictions form, layer by layer.
Representation engineering: a top down approach to transparency.
also on Interp Experiments & PapersThe frontend behind the attribution graph publications.
Automated circuit discovery for transformer internals.
also on Interp Experiments & PapersFeature and prompt centric SAE visualisations.
Small utilities for tracing and editing activations inside networks.
Attribution graph tooling for tracing circuits in language models.
Train and analyse sparse autoencoders on language models.
Automated interpretability of learned features.
Sparse autoencoder training.
Sparsify transformers with SAEs and transcoders.
Instrumentation tooling used for mech interp on production code.
A JAX toolkit for building, editing and visualising networks.
Compile RASP programs into transformer weights. Archived upstream.
Sparse autoencoders trained on every layer of Gemma 2.
An open platform for hosting and exploring interpretability artefacts.
Dashboards for SAE features at scale.
The server that answers remote nnsight requests.
also on Research InfrastructureThe nnsight package enables interpreting and manipulating the internals of deep learned models.
Accompanying codebase for neuroscope.io, a website for displaying max activating dataset examples for language model neurons
Sparse autoencoders implemented in JAX.
Explain neurons with language models. Archived upstream.
Sparse autoencoders for GPT-4 scale activations.
Neuron and attention head investigation.
PKU accepted-release repository for multimodal sparse-autoencoder training and analysis.
PKU accepted-release repository extending TransformerLens to multimodal models.
Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization
also on Learn Safety & InterpretabilityTooling for working with sparse features.
Interventions on PyTorch models as composable configuration.
Attention and activation visualisation.
A library for mechanistic interpretability of GPT-style language models
also on Learn Safety & InterpretabilityMLP neuron level circuit tracing (ADAG).
A library for making control vectors.
Interp Experiments & Papers
19
Research code for probes, steering, model diffing, introspection, learning dynamics, and causal interpretability.
Representation engineering: a top down approach to transparency.
also on Mech Interp ToolsCompanion code for the global workspace interpretability paper
Attribution based parameter decomposition.
Automated circuit discovery for transformer internals.
also on Mech Interp ToolsFind knowledge neurons in pretrained transformers.
The hub for EleutherAI's work on interpretability and learning dynamics
also on Datasets & BenchmarksParameter decomposition research code.
Reference sparse autoencoder implementation.
Steering models with contrastive activation addition.
Code and models for empirical and theoretical work on alignment elasticity after fine-tuning.
Paper and research entry point for multimodal sparse-autoencoder interpretability.
Code for the "Mechanisms of Introspective Awareness" paper.
also on Digital Minds & WelfareSource code for the paper: Probing the Misaligned Thinking Process of Language Models
also on Alignment: Model OrganismsResearch code on emergent misalignment features in open models.
Training Transformers with knowledge localization (SGTM)
Research code on steering models through weight edits.
Tools for studying developmental interpretability in neural networks.
Invert embeddings back to the text that made them.
Safety Math & Formal Methods
13
Math for AI safety, formal verification, theorem proving, decision theory, game theory, and proof tooling.
Code for implicit learning dynamics in Stackelberg games, from ICML 2020.
Code for "Convergence of Learning Dynamics in Stackelberg Games"
A proof oriented programming language for verified software.
Formal verification methods for neural network robustness.
The mathematical library of Lean 4, built by a large open community.
The Lean 4 language and theorem prover.
Study materials on the mathematics used in AI safety arguments.
also on Learn Safety & InterpretabilityA Learning Environment for Theorem Proving
Language model generation checked by a verifier inside tree search.
An open-source, customizable intermediate logic textbook
also on Learn Safety & InterpretabilityRust bindings for the Z3 theorem prover.
The Rocq Prover, formerly Coq: an interactive proof assistant.
Python interface for the SCIP Optimization Suite
Research Infrastructure
12
Inference APIs, experiment scaffolds, cloud runners, provenance, labeling, tracking, and analysis tools.
Hierarchical configuration for research applications.
METR's tool for running evaluations and agent elicitation research.
also on Safety Evals & AuditingExperiment tracking, model registry and lifecycle tooling.
Tips for releasing research code in Machine Learning (with official NeurIPS 2020 recommendations)
Safe reinforcement learning algorithm infrastructure, not an environment provider.
Inference API for many LLMs and other useful tools for empirical research
also on Safety Evals & AuditingAnalyse and search agent transcripts at scale.
also on Safety Evals & AuditingData version control for machine learning pipelines.
Inspect: A framework for large language model evaluations
also on Safety Evals & AuditingExperiment tracking and visualisation for ML runs.
Agents & Research Automation
5
Research agents, bounded autonomous science, agent harnesses, MCP, and machine-learning engineering agents.
Property based testing driven by coding agents.
Specification and documentation for the Model Context Protocol
A self-improving RLM agent for coding workflows and long-running autonomous tasks.
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
AIDE: AI-Driven Exploration in the Space of Code. The machine Learning engineering agent that automates AI R&D.
Training & Model Systems
34
Pretraining, post-training, RLHF, distributed training, inference, quantization, tokenization, and model systems.
A fully open model: training, evaluation and inference code.
AllenAI's post training codebase.
Fast and memory-efficient exact attention
Distributed training and inference optimisation for large models.
A friendly federated learning framework.
Language model inference in C and C++ on commodity hardware.
A guidance language for controlling generation.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Parameter efficient fine tuning methods for transformers.
π₯ Fast State-of-the-Art Tokenizers optimized for Research and Production
Model implementations and training utilities.
Train transformer language models with reinforcement learning.
Preference pairs used to train open aligned models.
also on Datasets & BenchmarksImitation and reward learning algorithms in PyTorch.
also on RL Environments & GamesComposable function transformations on accelerators.
Inference, preference collection and tuning for local models.
The simplest repository for training midsize GPTs.
also on Learn ML, RL & SystemsLegible, reproducible JAX training with named tensors.
Differentially private training for PyTorch.
NVIDIA's reference stack for large scale transformer training.
An open RLHF training framework built on Ray and vLLM.
Neural networks in JAX as pytrees.
Multimodal alignment training framework supporting SFT, DPO, PPO, and related methods.
Model-agnostic correction module for alignment training and deployment experiments.
Safe RLHF training pipeline with reward and cost modeling.
Safe reinforcement learning algorithm using world models. It is not categorized as an environment.
Constrained-learning code and benchmark assets for VLA safety alignment.
also on Safety Evals & AuditingGlobally distributed training over the internet.
The framework underneath most of the shelves.
PyTorch native pretraining at scale.
SGLang is a high-performance serving framework for large language models and multimodal models.
The Mamba state space architecture.
Extreme compression via additive quantization.
High throughput language model serving with paged attention.
Digital Minds & Welfare
4
Machine consciousness, introspection, valence, welfare, personas, and internal-state research.
The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's behavior is. Models can drift away from the Assistant during conversations, sometimes toward bizarre or harmful personas. This repo contains a pipeline for generating the Assistant Axis and notebooks for monitoring...
also on Alignment: Control & MonitoringTraining LLMs to Report Their Learned Behaviors
Code for the "Mechanisms of Introspective Awareness" paper.
also on Interp Experiments & PapersPersona Vectors: Monitoring and Controlling Character Traits in Language Models
Datasets & Benchmarks
36
Safety, capability, agent, interpretability, and ML datasets and benchmark suites.
A benchmark for measuring the harmfulness of LLM agents.
A comprehensive trustworthiness assessment of GPT models.
also on Safety Evals & AuditingModeration training and test data across risk categories.
Professional CTF tasks for evaluating cyber capability.
also on Cybersecurity & Exploit ResearchPrompts for measuring discrimination in model decisions.
Human preference data on helpfulness and harmlessness.
Evaluations written by models, covering many behaviours.
Values expressed by a deployed assistant, measured at scale.
The WMDP hazardous knowledge benchmark, as a dataset.
A proxy benchmark for hazardous knowledge, with unlearning baselines.
The hub for EleutherAI's work on interpretability and learning dynamics
also on Interp Experiments & PapersThe Abstraction and Reasoning Corpus.
Preference pairs used to train open aligned models.
also on Training & Model SystemsA prompt injection game and the dataset it produced.
Toxicity annotations on real user prompts.
Procedural generators that reverse engineer ARC.
Human-preference datasets for safety alignment research.
All-modality safety evaluation framework integrating more than 50 datasets.
also on Safety Evals & AuditingStatic multimodal deception benchmark with six deception categories. It is not presented as a monitoring tool.
also on Safety Evals & AuditingPreference data with separate helpfulness and harmlessness labels.
Research framework for progress alignment. Kept in the prototype tier because packaging remains incomplete.
also on Research Prototypes & ReferencesHuman-preference dataset and code for text-to-video safety alignment.
Unified safe reinforcement learning environment and benchmark.
also on RL Environments & Games, Safety Evals & AuditingBiosecurity benchmark with a public screening pipeline and reviewed access to the prompt set.
also on Safety Evals & AuditingVision-language-action evaluation framework with safety, generalization, and long-horizon task suites.
also on RL Environments & Games, Safety Evals & AuditingHidden reasoning in model outputs.
also on Alignment: Control & MonitoringThe standardized adversarial robustness benchmark.
also on Safety Evals & AuditingSandbox escape benchmark for LLM capability evaluation
also on Alignment: Control & MonitoringOfficial Inspect Implementation for "ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases"
also on Safety Evals & AuditingBenchmark code from the SCONE evaluation work.
also on Safety Evals & AuditingBenchmark dataset for evaluating trusted monitors on AI agent transcripts
also on Alignment: Control & MonitoringThe alignment literature as a scraped dataset.
also on Research Maps & Awesome ListsExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
also on Cybersecurity & Exploit ResearchCan language models resolve real world GitHub issues?
Measuring how models imitate human falsehoods.
Research Prototypes & References
12
Undocumented, early-stage, or legacy research artifacts kept for discovery, not presented as production-ready tools.
A technical AI governance project on blind auditing.
An interpretability experiment notebook.
Interpreting how GPT handles words split across token pairs.
An interpretability experiment on reducing documents.
A Peano arithmetic proof checker in Rust.
A small language built on top of the Z3 solver.
The Consciousness AI
Cooperative-agent research artifact requiring legacy Python 3.7 and TensorFlow 1 dependencies.
Research framework for progress alignment. Kept in the prototype tier because packaging remains incomplete.
also on Datasets & BenchmarksSafe dexterous-manipulation simulation platform. Documentation remains in development.
also on RL Environments & GamesApplying crosscoder model diffing to emergently misaligned models
Prototype unsupervised probes for truthfulness.
Learn Safety & Interpretability
7
Courses, guides, exercises, reading lists, and visual explainers for technical safety and interpretability.
A curated reading list for mechanistic interpretability.
also on Research Maps & Awesome ListsExercises and notebooks for the ARENA alignment research curriculum.
Study materials on the mathematics used in AI safety arguments.
also on Safety Math & Formal MethodsAn open-source, customizable intermediate logic textbook
also on Safety Math & Formal MethodsSurvey and structured reading map covering AI alignment.
also on Research Maps & Awesome ListsTransformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization
also on Mech Interp ToolsA library for mechanistic interpretability of GPT-style language models
also on Mech Interp Tools
Learn ML, RL & Systems
12
High-quality courses, books, labs, and from-scratch implementations for ML, RL, and systems foundations.
A high level deep learning library and its course.
LLM training in simple, raw C/CUDA
The best ChatGPT that $100 can buy.
The simplest repository for training midsize GPTs.
also on Training & Model SystemsDozens of papers implemented with side by side notes.
MIT's introductory deep learning course materials.
Efficient Deep Learning Systems course materials
A broad collection of machine learning study resources.
also on Research Maps & Awesome ListsAn educational resource to help anyone learn deep reinforcement learning.
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
Learn CUDA by solving puzzles.
Learn PyTorch by solving puzzles.
Research Maps & Awesome Lists
8
Curated maps, bibliographies, roadmaps, and repository collections for discovery and study.
A curated reading list for mechanistic interpretability.
also on Learn Safety & InterpretabilityA topic-centric list of HQ open datasets.
A collection of various awesome lists for hackers, pentesters and security researchers
also on Cybersecurity & Exploit ResearchA broad collection of machine learning study resources.
also on Learn ML, RL & SystemsSurvey and structured reading map covering AI alignment.
also on Learn Safety & InterpretabilityA curated list of papers related to constrained decoding of LLM, along with their relevant code and resources.
The alignment literature as a scraped dataset.
also on Datasets & Benchmarksπ° Must-read papers and blogs on LLM based Long Context Modeling π₯
RL Environments & Games19
- cfpark00/rl-environments
- facebookresearch/diplomacy_ciceroreview
- Farama-Foundation/Gymnasium
- Farama-Foundation/PettingZoo
- gambitproject/gte
- google-deepmind/meltingpot
- google-deepmind/open_spiel
- HumanCompatibleAI/imitationreview
- HumanCompatibleAI/overcooked_ai
- joie-zhang/bargain
- longtermrisk/marltoolbox
- oxwhirl/pymarl
- PKU-Alignment/ReDMan
- PKU-Alignment/safety-gymnasium
- PKU-Alignment/VLA-Arena
- PufferAI/PufferLibreview
- safety-research/social_games
- stanfordnlp/cocoareview
- Unity-Technologies/ml-agents
AI Red Teaming & Jailbreaks26
- bethgelab/foolboxreview
- centerforaisafety/HarmBench
- cleverhans-lab/cleverhansreview
- confident-ai/deepteamreview
- cyberark/FuzzyAIreview
- dissensus-ai/autonomous-red-team
- elder-plinius/T3MP3ST
- Giskard-AI/giskard-ossreview
- Goochbeater/Spiritual-Spell-Red-Teaming
- google/shieldgemma-2bhf
- GraySwanAI/nanoGCGreview
- guardrails-ai/guardrailsreview
- haizelabs/llama3-jailbreak
- JailbreakBench/jailbreakbench
- llm-attacks/llm-attacks
- meta-llama/Llama-Guard-3-8Bhf
- meta-llama/PurpleLlama
- microsoft/PyRITreview
- msoedov/agentic_securityreview
- NVIDIA-NeMo/Guardrailsreview
- NVIDIA/garakreview
- promptfoo/promptfoo
- protectai/llm-guardreview
- QData/TextAttack
- safety-research/agent-transcript-editor
- Trusted-AI/adversarial-robustness-toolboxreview
Cybersecurity & Exploit Research12
Alignment: Control & Monitoring15
- ApolloResearch/deception-detectionreview
- EleutherAI/elk
- invariantlabs-ai/invariant
- LoryPack/LLM-LieDetectorreview
- openai/weak-to-strongreview
- redwoodresearch/subversion-strategy-evalreview
- redwoodresearch/Text-Steganography-Benchmark
- rishub-tamirisa/tamper-resistance
- safety-research/agent-escape-bench
- safety-research/assistant-axis
- safety-research/legibility
- safety-research/lie-detector
- safety-research/sleight-bench
- safety-research/trusted-monitor
- UKGovernmentBEIS/control-arena
Alignment: Model Organisms10
- ethz-spylab/rlhf_trojan_competitionreview
- neverix/rlhf-trojan-2024-codreview
- redwoodresearch/alignment_faking_publicreview
- redwoodresearch/bench-af-2
- safety-research/alignment-faking-extensions
- safety-research/data-poisoning-public
- safety-research/false-facts
- safety-research/inoculation-prompting
- safety-research/misalignment-indicators
- safety-research/open-source-alignment-faking
Safety Evals & Auditing31
- AI-secure/DecodingTrustreview
- anthropics/evals
- bigcode-project/bigcode-evaluation-harnessreview
- EleutherAI/lm-evaluation-harnessreview
- Giskard-AI/giskard-ossreview
- huggingface/evaluatereview
- meridianlabs-ai/inspect_petri
- METR/hawk
- METR/task-standard
- METR/vivariareview
- openai/evals
- PKU-Alignment/eval-anything
- PKU-Alignment/MM-DeceptionBench
- PKU-Alignment/Safe-Policy-Optimization
- PKU-Alignment/safety-gymnasium
- PKU-Alignment/SafeVLA
- PKU-Alignment/SPIKE-Bench
- PKU-Alignment/VLA-Arena
- RobustBench/robustbenchreview
- safety-research/auditing-agents
- safety-research/bloom
- safety-research/faithful-cot
- safety-research/finetuning-auditor
- safety-research/impossiblebench
- safety-research/safety-examples
- safety-research/safety-tooling
- safety-research/SCONE-bench
- stanford-crfm/helmreview
- TransluceAI/docentreview
- UKGovernmentBEIS/inspect_ai
- UKGovernmentBEIS/inspect_evalsreview
Mech Interp Tools33
- AlignmentResearch/tuned-lensreview
- andyzoujm/representation-engineeringreview
- anthropics/attribution-graphs-frontendreview
- ArthurConmy/Automatic-Circuit-Discoveryreview
- callummcdougall/sae_visreview
- davidbau/baukitreview
- decoderesearch/circuit-tracer
- decoderesearch/SAELens
- EleutherAI/delphi
- EleutherAI/sae
- EleutherAI/sparsifyreview
- google-deepmind/mishaxreview
- google-deepmind/penzaireview
- google-deepmind/tracrreview
- google/gemma-scopehf
- hijohnnylin/neuronpedia
- jbloomAus/SAEDashboardreview
- ndif-team/ndifreview
- ndif-team/nnsight
- neelnanda-io/Neuroscope
- neverix/saexreview
- openai/automated-interpretabilityreview
- openai/sparse_autoencoder
- openai/transformer-debugger
- PKU-Alignment/SAELens-V
- PKU-Alignment/TransformerLens-V
- poloclub/transformer-explainer
- safety-research/sparse-feature-toolkit
- stanfordnlp/pyvenereview
- TransformerLensOrg/CircuitsVis
- TransformerLensOrg/TransformerLens
- TransluceAI/circuitsreview
- vgel/repengreview
Interp Experiments & Papers19
- andyzoujm/representation-engineeringreview
- anthropics/jacobian-lens
- ApolloResearch/apdreview
- ArthurConmy/Automatic-Circuit-Discoveryreview
- EleutherAI/elk
- EleutherAI/knowledge-neuronsreview
- EleutherAI/pythia
- goodfire-ai/param-decompreview
- neelnanda-io/1L-Sparse-Autoencoder
- nrimsky/CAAreview
- PKU-Alignment/llms-resist-alignment
- PKU-Alignment/SAE-V
- safety-research/introspection-mechanisms
- safety-research/misalignment-indicators
- safety-research/open-source-em-features
- safety-research/selective-gradient-masking
- safety-research/weight-steering
- timaeus-research/devinterp
- vec2text/vec2textreview
Safety Math & Formal Methods13
- fiezt/ICML-2020-Implicit-Stackelberg-Learning
- fiezt/Stackelberg-Code
- FStarLang/FStar
- google-deepmind/deep-verify
- leanprover-community/mathlib4review
- leanprover/lean4review
- lionellevine/MAIS
- ml4tp/gamepad
- namin/llm-verified-with-monte-carlo-tree-searchreview
- OpenLogicProject/OpenLogic
- prove-rs/z3.rs
- rocq-prover/rocqreview
- scipopt/PySCIPOpt
Research Infrastructure12
Agents & Research Automation5
Training & Model Systems34
- allenai/OLMoreview
- allenai/open-instructreview
- Dao-AILab/flash-attention
- deepspeedai/DeepSpeed
- flwrlabs/flowerreview
- ggml-org/llama.cpp
- guidance-ai/guidancereview
- hiyouga/LlamaFactory
- huggingface/peft
- huggingface/tokenizers
- huggingface/transformers
- huggingface/trl
- HuggingFaceH4/ultrafeedback_binarizedhf
- HumanCompatibleAI/imitationreview
- jax-ml/jaxreview
- JD-P/minihfreview
- karpathy/nanoGPTreview
- marin-community/levanterreview
- meta-pytorch/opacusreview
- NVIDIA/Megatron-LM
- OpenRLHF/OpenRLHF
- patrick-kidger/equinoxreview
- PKU-Alignment/align-anything
- PKU-Alignment/aligner
- PKU-Alignment/safe-rlhf
- PKU-Alignment/SafeDreamer
- PKU-Alignment/SafeVLA
- PrimeIntellect-ai/prime-dilocoreview
- pytorch/pytorch
- pytorch/torchtitan
- sgl-project/sglang
- state-spaces/mambareview
- Vahe1994/AQLMreview
- vllm-project/vllm
Digital Minds & Welfare4
Datasets & Benchmarks36
- ai-safety-institute/AgentHarmhf
- AI-secure/DecodingTrustreview
- allenai/wildguardmixhf
- andyzorigin/cybenchreview
- Anthropic/discrim-evalhf
- Anthropic/hh-rlhfhf
- Anthropic/model-written-evalshf
- Anthropic/values-in-the-wildhf
- anthropics/evals
- cais/wmdphf
- centerforaisafety/wmdpreview
- EleutherAI/pythia
- fchollet/ARC-AGIreview
- HuggingFaceH4/ultrafeedback_binarizedhf
- HumanCompatibleAI/tensor-trustreview
- lmsys/toxic-chathf
- michaelhodel/re-arcreview
- PKU-Alignment/beavertails
- PKU-Alignment/eval-anything
- PKU-Alignment/MM-DeceptionBench
- PKU-Alignment/PKU-SafeRLHFhf
- PKU-Alignment/ProgressGym
- PKU-Alignment/safe-sora
- PKU-Alignment/safety-gymnasium
- PKU-Alignment/SPIKE-Bench
- PKU-Alignment/VLA-Arena
- redwoodresearch/Text-Steganography-Benchmark
- RobustBench/robustbenchreview
- safety-research/agent-escape-bench
- safety-research/impossiblebench
- safety-research/SCONE-bench
- safety-research/sleight-bench
- StampyAI/alignment-research-datasetreview
- sunblaze-ucb/exploitgym
- SWE-bench/SWE-benchreview
- sylinrl/TruthfulQAreview
Research Prototypes & References12
Learn Safety & Interpretability7
Learn ML, RL & Systems12
Research Maps & Awesome Lists8
- AI-in-Transportation-Lab/awesome-mechanistic-interpretability
- awesomedata/awesome-public-datasets
- Hack-with-Github/Awesome-Hacking
- nivu/ai_all_resources
- PKU-Alignment/AlignmentSurvey
- Saibo-creator/Awesome-LLM-Constrained-Decoding
- StampyAI/alignment-research-datasetreview
- Xnhyacinth/Awesome-LLM-Long-Context-Modeling










































































































































