Work on this with us
The institute is built for collaboration: shared instruments, explicit protocols, and research objects designed to be replicated, criticised, and extended.
01 Areas
Collaboration is welcome across the institute's four programmes: mechanistic interpretability (probes, SAEs, circuit analysis), model cognition (situational and evaluation awareness, chain of thought faithfulness, introspection), alignment auditing and activation level experimentation, research interfaces (the instruments this site is built around), model entrenchment (case coding, institutional analysis, compute economics), and research engineering (TransformerLens/SAELens pipelines, in browser model tooling).
02 Contact
[email protected]. Include what you want to work on and what you've built or written.