AI Watch

Welcome! This is a website to track people and organizations in the AI safety/alignment/AI existential risk communities. A position or organization being on AI Watch does not indicate an assessment that that position or organization is actually making AI safer or that the position or organization is good for the world in any way. It is mostly a sociological indication that the position or organization is associated with these communities, as well as an indication that the position or organization claims to be working on AI safety or alignment. (There are some plans to eventually introduce such assessments on AI Watch, but for now there are none.) See the code repository for the source code and data of this website.

This website is developed by Issa Rice with data contributions from Sebastian Sanchez, Amana Rice, and Vipul Naik, and has been partially funded by Vipul Naik and Mati Roy (who in July 2023 paid for the time Issa had spent answering people’s questions about AI Watch up until that point).

Last updated on 2026-08-16; see here for a full list of recent changes.

Table of contents

Agendas

Agenda name Associated people Associated organizations
Iterated amplification Paul Christiano, Buck Shlegeris, Dario Amodei OpenAI
Embedded agency Eliezer Yudkowsky, Scott Garrabrant, Abram Demski Machine Intelligence Research Institute
Comprehensive AI services Eric Drexler Future of Humanity Institute
Ambitious value learning Stuart Armstrong Future of Humanity Institute
Factored cognition Andreas Stuhlmüller Ought
Recursive reward modeling Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg Google DeepMind
Debate Paul Christiano OpenAI
Interpretability Christopher Olah
Inverse reinforcement learning
Preference learning
Cooperative inverse reinforcement learning
Imitation learning
Alignment for advanced machine learning systems Jessica Taylor, Eliezer Yudkowsky, Patrick LaVictoire, Andrew Critch Machine Intelligence Research Institute
Learning-theoretic AI alignment Vanessa Kosoy
Counterfactual reasoning Jacob Steinhardt

Positions grouped by person

Showing 0 people with positions.

Name Number of organizations List of organizations

Positions grouped by organization

Showing 27 organizations.

Organization Number of people List of people
AI Security Institute 72 Orazio Angelini, Varad Vishwarupe, Advik Raj, Paul Röttger, Seoirse Murray, Konstantin Sietzy, Jessica McFadyen, Kai Fronsdal, Ivo Andrews, Cameron Holmes, Caroline Wagner, Jessica Wang, Jorge Perez-Alvarez-Pallete, Konstantinos V., Ben Dixon, Cate Heine, James Aung, Keno Jüchems, Adam R. M. Moore, Ekin Zorer, Sarenne Wallbridge, Satvik Golechha, Aleksandr Bowkis, David Demitri Africa, Cecilia Wood, Vanessa Cheung, Jan Michelfeit, Thomas Read, Matthew Clarke, Elliot Jones, Giorgi Giglemiani, Toby D. Pilditch, Steph Suddell, Andrew Strait, Jordan Taylor, Magda Dubois, Max Heitmann, Martín Soto, Charlie Griffin, Kimberly Mai, Benjamin A. R. Hilton, Merlin Stein, Rogan Inglis, Vy Hong, Felix Jackson, Joseph Bloom, Marie Buhl, Hadrien Pouget, Jonathan Richard Schwarz, Lennart Luettgau, Kola Ayonrinde, Asa Cooper Stickland, Alex Remedios, Anna Gausen, Hannah Rose Kirk, Philippos Maximos Giavridis, Robert Kirk, Kobi Hackenburg, Sophie Rose, Mahmoud Ghanem, Timo Flesch, Ishan Mishra, Alan Cooney, Kwan Yee Ng, Joe Skinner, Jake Pencharz, Lexi Keegan, Alexandra Souly, Harry Coppock, Henry Ogden, Xander Davies, John Wilkinson
Fund for Alignment Research 32 Tigist Diriba, Matt Pallissard, Jasper Timm, Thomas Costello, Stefan Heimersheim, Sam Adam-Day, Matthew Kowal, Lukas Struppek, Levon Avagyan, Lars Yencken, Jean-François Godbout, Gordon Pennycook, David Rand, Antonio Arechar, Oskar Hollinsworth, Tony Wang, Chris MacLeod, Chris Cundy, Ann-Kathrin Dombrowski, Aaron Tucker, Niki Howe, Adrià Garriga-Alonso, Ethan Perez, Claudia Shi, Kellin Pelrine, Mohammad Taufeeque, Tom Tseng, Nora Belrose, Tomasz Korbak, Jun Shern Chan, Jérémy Scheurer, Ian McKenzie
Leverhulme Centre for the Future of Intelligence 28 Raphael Hernandes, Harriet C., Minja Axelsson, Leah Madelaine Schmidt, Christoffer Koch Andersen, Connor Wright, Suren Pahlevan, Helen Leung, Konstantinos V., Aanya Niaz, Sammy McKinney, Julien Porquet, Benjamin Henke, Seraphina Zhang, Muhammed Alakitan, Wout Schellaert, Emily Elstub, Flavia Saxler, Beryl Pong, Anna Odynets, Xiang Li, Matthijs M. Maas, Tomasz Hollanek, Cassie Robinson, Daniel White, Adrian Weller, Carla Zoe Cremer, José Hernández-Orallo
Machine Intelligence Research Institute 8 Alex Zhu, Alex Mennen, Alex Appel, Linda Linsefors, Evan Hubinger, David Simmons, Daniel Demski, Patrick LaVictoire
Convergence Analysis 6 Anna Schuh, Elliot McKernon, Christopher DiCarlo, Michael Aird, Siebe Rozendal, Eric Easley
OpenAI 4 Christopher Olah, Geoffrey Irving, Paul Christiano, Dario Amodei
Center for Human-Compatible AI 3 Christopher Cundy, Beth Barnes, Dmitrii Krasheninnikov
Heron AI Security 3 Zhuang Ye, Christine Lai, Nitzan Shulman
AIDEUS 2 Sergey Rodionov, Alexey Potapov
Association for Long Term Existence and Resilience 2 Vanessa Kosoy, Ram Rachum
Athena 2 Yulia Volkova, Isabel MacGinnitie
Google DeepMind 2 Pedro A. Ortega, Chris Maddison
Learning Intelligent Distribution Agent 2 Tamas Madl, Stan Franklin
University of Oxford 2 Ruth Fong, Chris Maddison
Australian National University 1 Jarryd Martin
Carnegie Mellon University 1 Noam Brown
Center on Long-Term Risk 1 Caspar Oesterheld
EleutherAI 1 Jonas Müller
ETH Zurich 1 Felix Berkenkamp
EthicsNet 1 Adam Alonzi
Future of Humanity Institute 1 Sören Mindermann
Heron AI Security X Apart Research 1 Neta Ravid
Massachusetts Institute of Technology 1 Jon Gauthier
Oregon State University 1 Thomas Dietterich
Stanford University 1 Aditi Raghunathan
University of California, Berkeley 1 Michael Janner
University of Toronto 1 Roger Grosse