Information for Shane Legg

Basic information

Item	Value
Agendas	Recursive reward modeling

Organization	Title	Start date	End date	AI safety relation	Subject	Employment type	Source	Notes
Google DeepMind	Co-Founder and Chief Scientist	2010-09-23					[1], [2], [3], [4], [5]

Name	Creation date	Description

Title	Publication date	Author	Publisher	Affected organizations	Affected people	Document scope	Cause area	Notes

Title	Publication date	Author	Publisher	Affected organizations	Affected people	Affected agendas	Notes
Scalable agent alignment via reward modeling: a research direction	2018-11-19	Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg	arXiv	Google DeepMind		Recursive reward modeling, Imitation learning, inverse reinforcement learning, Cooperative inverse reinforcement learning, myopic reinforcement learning, iterated amplification, debate	This paper introduces the (recursive) reward modeling agenda, discussing its basic outline, challenges, and ways to overcome those challenges. The paper also discusses alternative agendas and their relation to reward modeling.

Showing at most 20 people who are most similar in terms of which organizations they have worked at.

Person	Number of organizations in common	List of organizations in common
Nick Bostrom	1	Google DeepMind
Chris Maddison	1	Google DeepMind
Laurent Orseau	1	Google DeepMind
Jan Leike	1	Google DeepMind
Demis Hassabis	1	Google DeepMind
Tom Everitt	1	Google DeepMind
Pedro A. Ortega	1	Google DeepMind
Vishal Maini	1	Google DeepMind
Thore Graepel	1	Google DeepMind
Stanislav Fort	1	Google DeepMind
Victoria Krakovna	1	Google DeepMind
Andrew Lefrancq	1	Google DeepMind
Azade Nova	1	Google DeepMind
Christiana Figueres	1	Google DeepMind
Diane Coyle	1	Google DeepMind
Edward W. Felten	1	Google DeepMind
Hanie Sedghi	1	Google DeepMind
Ian Goodfellow	1	Google DeepMind
Isabel Leal	1	Google DeepMind
James Manyika	1	Google DeepMind