Teaching
Rational Approaches to Cooperative Intelligence
What are the computational principles underlying human-like cooperative intelligence, and how can we use them to engineer cooperative and human-aligned machines? This reading-based seminar introduces a rational approach to answering these questions: one where both humans and AI are treated as approximately rational agents with coherent, probabilistic models of the social world, allowing them to act and cooperate on the basis of good reasons.
We begin with fundamental cooperative capacities like theory of mind and inverse planning, then explore how these enable forms of cooperation from assistance and teamwork to communication and teaching. We then study how many agents can cooperate even when they have different interests and goals—via norms, institutions, and negotiation—and the implications of all this for human-AI alignment.