Coming later
CIS 7000 Graduate Topics Course
Mathematical Foundations of AI Alignment
- Instructor
- Aaron Roth
- Meetings
-
TR 12:00–1:29 PM
3401 Walnut, #401B
(Active Learning Classroom) - Dates
- August 25–December 7
- Communication
- Join course Slack
Course outline
Topics
-
01
From feedback to social objectives
- RLHF as social aggregation
- Heterogeneous preferences and cycles
- Arrow’s impossibility theorem
- Cardinal utility and Harsanyi aggregation
-
02
Preference optimization
- Scalar reward fitting and DPO
- Maximal lotteries and no-regret self-play
- Regularized Nash learning
- Welfare distortion and identifiability
-
03
Decision-aligned prediction
- Calibration and decision regret
- Omniprediction and indistinguishability
- Blackwell approachability
- Human–model agreement
-
04
Oversight and assistance
- AI debate and interactive verification
- Bayesian persuasion
- Deference and the off-switch game
- Cooperative inverse reinforcement learning
-
05
Institutions and competition
- Best-AI selection
- Emergent alignment through competition
- Equilibrium welfare guarantees
- Evaluation error and assumption audits
Possible extensions: Blackwell’s comparison of experiments; cheap talk without commitment.
Working list
Readings
Feedback and social objectives
- Deep Reinforcement Learning from Human Preferences
- Training Language Models to Follow Instructions with Human Feedback
- Three Brief Proofs of Arrow’s Impossibility Theorem
- Cardinal Welfare, Individualistic Ethics, and Interpersonal Comparisons of Utility
- Servant of Many Masters: Shifting Priorities in Pareto-Optimal Sequential Decision-Making
- Axioms for AI Alignment from Human Feedback
- Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Preference optimization
- Direct Preference Optimization
- A Minimaximalist Approach to Reinforcement Learning from Human Feedback
- Nash Learning from Human Feedback
- Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
- Distributional Preference Learning
- Distortion of AI Alignment
Prediction and agreement
- Companion tutorialCalibration, Decisions, and Collaboration in Learning
- The Well-Calibrated Bayesian
- Calibrated Language Models Must Hallucinate
- Evaluating Large Language Models for Accuracy Incentivizes Hallucinations
- Multicalibration: Calibration for the (Computationally-Identifiable) Masses
- Outcome Indistinguishability
- Calibrating Predictions to Decisions: A Novel Approach to Multi-Class Calibration
- Omnipredictors
- High-Dimensional Prediction for Sequential Decision Making
- Forecasting for Swap Regret for All Downstream Agents
- An Analog of the Minimax Theorem for Vector Payoffs
- Agreeing to Disagree
- Tractable Agreement Protocols
Oversight and assistance
- AI Safety via Debate
- Supervising Strong Learners by Amplifying Weak Experts
- AI Alignment via Incentives and Correction
- Bayesian Persuasion
- Friend or Foe: Delegating to an AI Whose Alignment is Unknown
- Robust Trust
- Scalable AI Safety via Doubly-Efficient Debate
- Cooperative Inverse Reinforcement Learning
- The Off-Switch Game
Institutions and competition
Course materials
Lectures
- Lecture 1 Overview PowerPoint Recording
- Lecture 2 Arrow’s Impossibility Theorem PDF
- Lecture 3 Expected Utility and Harsanyi Aggregation PDF