CIS 7000 Graduate Topics Course

Mathematical Foundations of AI Alignment

Instructor
Aaron Roth
Meetings
TR 12:00–1:29 PM 3401 Walnut, #401B
(Active Learning Classroom)
Dates
August 25–December 7

Course outline

Topics

  1. 01

    From feedback to social objectives

    • RLHF as social aggregation
    • Heterogeneous preferences and cycles
    • Arrow’s impossibility theorem
    • Cardinal utility and Harsanyi aggregation
  2. 02

    Preference optimization

    • Scalar reward fitting and DPO
    • Maximal lotteries and no-regret self-play
    • Regularized Nash learning
    • Welfare distortion and identifiability
  3. 03

    Decision-aligned prediction

    • Calibration and decision regret
    • Omniprediction and indistinguishability
    • Blackwell approachability
    • Human–model agreement
  4. 04

    Oversight and assistance

    • AI debate and interactive verification
    • Bayesian persuasion
    • Deference and the off-switch game
    • Cooperative inverse reinforcement learning
  5. 05

    Institutions and competition

    • Best-AI selection
    • Emergent alignment through competition
    • Equilibrium welfare guarantees
    • Evaluation error and assumption audits

Possible extensions: Blackwell’s comparison of experiments; cheap talk without commitment.

Working list

Readings

01

Feedback and social objectives

02

Preference optimization

03

Prediction and agreement

04

Oversight and assistance

05

Institutions and competition

Coming later

Project ideas