The course covers foundational model training pipelines, mechanistic interpretability, RLHF and goal misgeneralization, safety evaluations and red teaming, scalable oversight and control, and policy and career pathways in AI safety.
The format emphasizes hands-on notebooks, paper-driven discussion, and a final project to help students build both conceptual understanding and practical skills.