Enroll today!
CS 1998, PRJ 608 · Class number 18589
1 Credit · 7 Weeks First · S/U Grading · Open Enrollment
CS 1998, PRJ 608 · Class number 18589

Organizer and Lead Instructor

Head TA

Teaching Assistant

Teaching Assistant

Advisor

Advisor

Advisor

Faculty Advisor
The course is led by CAIA members with faculty advising from Cornell.

Organizer and Lead Instructor

Head TA

Teaching Assistant

Teaching Assistant

Advisor

Advisor

Advisor

Faculty Advisor
CS 1998: Intro to AI Safety & Alignment is a student-led course that explores why advanced AI systems can fail in unexpected and dangerous ways. We begin by building a solid understanding of how modern language models are trained, from pretraining on web-scale data through supervised fine-tuning and reinforcement learning from human feedback. From there, we turn to the core question: how do we ensure these systems do what we actually want? Students will learn key technical ideas in mechanistic interpretability (reverse-engineering model internals to understand what they've learned), reward learning (how optimization pressure can produce unintended behaviors like sycophancy and reward hacking), red teaming and adversarial evaluation (systematically probing models for failure modes), and scalable oversight (supervising systems that may exceed human-level performance on the tasks we're evaluating them on).
The course includes Friday lectures, guided notebooks, Monday paper discussions, and a final project. We use point-based S/U grading: 145 points are available, and 100 points earns a pass.
Lecture. Each of seven Friday lectures earns 5 points, up to 35 points. Lectures introduce the week's core ideas, technical foundations, and research context.
Notebook. Weeks 2–6 each include a guided take-home notebook. Satisfactory completion earns 10 points, up to 50 points. Expect roughly 30 lines of student-written code within a provided framework, focused on a hands-on experiment and short analysis.
Discussion. Monday discussions examine a frontier or recent paper related to the preceding lecture. Each attendance earns 5 points, capped at 15 points.
Project proposal. Earn 5 points for defining an AI safety question, hypothesis, method, and expected result. You'll receive feedback before final-project work begins.
Final project. Earn 40 points for reproducing or extending a result, building an evaluation, comparing methods, or testing a safety hypothesis. Agentic coding tools are encouraged, but you remain responsible for understanding and validating the work. Compute credits will be provided.
Week 1
Introduction to AI Safety
Topics
Week 2
AI Alignment and RLHF
Topics
Slides
Discussion Readings
Week 3
Reward Hacking and Goal Misgeneralization
Topics
Slides
Slides (TBD)
Discussion Readings
Week 4
Interpretability
Topics
Slides
Slides (TBD)
Discussion Readings
Week 5
Evaluations—Evaluating Dangerous Capabilities
Topics
Slides
Slides (TBD)
Discussion Readings
Week 6
Control and Scalable Oversight
Topics
Slides
Slides (TBD)
Discussion Readings
Week 7
Policy, Governance, and Forecasting
Topics
Slides
Slides (TBD)
Discussion Readings