Cornell AI Alignment Club
Fall 2026

CS 1998: Intro to AI Safety & Alignment

1 Credit · 7 Weeks First · S/U Grading · Open Enrollment

Course StaffMeet the course team
Jinzhou Wu's profile photo

Jinzhou Wu

Organizer and Lead Instructor

Daniel Lee's profile photo

Daniel Lee

Head TA

Arya Datla's profile photo

Arya Datla

Teaching Assistant

Karan Verma's profile photo

Karan Verma

Teaching Assistant

Jasmine Li's profile photo

Jasmine Li

Advisor

Jonathn Chang's profile photo

Jonathn Chang

Advisor

Suvadip Sana's profile photo

Suvadip Sana

Advisor

Éva Tardos's profile photo

Éva Tardos

Faculty Advisor

Content

CS 1998: Intro to AI Safety & Alignment is a student-led course that explores why advanced AI systems can fail in unexpected and dangerous ways. We begin by building a solid understanding of how modern language models are trained, from pretraining on web-scale data through supervised fine-tuning and reinforcement learning from human feedback. From there, we turn to the core question: how do we ensure these systems do what we actually want? Students will learn key technical ideas in mechanistic interpretability (reverse-engineering model internals to understand what they've learned), reward learning (how optimization pressure can produce unintended behaviors like sycophancy and reward hacking), red teaming and adversarial evaluation (systematically probing models for failure modes), and scalable oversight (supervising systems that may exceed human-level performance on the tasks we're evaluating them on).

Logistics

  • This is a 1-credit, 7-week, first S/U course.
  • Location: Phillips Hall 203.
  • This course is open enrollment (without application), with around 75 seats.

Course Structure and Grading

The course includes Friday lectures, guided notebooks, Monday paper discussions, and a final project. We use point-based S/U grading: 145 points are available, and 100 points earns a pass.

Lecture. Each of seven Friday lectures earns 5 points, up to 35 points. Lectures introduce the week's core ideas, technical foundations, and research context.

Notebook. Weeks 2–6 each include a guided take-home notebook. Satisfactory completion earns 10 points, up to 50 points. Expect roughly 30 lines of student-written code within a provided framework, focused on a hands-on experiment and short analysis.

Discussion. Monday discussions examine a frontier or recent paper related to the preceding lecture. Each attendance earns 5 points, capped at 15 points.

Project proposal. Earn 5 points for defining an AI safety question, hypothesis, method, and expected result. You'll receive feedback before final-project work begins.

Final project. Earn 40 points for reproducing or extending a result, building an evaluation, comparing methods, or testing a safety hypothesis. Agentic coding tools are encouraged, but you remain responsible for understanding and validating the work. Compute credits will be provided.

Course Calendar

Syllabus

Week 1

Introduction to AI Safety

+

Topics

  • Why increasingly capable AI systems do not become safe or aligned by default.
  • Evidence from current systems: sycophancy, deception, specification gaming, and agentic failure modes.
  • Risk models spanning misuse, accidents, systemic harms, and loss of control.
  • How to reason about uncertain, low-probability, high-impact outcomes.
  • Capabilities forecasts, scenarios, and the assumptions behind them.

Week 2

AI Alignment and RLHF

+

Topics

  • The post-training pipeline: supervised fine-tuning, preference data, reward models, and policy optimization.
  • RLHF with PPO and direct preference methods such as DPO.
  • Constitutional AI, reinforcement learning from AI feedback, and rule-based alignment.
  • What present-day alignment methods accomplish in practice.
  • Known limitations: reward misspecification, overoptimization, plural values, and dependence on human oversight.

Week 3

Reward Hacking and Goal Misgeneralization

+

Topics

  • Goodhart's law, proxy objectives, and specification gaming.
  • Reward hacking and reward tampering in agents and language models.
  • Goal misgeneralization versus ordinary capability failures under distribution shift.
  • Sycophancy and emergent misalignment after narrow fine-tuning.
  • Mitigation strategies and why behavioral training may not remove the underlying failure mode.

Week 4

Interpretability

+

Topics

  • What interpretability is for and what kinds of evidence it can provide.
  • Probes, logit attribution, activation patching, and causal interventions.
  • Sparse features, circuit tracing, and representation engineering.
  • The Jacobian lens and the J-space hypothesis of a verbalizable global workspace.
  • Limits of current methods and the challenge of scalable, safety-relevant auditing.

Week 5

Evaluations—Evaluating Dangerous Capabilities

+

Topics

  • Threat modeling and capability elicitation before benchmark design.
  • Evaluations for cyber, biological, persuasion, deception, autonomy, and self-proliferation capabilities.
  • Construct validity, contamination, sandbagging, and evaluation awareness.
  • Manual red teaming, automated red teaming, and scalable evaluation pipelines.
  • Using evaluation evidence to inform deployment safeguards and safety cases.

Week 6

Control and Scalable Oversight

+

Topics

  • The supervision gap and weak-to-strong generalization.
  • Scalable oversight through debate, critique, decomposition, and recursive supervision.
  • AI control protocols: trusted monitoring, untrusted monitoring, trusted editing, and defer-to-human strategies.
  • Control evaluations against intentionally subversive model behavior.
  • Safety-usefulness tradeoffs and evidence for an AI control safety case.

Week 7

Policy, Governance, and Forecasting

+

Topics

  • Capability, algorithmic, and compute trends and what they can and cannot predict.
  • Forecasting timelines and takeoff under deep uncertainty.
  • Governance tools: evaluations, safety cases, transparency, incident reporting, and compute governance.
  • Domestic regulation and international coordination at the frontier.
  • Open research questions and technical, policy, and governance career paths.