_index.org

non-dictatorship

Last edited: September 9, 2026

no single person can decide aggregation

Stanford UG Courses Index

Last edited: September 9, 2026

Stanford UG Y1, Aut

Stanford UG Y1, Win

Stanford UG Y1, Spr

Stanford UG Y2, Aut

Stanford UG Y2, Win

Stanford UG Y2, Spr

Stanford UG Y3, Aut

Stanford UG Y3, Win

Stanford UG Y3, Spr

Stanford GR Y1, Aut

Stanford Talks

DateTopicPresenterLink
<2023-09-20 Wed>UG Research ProgramBrian ThomasStanford UG Research Program
<2023-09-28 Thu>Bld an Ecosystem, Not MonolithColin RaffelBuild a System
<2023-10-05 Thu>Training Helpful CHatbotsNazeen RajaniTraining Helpful Chatbots
<2023-10-26 Thu>AI Intepretability for BioGasper BegusAI Intepretability
<2023-11-02 Thu>PT Transformers on Long SeqsMike LewisPretraining Long Transformers
<2023-11-07 Tue>Transformers!A. VaswaniTransformers
<2023-11-09 Thu>Towards Interactive AgentsJessy LinInteractive Agent
<2023-11-16 Thu>Dissociating Language and ThoughtAnna IvanovaDissociating Language and Thought
<2024-01-11 Thu>Language AgentsKarthik NarasimhanLanguage Agents with Karthik
<2024-02-01 Thu>Pretraining Data
<2024-02-08 Thu>value alignmentBeen KimLM Alignment
<2024-02-15 Thu>model editingPeter HaseKnowledge Editing
<2024-07-18 Thu>Knowledge Localization
<2024-11-11 Mon>PresentationsSydney KatzPresentations
<2025-01-06 Mon>Video Generation with Learned PriorMeenakshi SarkarPriors
<2025-01-06 Mon>Theoretical Drone ControlSliding Mode UAV Control
<2025-01-09 Thu>VLM to AgentsTao YuVLM to Agents
<2025-01-13 Mon>Social RLNatasha JaquesSocial Reinforcement Learning
<2025-02-10 Mon>Model Predictive Control + PromptingGabriel MaherLLM MPC
<2025-03-03 Mon>Planning for Learning
<2025-03-06 Thu>Theorem ProvingSelf-Play Conjection Generalization
<2025-04-10 Thu>Safety for TrucksSafety for Autonomous Trucking
<2025-08-04 Mon>Collaborate Multiagent DMCollaborative Multiagent DM
<2025-09-22 Mon>AI Safety TalksAI Safety Annual Meeting
<2025-10-02 Thu>Pretraining under infinite computeLimited Samples and Infinite Compute
<2025-10-06 Mon>Mel KrusniakDecisions.jl
<2025-10-11 Sat>SISL Flash TalksSISL Talks
<2025-10-16 Thu>Predicting Scaling Performance
<2025-12-08 Mon>mixed-autonomy traffic with LLMSmixed-autonomy traffic with LLMs
<2026-01-05 Mon>AI Incidents PolicyAI Incidents Policy
<2026-01-12 Mon>Reliable RLReliable RL
<2026-01-15 Thu>Words to ConceptsWords to Concepts
Zen’s Defense
<2026-03-30 Mon>multi-agent LLMMulti-Agent LLMs
<2026-04-27 Mon>Alex’s Defense
<2026-09-28 Mon>Yi Chengen ai safety model and datasets
<2026-09-28 Mon>Anastasia KoslovaFrom Security to Trustworthy AI

Contacts

Talk Contacts

SU-CS224N MAY022024

Last edited: September 9, 2026

Zero-Shot Learning

GPT-2 is able to do many tasks with not examples + no gradient updates.

Instruction Fine-Tuning

Language models, by default, are not aligned with user intent.

  1. collect paired examples of instruction + output across many tasks
  2. then, evaluate on unseen tasks

~3 million examples << n billion examples

dataset: MMLU

You can generate an Instruction Fine-Tuning dataset by asking a larger model for it (see Alpaca).

Pros + Cons

  • simple and straightforward + generalize to unseen tasks
  • but, its EXPENSIVE to collect ground truth data
    • ground truths maybe wrong
    • creative tasks may not have a correct answer
    • LMs penalizes all token-level mistakes equally, but some mistakes are worse than others
    • humans may generate suboptimal answers

Human Preference Modeling

Imagine if we have some input \(x\), and two output trajectories, \(y_{1}\) and \(y_{2}\).

SU-CS329H SEP232026

Last edited: September 9, 2026

MLHF!!!

Optimizing system used by humans, from data on what they chose

New Concepts

Important Results / Claims

Rough Fields / Mapping

Book: mlhp.stanford.edu

Foundations

“where does the data come from”

  • Psycometrics: how to measure people
  • Discrete choice: how to do the math

Learning

“how do we model and fit”

  • Reinforcement learning: how to use what you’ve measured
  • Decision theory / inverse RL: POMDP stuff, since RL model need to actually map to actions

action

“what do we do with the fit”

SU-CS329Z SEP282026

Last edited: September 9, 2026

> nit

Slide numbers :)

> How having a probability related to generation? / Why is LM discriminative.

Re the student question about why a language model isn’t descriminative. It’s sorta true that LMs are a “classifier”, as are many descriminative models, but notice that there’s no “correct” classification.

Instead, you are normatively “supposed” to sample the distribution instead of simply outputting the “correct” next token because no such token exist. In this sense, its generative.