SU-CS329H SEP232026

MLHF!!!

Optimizing system used by humans, from data on what they chose

New Concepts

Important Results / Claims

Rough Fields / Mapping

Book: mlhp.stanford.edu

Foundations

“where does the data come from”

  • Psycometrics: how to measure people
  • Discrete choice: how to do the math

Learning

“how do we model and fit”

  • Reinforcement learning: how to use what you’ve measured
  • Decision theory / inverse RL: POMDP stuff, since RL model need to actually map to actions

action

“what do we do with the fit”

Inversion

“do choices mean what we think”

  • Behavior economics / HCI: how to actually capture human behavior now that your model is turned
  • Social choice / normative theory: how to actually combine preferences of population for what exactly people want

Aggregation

“how do we combine across population”

Why human feedback at all?

Human Preferences as Training Signal

Christiano et al., 2017. Figure learn to backflip.

  • rollouts from trajectories
  • humans give feedback on whether the backflip was imminant
  • train reward model
  • then use reward model for RL algorithm

Human-Preferences is better at big models

Stiennon et al., 2020. Scaling human feedback outperforms just scaling the model bigger (e.g., 1.3B beats 6B).

Why not human feedback

“Targets are often ill-defined, optimized anyways.” We don’t know…

  • is a summary good?
  • is it “helpful”? what does it mean to be helpful to one human vs. another?
    • …operative example: “what is funny?”

context-dependent, subjective, potentially inconsistent.

different people have different preferences

Santurkar et al., 2023

Models shift with one default, someone’s desires loose. This is called a preference aggregation problem.

Voting, social choice, etc.

era of experience

(Silver and Sutton 2025)

  • Relying on human leads to a ceiling on agent performance
  • but grounded rewards may arise from humans just sitting around the agent’s environment instead of objective preferences (such as the humans being happy)

grounding

“signal is tied to reality”