_index.org

random utility

Last edited: September 9, 2026

Let’s create \(N\) menus \(U_{1} … U_{N}\). And let’s write a decision rule \(U_{j} = u_{j} + \varepsilon_{j}\) (“each person picks some best available option \(u_{j}\) their menu \(U_{j}\), up to noise”)

The logit model asks that \(\varepsilon\) comes from a Gumbel distribution.

An alternative model probit model where \(\varepsilon \sim \mathcal{N}\qty(0, \Sigma)\)

Scale vs Population

Last edited: September 9, 2026

You can think of prference learning as two things:

“Scale People” - pairs

“We are observing the same population, but with different types of noise.”

simple scalability, substitutability, independence. In order of increasingly specific: Transitivity and Simple Scalability => Fechner and Quadrupel Condition => Bradley-Terry Preference Model

“Population People” - menus

“We are observing a heterogeneous population, and we are trying to characterize how they are heterogeneous.”

random utility, elimination by aspects, recommender systems => A softmax predictor follows Luce’s Axiom / independence to irrelavent alternatives

SU-CS329H SEP282026

Last edited: September 9, 2026

Structures

Intro

Rich Sutton and David Silver seems to think that rollout of environments is all we need, and that we are “done” with human generated supervizing data, so… What’s special about human data?

How to get Inductive Bias into your system?

So we learned from what’s special about human data? that if you want to make inference

Content

Important Results / Claims

positively vs. normatively

  • positively: “what is true” - facts, cause and effect, measurable reality
  • normatively: “what ought to be” - what is good, bad, etc.

Questions

Interesting Factoids

SU-CS329H SEP302026

Last edited: September 9, 2026

New Concepts

pairs

Interesting Factoids

proper scoring rule

What’s doing the best is maximized at 0.

SU-CS329Z SEP302026

Last edited: September 9, 2026