_index.org

Bradley-Terry Preference Model

Last edited: October 10, 2026

constituents

  • with \(\sigma\qty(z) = \frac{1}{1+e^{-z}}\) sigmoid model…
  • \(i\) user
  • \(x\) context (prompt, etc.)
  • \(a,b\) candidate pairwise items
  • \(r\qty(i,x,a)\) the representation of the choice

requirements

The Brady-Terry gives:

\begin{equation} P\qty(a \succ b | i, x) = \sigma \qty(r\qty(i,x,a) - r\qty(i,x,b)) \end{equation}

additional information

properties of Bradly-Terry

Suppose you are only observing binary cases, how do we induce Choice Data that are ordered preferences?

The following three statements are equivalent.

Brady Terry For \(r: X \to \qty(0, \infty)\) with

Bradly-Terry model fitting

Last edited: October 10, 2026

For user \(i\), context \(x\), candidates \(a,b\). Let us get a binary preference label: \(y=1\) if \(a\) wins, and \(y=0\) if \(b\) wins.

Under a Bradley-Terry Preference Model:

  • utility of one candidate: \(r_{\theta}\qty(i,x,a) \in \mathbb{R}\)
  • comparison logit: \(z_{\theta} = r_{\theta}\qty(i, x,a) - r_{\theta}\qty(i,x,b)\)
  • choice probability under Bradley-Terry Preference Model: \(p = \sigma\qty(z_{\theta})\)

additional information

maximum likelihood estimation given data

Suppose you believe that the probability that someone chooses \(a\) over \(b\) is distributed by a Bernoulli distribution, and you had a bunch of data on that. How would you parametrize it? Well, you’d ask:

Choice Data

Last edited: October 10, 2026

Three shapes of how humans express preferences:

accept-reject

constituents

Rater \(i\) sees \(j\), and gives \(y_{ij} \in \qty {1,0}\). (thumbs up or down)

requirements

We can write down this as the following model:

\begin{equation} P\qty(y_{ij} = 1) = \sigma \qty(\theta_{i} - b_{j}) \end{equation}

where \(\theta_{i}\) is some notion, and \(b_{j}\) is how “hard” \(j\) is to like, and \(\sigma\) is sigmoid.


Alternatively you can write \(b_{j}\) as appeal and make it additive.

factor model

Last edited: October 10, 2026

constituents

  • \(u_{i}, v_{j} \in \mathbb{R}^{K}\)
  • \(u_{i}\), the user \(i\) preferences in \(K\) factors
  • \(v_{j}\), how the item \(j\) scores across these \(K\) factors
  • \(b_{j}\), general appeal of item \(j\)

requirements

\begin{equation} r_{i}\qty(j) = u_{i}^{T} v_{j} + b_{j} \end{equation}

additional information

applied to pairwise feedback in Bradley-Terry Preference Model

\begin{equation} P\qty(j \succ k \mid i) = \sigma \qty(u_{i}^{T}\qty(v_{j} - v_{k}) + b_{j} - b_{k}) \end{equation}

factor geometry

Now, the left thing is a big ol outper product between everybody’s \(u\) everybody’s \(v\).

independence to irrelavent alternatives

Last edited: October 10, 2026

aggregating x vs. y should ignore any possible z

see also softmax predictor follows Luce’s Axiom