Bradly-Terry model fitting

For user \(i\), context \(x\), candidates \(a,b\). Let us get a binary preference label: \(y=1\) if \(a\) wins, and \(y=0\) if \(b\) wins.

Under a Bradley-Terry Preference Model:

  • utility of one candidate: \(r_{\theta}\qty(i,x,a) \in \mathbb{R}\)
  • comparison logit: \(z_{\theta} = r_{\theta}\qty(i, x,a) - r_{\theta}\qty(i,x,b)\)
  • choice probability under Bradley-Terry Preference Model: \(p = \sigma\qty(z_{\theta})\)

additional information

maximum likelihood estimation given data

Suppose you believe that the probability that someone chooses \(a\) over \(b\) is distributed by a Bernoulli distribution, and you had a bunch of data on that. How would you parametrize it? Well, you’d ask:

What’s the \(p\) that goes into the Bernoulli? Since \(x \sim \text{Bern}\qty(p)\) takes a \(p\); well if you believe in Bradley-Terry Preference Model, then, we can write:

\begin{equation} x\sim \text{Bern}\qty(\sigma\qty(z_{\theta})) \end{equation}

So, let’s estimate parameters for \(\theta\) given some data. Standard Maximum Likelihood Parameter Learning:

\begin{align} \theta &= \arg\max_{\theta} P\qty(D | \theta) \\ &= \arg\max_{\theta} \prod_{i=1}^{n} \sigma\qty(z_{\theta})^{y_{i}} \qty(1-\sigma\qty(z_{\theta}))^{1-y_{i}} \\ &= \arg\min_{\theta} - \sum_{i=1}^{n} y_{i}\log \sigma\qty(z_{\theta}) + \qty(1-y_{i}) \log \qty(1-\sigma\qty(z_{\theta})) \end{align}

Now recall:

\begin{equation} \text{softplus}\qty(z) = \log \qty(1+e^{-z}) \end{equation}

\begin{equation} \text{softplus}\qty(-z) = \text{softplus}\qty(z) - z \end{equation}

\begin{equation} \dv{z}\text{softplus}\qty(z) = \sigma\qty(z) \end{equation}

and that:

\begin{equation} 1-\sigma\qty(z) = \sigma\qty(-z) \end{equation}

since sigmoid is symmetric, so let’s write:

\begin{align} \theta &= \arg\max_{\theta} P\qty(D | \theta) \\ &= \arg\max_{\theta} \prod_{i=1}^{n} \sigma\qty(z_{\theta})^{y_{i}} \qty(1-\sigma\qty(z_{\theta}))^{1-y_{i}} \\ &= \arg\min_{\theta} \sum_{i=1}^{n} y_{i}\log \qty(1+e^{-z_{\theta}}) + \qty(1-y_{i}) \log \qty(1+e^{z_{\theta}}) \\ &= \arg\min_{\theta} \sum_{i=1}^{n} y_{i}\text{softplus}\qty(-z_{\theta}) + \qty(1-y_{i}) \text{softplus}\qty(z_{\theta}) \\ &= \arg\min_{\theta} \sum_{i=1}^{n} y_{i}\qty [\text{softplus}\qty(z_{\theta}) -z_{\theta}] + \qty(1-y_{i}) \text{softplus}\qty(z_{\theta}) \\ &= \arg\min_{\theta} \sum_{i=1}^{n} \text{softplus}\qty(z_{\theta}) - y_{i}z_{\theta} \end{align}

And the gradient for this, for completeness, is:

\begin{equation} L’ = \sum_{i=1}^{n}\sigma\qty(z_{\theta}) - y_{i} \end{equation}