For user \(i\), context \(x\), candidates \(a,b\). Let us get a binary preference label: \(y=1\) if \(a\) wins, and \(y=0\) if \(b\) wins.
Under a Bradley-Terry Preference Model:
- utility of one candidate: \(r_{\theta}\qty(i,x,a) \in \mathbb{R}\)
- comparison logit: \(z_{\theta} = r_{\theta}\qty(i, x,a) - r_{\theta}\qty(i,x,b)\)
- choice probability under Bradley-Terry Preference Model: \(p = \sigma\qty(z_{\theta})\)
additional information
maximum likelihood estimation given data
Suppose you believe that the probability that someone chooses \(a\) over \(b\) is distributed by a Bernoulli distribution, and you had a bunch of data on that. How would you parametrize it? Well, you’d ask:
What’s the \(p\) that goes into the Bernoulli? Since \(x \sim \text{Bern}\qty(p)\) takes a \(p\); well if you believe in Bradley-Terry Preference Model, then, we can write:
\begin{equation} x\sim \text{Bern}\qty(\sigma\qty(z_{\theta})) \end{equation}
So, let’s estimate parameters for \(\theta\) given some data. Standard Maximum Likelihood Parameter Learning:
\begin{align} \theta &= \arg\max_{\theta} P\qty(D | \theta) \\ &= \arg\max_{\theta} \prod_{i=1}^{n} \sigma\qty(z_{\theta})^{y_{i}} \qty(1-\sigma\qty(z_{\theta}))^{1-y_{i}} \\ &= \arg\min_{\theta} - \sum_{i=1}^{n} y_{i}\log \sigma\qty(z_{\theta}) + \qty(1-y_{i}) \log \qty(1-\sigma\qty(z_{\theta})) \end{align}
Now recall:
\begin{equation} \text{softplus}\qty(z) = \log \qty(1+e^{-z}) \end{equation}
\begin{equation} \text{softplus}\qty(-z) = \text{softplus}\qty(z) - z \end{equation}
\begin{equation} \dv{z}\text{softplus}\qty(z) = \sigma\qty(z) \end{equation}
and that:
\begin{equation} 1-\sigma\qty(z) = \sigma\qty(-z) \end{equation}
since sigmoid is symmetric, so let’s write:
\begin{align} \theta &= \arg\max_{\theta} P\qty(D | \theta) \\ &= \arg\max_{\theta} \prod_{i=1}^{n} \sigma\qty(z_{\theta})^{y_{i}} \qty(1-\sigma\qty(z_{\theta}))^{1-y_{i}} \\ &= \arg\min_{\theta} \sum_{i=1}^{n} y_{i}\log \qty(1+e^{-z_{\theta}}) + \qty(1-y_{i}) \log \qty(1+e^{z_{\theta}}) \\ &= \arg\min_{\theta} \sum_{i=1}^{n} y_{i}\text{softplus}\qty(-z_{\theta}) + \qty(1-y_{i}) \text{softplus}\qty(z_{\theta}) \\ &= \arg\min_{\theta} \sum_{i=1}^{n} y_{i}\qty [\text{softplus}\qty(z_{\theta}) -z_{\theta}] + \qty(1-y_{i}) \text{softplus}\qty(z_{\theta}) \\ &= \arg\min_{\theta} \sum_{i=1}^{n} \text{softplus}\qty(z_{\theta}) - y_{i}z_{\theta} \end{align}
And the gradient for this, for completeness, is:
\begin{equation} L’ = \sum_{i=1}^{n}\sigma\qty(z_{\theta}) - y_{i} \end{equation}
