factor model

constituents

  • \(u_{i}, v_{j} \in \mathbb{R}^{K}\)
  • \(u_{i}\), the user \(i\) preferences in \(K\) factors
  • \(v_{j}\), how the item \(j\) scores across these \(K\) factors
  • \(b_{j}\), general appeal of item \(j\)

requirements

\begin{equation} r_{i}\qty(j) = u_{i}^{T} v_{j} + b_{j} \end{equation}

additional information

applied to pairwise feedback in Bradley-Terry Preference Model

\begin{equation} P\qty(j \succ k \mid i) = \sigma \qty(u_{i}^{T}\qty(v_{j} - v_{k}) + b_{j} - b_{k}) \end{equation}

factor geometry

Now, the left thing is a big ol outper product between everybody’s \(u\) everybody’s \(v\).

So \(U^{T}V\) you can do a eigenvalue decomposition to help you explain how many factors required to explain the data.

In some sense, the dot product scores how an item is aligned with the user’s preference direction.

You are essentially “decomposing” each user-item preferences into a rank-\(K\) for preferences:

\begin{equation} \text{rank}\qty( R) = \text{rank}\qty(UV^{T}) \leq K+1 \end{equation}

collecting feedback data

How do you collect data about users? You very rarely select itemwise binary feedback.

Instead, you collect pairwise feedback and then fit the model using something like Bradly-Terry.

cold start

  • new pair, known users, known items: we have \(u_{i}, v_{j}, v_{k}\), so just do the math
  • new item: a fresh vector \(v_{\text{new}}\) which we predict from item content, feature, calibration
  • new user: a fresh vector \(u_{\text{new}}\) which we predict from user information, history, calibration