Three shapes of how humans express preferences:
accept-reject
constituents
Rater \(i\) sees \(j\), and gives \(y_{ij} \in \qty {1,0}\). (thumbs up or down)
requirements
We can write down this as the following model:
\begin{equation} P\qty(y_{ij} = 1) = \sigma \qty(\theta_{i} - b_{j}) \end{equation}
where \(\theta_{i}\) is some notion, and \(b_{j}\) is how “hard” \(j\) is to like, and \(\sigma\) is sigmoid.
additional information
“test model”
The above model makes an independence assumption, sorta like the .
generalization
This already allow us to learn \(\theta\) and \(b\) separately.
ranking
net promoter score
latent score \(\theta_{i} - b_{j} + \epsilon\), and cut points \(c_1 < c_2 < c_3< c_4\), such that
\begin{equation} P\qty(Y \leq k) = F\qty(c_{k} - \theta_{i} + b_{j}) \end{equation}
choices form sets
Rater \(i\) picks from menu \(T = \qty {j,j’, \dots}\), such that \(C_{i}\qty(T) = j\) meaning element \(j\) was chosen from \(T\).
Write \(j \succ_{i} j’\) if:
\begin{equation} C_{i}\qty(\qty {j,j’}) = j \end{equation}
where \(C\) is a choice function of agent \(i\). We often ask questions about say \(P\qty(j \succ_{i} j’)\) and figt models for it.
two types of inferences
- known: infer \(C_{i} \qty(T)\) when we have observation on \(i\) and \(T\)
- unknown: infer \(C_{i} \qty(T)\) when we have observation on \(i\) and some under “bundle” \(T’\)
axioms of set choices
- infinite data: that is, claims of probability is frequentist exact
- no ties, so \(P\qty(j \succ j’) = 1- P\qty(j \succ j’)\)
-\(P\qty(j \succ j’) > 0\)
pop quiz
Question: \(P\qty(j \succ j’) = P\qty(j’ \succ j’’) = \frac{1}{2}\), when what is \(P\qty(j \succ j’’)\)?
IDK! Anything. You can take two such independetn populations and mix them arbitrarily to get any \(P\qty(j \succ j’’)\) value.
additional information
A softmax predictor follows Luce’s Axiom
All three statements are equivalent.
- Logits: There exists \(u : X \to \mathbb{R}\) such that: \(P\qty(C\qty(S) = j) = \text{softmax}_{j} \qty(u\qty(j’)_{j’ \in S})\)
- Independence of irrelevant alternatives: for \(j, j’\) in both \(S\) and \(T\), we have: \(\frac{P\qty(C\qty(S) = j)}{P\qty(C\qty(S) = j’)} = \frac{P\qty(C\qty(T) = j)}{P\qty(C\qty(T) = j’)}\)
- Chain Rule: \(P\qty [C\qty(S) = j] = P\qty [C\qty(T) = j] \qty(\sum_{j’ \in T}^{} P\qty[C\qty(S) = j’])\) “the probability of \(j\) being chosen in \(S\) is the probabliity of \(j\) being chosen in \(T\) times the probability of anybody in \(S\) being chosen”
(1 => 3)
Proof idea: just expand the chain rule by cancelling out softmax denominator:
Write \(P_{s}\qty(j)\) to mean \(P\qty(C\qty(S) = j)\).
Define \(Z_{A} = \sum_{k \in A}^{} e^{u_{k}}\) the bottom of softmax, then for any \(j \in T \subseteq S\), we have
\begin{equation} P_{T}\qty(j)\sum_{j \in T}^{} P_{s}\qty(j) = \frac{e^{u_{j}}}{Z_{T}} \sum_{j \in T}^{} \frac{e^{u_{j}}}{Z_{S}} = \frac{e^{u_{j}}}{Z_{S}} = P_{S}\qty(j) \end{equation}
(3 => 2)
\(j, j’ \in S,T\), define the “overlap” \(R = S \cup T\), then chain rule gives:
\begin{equation} \frac{P_{s}\qty(j)}{P_{s}\qty(j’)} = \frac{P_{R}\qty(j)}{P_{R}\qty(j’)} \frac{\sum_{k \in R}^{} P_{S}\qty(k)}{\sum_{k \in R}^{} P_{T}\qty(k)} = \frac{P_{R}\qty(j)}{P_{R}\qty(j’)} = \frac{P_{T}\qty(j)}{P_{T}\qty(j’)} \end{equation}
where the last equality is applying the chain rule one again in a similar manner.
Ok, so this is great, but suppose your choice function is restricted to a pair e.g., \(C\qty(\qty{j, j’})\) the above is hard to talk about because then the sets \(S\) and \(T\) (maybe disjoint? either way it desn’t typecheck)
So we turn to…
