Choice Data
Last edited: September 9, 2026Three shapes of how humans express preferences:
accept-reject
constituents
Rater \(i\) sees \(j\), and gives \(y_{ij} \in \qty {1,0}\). (thumbs up or down)
requirements
We can write down this as the following model:
\begin{equation} P\qty(y_{ij} = 1) = \sigma \qty(\theta_{i} - b_{j}) \end{equation}
where \(\theta_{i}\) is some notion, and \(b_{j}\) is how “hard” \(j\) is to like, and \(\sigma\) is sigmoid.
additional information
“test model”
The above model makes an independence assumption, sorta like the .
elimination by aspects
Last edited: September 9, 2026Each item is a set of “aspect.” Aspect \(\alpha\) has weight \(w\qty(\alpha) > 0\). And the way to get probability distribution of which items to pick:
given menu \(S\), discard aspects shared by every item in \(S\). then:
- choose an aspects with probability \(\propto w\qty(\alpha)\)
- keep only the items that remain, repeat until the list is broken down
Insight: “similar thing being added increasingly should not as much more decision share. Red bus + green ice cream adds more decision share than Red bus + blue bus + green ice cream which adds more than Red + Blue + green …. + purple bus + green ice cream”
Fechner and Quadrupel Condition
Last edited: September 9, 2026Fechner / strong utility: there are \(u : X \to \mathbb{R}\) and \(F: \mathbb{R} \to \qty(0,1)\), continuous, strictly increasing, and \(F\qty(-t) = 1 - F\qty(t)\), and
\begin{equation} P\qty(j \succ j’) = F\qty(u_{j} - u_{j’}) \end{equation}
Quadruple Condition: for all choices of \(j, j’, j’’, j’’’\), we have:
\begin{equation} P\qty(j \succ j’) \geq P\qty(j’’ \succ j’’’) \Leftrightarrow P\qty(j \succ j’’) \geq P\qty(j’ \succ j’’’) \end{equation}
Solvability: for \(j, j’, j’’\) and \(\lambda \in [0,1]\) with \(P\qty(j \succ j’) \leq \lambda \leq P\qty(j \succ j’’)\), some \(j’’’\) has \(P\qty(j \succ j’’’) = \lambda\)
knowledgebase testing page 2
Last edited: September 9, 2026holy shit \(48\) wow wow yes \(8\)
\begin{equation} \frac{12}{3} - \int_{i}^{j} 8 \end{equation}
thta’s crazy.
The simplest channel-specific baseline I’d recommend is an input-gated residual belief update:
\[ q=\operatorname{LN}_x(x) \]
\begin{equation} c=\operatorname{LN}_s(s),\qquad s= \begin{cases} e_{\text{boundary}} & \text{first pass}\\ x_{\text{loop}} & \text{later passes} \end{cases} \end{equation}
\[ g=\tanh(\theta),\qquad \theta\in\mathbb{R}^{d},\quad\theta_0=0 \]
\[ q_{\text{updated}}=q+g\odot c \]
\[ x\leftarrow x+\operatorname{Attn}(q_{\text{updated}}) \]
Then keep the ordinary MLP update unchanged.
This is essentially the input-gate portion of an LSTM:
