_index.org

expectation maximization

Last edited: January 1, 2026

Sorta like “distribution-based k-means clustering”. guarantees convergence (i.e. each parameter will converge to the maximum possible parameter).

constituents

A Gaussian mixture model!

requirements

Two steps:

e-step

“guess the value of \(z^{(i)}\); soft guesses of cluster assignments”

\begin{align} w_{j}^{(i)} &= p\qty(z^{(i)} = j | x^{(i)} ; \phi, \mu, \Sigma) \\ &= \frac{p\qty(x^{(i)} | z^{(i)}=j) p\qty(z^{(i)}=j))}{\sum_{l=1}^{k}p\qty(x^{(i)} | z^{(i)}=l) p\qty(z^{(i)}=l))} \end{align}

Where we have:

  • \(p\qty(x^{(i)} |z^{(i)}=j)\) from the Gaussian distribution, where we have \(\Sigma_{j}\) and \(\mu_{j}\) for the parameters of our Gaussian \(j\).
  • \(p\qty(z^{(i)} =j)\) is just \(\phi_{j}\) which we are learning

These weights \(w_{j}\) are how much the model believes it belongs to each cluster.

Jensen's Inequality

Last edited: January 1, 2026

linear edition

if \(f\) is convex, then for \(x,y \in \text{dom }f, 0 \leq \theta \leq 1\), then:

\begin{equation} f\qty(\theta x + \qty(1-\theta) y) \leq \theta f\qty(x) + \qty(1-\theta) f\qty(y) \end{equation}

probabilistic extension

Let \(f\) be a convex function; that is, \(f’’\qty(x) \geq 0\); let \(x\) be a random variable. Then, \(f\qty(\mathbb{E}[x]) \leq \mathbb{E}\qty [f\qty(x)]\).

Further, if \(f\) is strictly convex, that is \(f’’\qty(x) > 0\), then \(\mathbb{E}\qty [f\qty(x)] = f\qty(\mathbb{E}[x])\), that is, \(x\) is constant.

model-free inte

Last edited: January 1, 2026

model-free reinforcement learning

Last edited: January 1, 2026

In model-based reinforcement learning, we tried real hard to get \(T\) and \(R\). What if we just estimated \(Q(s,a)\) directly? model-free reinforcement learning tends to be quite slow, compared to model-based reinforcement learning methods.

\begin{equation} \frac{1}{2} \qty(\frac{1}{2}) \end{equation}

review: estimating mean of a random variable

we got \(m\) points \(x^{(1 \dots m)} \in X\) , what is the mean of \(X\)?

\begin{equation} \hat{x_{m}} = \frac{1}{m} \sum_{i=1}^{m} x^{(i)} \end{equation}

\begin{equation} \hat{x}_{m} = \hat{x}_{m-1} + \frac{1}{m} (x^{(m)} - \hat{x}_{m-1}) \end{equation}

norm

Last edited: January 1, 2026

The norm is the “length” of a vector, defined generally using the inner product as:

\begin{equation} \|v\| = \sqrt{\langle v,v \rangle} \end{equation}

additional information

properties of the norm

  1. nonnegativity: \(\norm{v} \geq 0\)
  2. zero: \(\|v\| = 0\) IFF \(v=0\)
  3. first-degree homogeneity: \(\|\lambda v\| = |\lambda|\|v\|\)
  4. triangle inequality: \(\norm{x+y} \leq \norm{x} + \norm{y}\)

inner product is a norm

Inner product is a norm:

  1. By definition of an inner product, \(\langle v,v \rangle = 0\) only when \(v=0\)
  2. See algebra:

\begin{align} \|\lambda v\|^{2} &= \langle \lambda v, \lambda v \rangle \\ &= \lambda \langle v, \lambda v \rangle \\ &= \lambda \bar{\lambda} \langle v,v \rangle \\ &= |\lambda |^{2} \|v\|^{2} \end{align}