_index.org

ICLR2026 Workshop: Weight-Space Symmatrices

Last edited: July 7, 2026

One-Liner

Novelty

Motivation

Notable Methods

Key Figs

New Concepts

Notes

ICLR2026 Zhou: DynMuon

Last edited: July 7, 2026

https://arxiv.org/pdf/2605.17109

Motivation

“How can we evolve the optimization algorithm itself during training?”

Notable Methods

For \(W_{t}\) weight matrix and \(M_{t}\) update matrix, we weight the singular directions:

\begin{equation} W_{t+1} = W_{t} - \eta U^{(t)} \Sigma^{\rho^{(t)}} V^{(t)} \end{equation}

and in particular we can dial up and down your choice of \(\rho\) and do stuff. some choices of \(p\):

  • \(p=0\), muon OG, the idea of muon is to set all singular values to 1 so that we only prioritize direction
  • \(p=1\), some aggressive weight-based updates, proposed idea,
  • \(p < 0\), then we prioritize adjusting small directions

generally key insight: “in the beginning, we should prioritize large directions; then, we should prioritize small directions more.”

ICML2026 Kakade: Pretraining

Last edited: July 7, 2026

Pretraining Hard

  • base model quality is hard
  • continual learning is hard
  • shapes may change

Pretraining’s solutions

  • When do we stop?: cosine bakes in an end time
  • How big a batch?: more parallelism (fewer steps)… what’s the limits such that batches are too big?
  • does momentum help?

quadratic models