ICLR2026 Workshop: Learning Dynamics
Last edited: July 7, 2026ICLR2026 Workshop: Weight-Space Symmatrices
Last edited: July 7, 2026One-Liner
Novelty
Motivation
Notable Methods
Key Figs
New Concepts
Notes
ICLR2026 Zhou: DynMuon
Last edited: July 7, 2026https://arxiv.org/pdf/2605.17109
Motivation
“How can we evolve the optimization algorithm itself during training?”
Notable Methods
For \(W_{t}\) weight matrix and \(M_{t}\) update matrix, we weight the singular directions:
\begin{equation} W_{t+1} = W_{t} - \eta U^{(t)} \Sigma^{\rho^{(t)}} V^{(t)} \end{equation}
and in particular we can dial up and down your choice of \(\rho\) and do stuff. some choices of \(p\):
- \(p=0\), muon OG, the idea of muon is to set all singular values to 1 so that we only prioritize direction
- \(p=1\), some aggressive weight-based updates, proposed idea,
- \(p < 0\), then we prioritize adjusting small directions
generally key insight: “in the beginning, we should prioritize large directions; then, we should prioritize small directions more.”
ICML2026 Index
Last edited: July 7, 2026Sessions
Posters
Workshops
ICML2026 Kakade: Pretraining
Last edited: July 7, 2026Pretraining Hard
- base model quality is hard
- continual learning is hard
- shapes may change
Pretraining’s solutions
- When do we stop?: cosine bakes in an end time
- How big a batch?: more parallelism (fewer steps)… what’s the limits such that batches are too big?
- does momentum help?
