ICML2026 Poster Day 1-2
Last edited: July 7, 2026ICML2026 Svete: revisiting padded transformer expressivity
How to transformer give TC0 full coverage, Luke transformer breaks out of T0 into TCD
(Ryan Kattwral and Will Merrill)
ICML2026 Su: rewiring experts on the fly
Take the prefill prefix at runtime and compute LM loss; use to build additive weight on moe routing
ICML2026 Kreitner: efficient numeracy in language models
Just encode numerical tokens in IEEE floats as the embedding + padding. So each <num> token is one token and Fourier math just works.
ICML2026 Poster Day 3
Last edited: July 7, 2026ICML2026 Zhang: generalist value model
Use embedding similarity model with respect to previous group performance as the value model, such the value model can take previous performance as well as rollout as input, such hard to predict values more generally
ICML2026 Jayalath: compute as teacher
Consensus voting among multiple rollout used as supervisory target for reward and gold derivation
ICML2026 Mirvakabova: dirchlet prior sampling
For expert upcycling, initialize router based on weighted dirchlet priors instead of uniform
Linear Representation Hypothesis
Last edited: July 7, 2026constituents
- \(z \in \mathbb{R}^{n}\) which encodes \(n\) distinct concepts, which is “sparse” \(\norm{z}_{1} < \epsilon_{1}\)
- residual steam \(x \in \mathbb{R}^{m}\) such that \(m \ll n\)
requirements
The Linear Representation Hypothesis states that representation in neural networks can be determined by some:
\begin{equation} \exists F \in \mathbb{R}^{m \times n} \end{equation}
such that \(Fz=x\) for any choice of \(x, z\). That is, neural networks encodes stream-concept mapping linearly.
Importantly, this representation \(F\) also admits a reverse mapping \(G\) which closely recovers the concept, that is:
Notes in Optimizing Thesus for Weight Offloading
Last edited: July 7, 2026Hidden
Notes on Aether Memory
Last edited: July 7, 2026no
