_index.org

ICML2026 Poster Day 1-2

Last edited: July 7, 2026

ICML2026 Svete: revisiting padded transformer expressivity

How to transformer give TC0 full coverage, Luke transformer breaks out of T0 into TCD

(Ryan Kattwral and Will Merrill)

ICML2026 Su: rewiring experts on the fly

Take the prefill prefix at runtime and compute LM loss; use to build additive weight on moe routing

ICML2026 Kreitner: efficient numeracy in language models

Just encode numerical tokens in IEEE floats as the embedding + padding. So each <num> token is one token and Fourier math just works.

ICML2026 Poster Day 3

Last edited: July 7, 2026

ICML2026 Zhang: generalist value model

Use embedding similarity model with respect to previous group performance as the value model, such the value model can take previous performance as well as rollout as input, such hard to predict values more generally

ICML2026 Jayalath: compute as teacher

Consensus voting among multiple rollout used as supervisory target for reward and gold derivation

ICML2026 Mirvakabova: dirchlet prior sampling

For expert upcycling, initialize router based on weighted dirchlet priors instead of uniform

Linear Representation Hypothesis

Last edited: July 7, 2026

constituents

  • \(z \in \mathbb{R}^{n}\) which encodes \(n\) distinct concepts, which is “sparse” \(\norm{z}_{1} < \epsilon_{1}\)
  • residual steam \(x \in \mathbb{R}^{m}\) such that \(m \ll n\)

requirements

The Linear Representation Hypothesis states that representation in neural networks can be determined by some:

\begin{equation} \exists F \in \mathbb{R}^{m \times n} \end{equation}

such that \(Fz=x\) for any choice of \(x, z\). That is, neural networks encodes stream-concept mapping linearly.

Importantly, this representation \(F\) also admits a reverse mapping \(G\) which closely recovers the concept, that is:

Notes in Optimizing Thesus for Weight Offloading

Last edited: July 7, 2026

Hidden

Notes on Aether Memory

Last edited: July 7, 2026

no