_index.org

ICLR2025 Kilani: MrT5 Tokenizer-Free

Last edited: August 8, 2025

Motivation

ByteT5 is very expensive (because you have to have a residual on every damn token)

MrT5

MrT5 uses a soft attention masking gate at pretraining time to delete unused tokens; at inference time we use a hard cut.

Cool: MrT5 learns language independent compression rate (different languages have different rates).

ICLR2025 Li: MoE is secretly an embedding

Last edited: August 8, 2025

motivation

Can we directly extract embeddings from MoE forwarding routing weights (i.e., compared to traditional residual stream information)?

Key Insight

Using residual states vs. forwarding weights as semantic searc embeddings offer complementary strengths (i.e., when one method fails, the other one succeeds more)

Method

Create an aggregate embedding:

\begin{equation} E_{j} = X_{j} + \alpha W_{j} \end{equation}

where \(W_{j}\) is the routing weight of the residual, and \(X_{j}\) is the residual.

ICLR2025 Mathur: MIND Adaptive Thinking with Dynamic Computation

Last edited: August 8, 2025

Motivation

Standard computation doesn’t adapt.

Fixed-Point Iteration for Adaptation

method: CNN

  1. for every layer, perform fixed-point iteration until convergence to mask out (what exactly?)
  2. supervise also an “introspection model” to skip the entire fixed point
  3. loss: LM + supervision for the introspection model

method: MIND-transformer

  1. for every layer, perform fixed-point iteration until attention activation convergence
  2. ditto introspection as above

ICLR2025 Neitemeier: Hierachical Autoregressive Transformers

Last edited: August 8, 2025

“A Byte Level transformer, with some compression”

Key insight: use a [CLS] token in front of every word to train a small “tokenizer”, and then do a normal transformer on the [CLS] tokens, and then autoregressive decode out the single bytes.

Method

Hierarchical Autoregressive Transformers

We put a [cls] in front of every word. So the input looks like

[CLS] M y _ [CLS] n a m e _ [CLS] i s

We then run a small encoder over each sequence. And then you take the encoded [CLS], and run