> nit
Slide numbers :)
> How having a probability related to generation? / Why is LM discriminative.
Re the student question about why a language model isn’t descriminative. It’s sorta true that LMs are a “classifier”, as are many descriminative models, but notice that there’s no “correct” classification.
Instead, you are normatively “supposed” to sample the distribution instead of simply outputting the “correct” next token because no such token exist. In this sense, its generative.
> Attention Figure
I think it may help students in terms of showing what one ROW / COL do in attn, which can make the figures less confusing.
> Linear Attention Figure
I’m not sure I’m following this figure. It may help reducing arrows / also labeling each chunk / zooming in.
> Context length / long context extension
224n teaches RoPE. could be useful to bring this up. which is technically context length agnostic. And frequently this means we don’t “train position embeddings” anymore.
> content
This is a very ambitious lecture! It is also going a but scattered… Could help if content is tied to agents at each stage. But also the intuitive instructions is AMAZING.
> “decode is memory bound”
Its memory bandwith bound because KV cache has to get constantly written, read, etc. and in long context positions since there’s not enough HBM etc., to store the entire KV cache.
> student questions
During one question pause there was a question at the other end of the room that was missed. Sorry I didn’t flag it to you clearly…
