Linear Representation Hypothesis
Last edited: July 7, 2026constituents
- \(z \in \mathbb{R}^{n}\) which encodes \(n\) distinct concepts, which is “sparse” \(\norm{z}_{1} < \epsilon_{1}\)
- residual steam \(x \in \mathbb{R}^{m}\) such that \(m \ll n\)
requirements
The Linear Representation Hypothesis states that representation in neural networks can be determined by some:
\begin{equation} \exists F \in \mathbb{R}^{m \times n} \end{equation}
such that \(Fz=x\) for any choice of \(x, z\). That is, neural networks encodes stream-concept mapping linearly.
Importantly, this representation \(F\) also admits a reverse mapping \(G\) which closely recovers the concept, that is:
Notes in Optimizing Thesus for Weight Offloading
Last edited: July 7, 2026Hidden
Notes on Aether Memory
Last edited: July 7, 2026no
AA228/CS238: Probability Review!
Last edited: June 6, 2026Random Variable
random variables takes on different values with different probabilities. Each value a random variable take on is an event.
For instance, here’s a random variable representing a die: \(X\). It can takes on the following values, with the following probabilities:
\begin{align} P(X=1) = \frac{1}{6}\\ P(X=2) = \frac{1}{6}\\ \dots \\ P(X=6) = \frac{1}{6} \end{align}
where each assignment \(X=k\) is what we refer to above as an event.
The set of assignments of a random variable and their associated probability is called a distribution: distributions “assigns probabilities to outcomes.” When we say a certain random variable \(X\) is “distributed” following a distribution \(D\), we say \(X \sim D\). Semantically, we say \(X\) is a \(D\) random variable.
adventuretime
Last edited: June 6, 2026- 1 Executive Summary
- 2 Core Design Principles
- 3 Non-Goals
- 4 System Overview
- 5 Hardware Architecture (Recommended)
- 6 Network & Communication Diagram (Textual)
- 7 Storage Model
- 7.1 Invariant
- 7.2 Checkpoint Flow
- 7.3 Artifacts & Logs
- 8 Unified Checkpointing (JAX Pytrees)
- 9 Repository Structure (Tech Spec)
- 10 Execution Model
- 10.1 Job Lifecycle
- 10.2 Backend Interface
- 11 DAG / “Ray-lite” Model
- 12 Example YAML Specifications
- 12.1 Backend Inventory
- 12.2 Storage
- 12.3 Single Run Spec
- 12.4 DAG Spec (Tokenize → Train → Rollouts)
- 13 Implementation Plan
- 13.1 Phase 0 (2–3 weeks)
- 13.2 Phase 1 (4–6 weeks)
- 13.3 Phase 2 (3–4 weeks)
- 13.4 Phase 3 (optional)
- 14 Cost Estimates
- 14.1 One-Time Hardware (Target ~$50k)
- 14.2 Ongoing
- 15 Risks & Mitigations
- 16 Success Criteria
- 17 Conclusion
After a conversation with an LM https://chatgpt.com/share/697143db-c3e0-8000-b56c-07cf7ca43795 the following proposal was generated.
