Did you know MHA, GQA, and MLA are all just special cases of a broader architecture? 🧩
Tucker Attention uses tensor factorization to unify them, drastically cutting parameters while staying fully compatible with FlashAttention.
Elegant math!
https://t.co/Vi3IEUs8eR
Master’s thesis vs. PhD thesis — they are not the same thing.
A Master’s thesis usually shows that you can conduct focused research using established methods.
A PhD thesis is expected to make an original contribution to knowledge through deeper, independent research.
The exact requirements vary by university and field.
Unpopular view: the "latent space" in many existing works is NOT really latent. Model's "hidden" states, activations, KV caches--these are just another representation of observed data like text, especially for opensource models. They are observations.
The true latents are what GENERATE the observational world: latent thoughts generate what we/agents say, abstract concepts generate what we see, etc. We need something like an encoder to get these.
Thought Communication recovers those latents from model "hidden" states, and shows how they improve reasoning and collaboration. https://t.co/2Omu0Y82rD
Introducing Flow Reasoning Models.
We developed a recurrent flow-based architecture to efficiently solve structured reasoning problems (e.g., Sudoku).
FRMs apply continuous flows to discrete data and recurrently refine their past mistakes through self-conditioning.
Next-token prediction is myopic. What if transformers learn to predict their own next latent state?
🌠 We present 𝗡𝗲𝘅𝘁-𝗟𝗮𝘁𝗲𝗻𝘁 𝗣𝗿𝗲𝗱𝗶𝗰𝘁𝗶𝗼𝗻 (𝗡𝗲𝘅𝘁𝗟𝗮𝘁): a self-supervised learning method that teaches transformers to form compact world models for reasoning and planning. It also unlocks up to 3.3x faster inference via self-speculative decoding! 🚀
Rosarians ❤️
To anyone studying, chasing a difficult dreams, working hard, or simply trying their very best to keep on going - I hope you know how incredibly proud I am of you for putting one foot infront of the other everyday - Keep going one step at a time! Wherever you are you carry the love and strength of the Bloodflame Kingdom wherever you go! Huzzah! 🔥
勉強を頑張っているあなたへ、難しい夢を追いかけているあなたへ、一生懸命努力しているあなたへ、そしてただ毎日を精一杯頑張って前に進み続けているあなたへ。
毎日、一歩一歩前へ進み続けていることを、私は心から誇りに思っています。
焦らず、一歩ずつ進んでいこう!
どこにいても、ブラッドフレイム王国の愛と強さは、いつもあなたと共にあります!HUZZAH!🌹🔥
🫂
every chinese frontier model now uses linear attention (except deepseek)
they all use (except kimi) sparse attention with similar indexer/compression designs to maximize efficiency
they all use "fancy" residuals (mHC, attention residual, gated residual) to maximize signal propagation
they all use Muon
very exciting time for frontier (and efficient) oss models, the beauty of open research :)
A matrix is not just a grid of numbers. Its structure determines what it can do.
Square matrices can be invertible or singular. Some have deeper structure: symmetric matrices, for example, can always be diagonalized using an orthogonal matrix:
S = QΛQᵀ
That is what makes matrix factorization so powerful. Instead of treating every matrix as a collection of numbers, linear algebra gives us ways to reveal the structure hidden inside it.
Hey Momos,
In August, I found out that certain companies or people I trusted have betrayed a lot of streamer’s and the audience's trust. This began a really intense downward spiral for me where I seriously reconsidered a lot of what I know about this industry and the companies inside it.
I’ve been an actual ghost this past few weeks, curtains drawn, phone muted and barely existing in a bad mental health induced cocoon. I burnt out HARD. couldn't work & could barely take care of myself.
I’m having my family visit and then I’ll be heading back to my parents place for a bit ( yes, at my big age ). I just need to be in a place where people can check on me and I don’t have to pretend that I’m okay. Absolutely no shame in it, but I’m just not okay to be alone right now.
I’m gonna try my best to stay strong, get back on my little bug toes and come back in September. I’ve really missed you all, I just need to get a little bit better mentally before I’m back consistently for you.
Please hold tight, I just need a little bit more time 🤍🪳
Transformer by hand ✍️ ~ 6 steps walkthrough below
Studying the Transformer architecture is like opening up the engine hood of your car. So many unknown parts: embeddings, positional encoding, attention weighting, self-attention, cross-attention, multi-head attention, layer norm, skip connections, softmax, linear, Nx, shifted right, query, key, value, masking.
But which of those actually make the car run?
I think at the core, there are two most essential components:
Attention weighting and the feed-forward network.
Everything else is an enhancement to make it run faster and longer, which is how we got from a car to a truck, and to the word "large" in large language model.
So I drew and calculated those two parts entirely by hand.
Goal: push five features through one transformer block, filling in every cell yourself.
1. Given
Five positions of input features, arriving from the previous block.
2. Attention matrix
Let us feed all five features to a query-key module (QK) and read back an attention weight matrix, A. The details of that module are a post of their own.
3. Attention weighting
We multiply the input features by A to get the attention weighted features, Z. Still five positions. The effect is to combine features *across positions*, horizontally: X1 becomes X1 + X2, X2 becomes X2 + X3, and so on.
4. First layer
Let us feed all five weighted features into the first layer of the FFN. Multiply by the weights and biases. This time the combining happens *across feature dimensions*, vertically, and each feature grows from 3 numbers to 4.
Note that every position goes through the same weight matrix. That is what "position-wise" means.
5. ReLU
We cross out the negatives. They become zeros.
6. Second layer
Let us bring it back down: 4 dimensions to 3. The output feeds the next block, which has a completely separate set of parameters, and the whole thing runs again.
You have just calculated a transformer block by hand. ✍️
The takeaway: the two parts are doing two different jobs, and neither one alone is enough. Attention mixes across positions, so a feature can see its neighbours. The FFN mixes across feature dimensions, so each position can think about itself. Horizontal, then vertical. Then that pattern repeats N times, each block with its own separate set of weights. That is the Nx from the list up top, and that is what makes the transformer run.
💾 Save this post!
#AIbyHand #Transformers #DeepLearning
"Latent On-Policy Self-Distillation"
Self-improving agents usually learn from experience through a teacher that gets extra information, but humans still have to decide how that experience should be represented.
This paper makes that representation learnable. Past trajectories are compressed into latent context that helps the model become a better teacher for itself on the exact states the student visits.
The student then distills that guidance into its own weights, so the extra context is only needed for training, while still improving tool use and coding with much fewer rollouts.
https://t.co/5kMhyCFLqw
What if every measurable quantity in physics occupied a vertex on a four-dimensional cube?
Dimensional analysis expresses any quantity as
[ Lᵃ Tᵇ Mᶜ Θᵈ ] where the integer exponents a,b,c,d range over {-3…3}.
Plotting these four independent powers as coordinates produces a tesseract: 16 vertices become pure dimensional states, 32 edges become the allowed transitions between them.
This geometric view of units is already at work inside physics-informed neural networks, automatic equation checkers, and multi-physics simulation codes that refuse to add meters to seconds.
The shape of measurement is itself a higher-dimensional object.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
The human brain is strikingly modular, with distinct networks for language, formal reasoning, social reasoning, and physical reasoning. Is this a fundamental principle of how intelligent systems are built, or an accident of biological evolution?
In our latest preprint, we find that a similar modular organization emerges in Large Language Models, another class of intelligent system.
Brains and LLMs are shaped by entirely different kinds of optimization (biological evolution vs. gradient descent). That they arrive at the same modular design anyway suggests modularity may be a fundamental property of intelligent systems.
🌐 Web: https://t.co/ZKrnTSSuSf
📄 Paper: https://t.co/ZibBXz3PUy
💻 Code & data: https://t.co/uBo5iOYNjy
Using circuit analyses across 46 tasks spanning four cognitive domains, we find:
1️⃣ Tasks that draw on the same network in humans recruit overlapping units in LLMs, while tasks drawing on different networks recruit distinct units.
2️⃣ These units are causally linked to model behavior. Ablating the units critical for one domain impairs performance in that domain (−26% accuracy) but barely touches the others (−2.5%).
This project has been in the works for a while :) Huge thanks to my advisors @jacobandreas@ev_fedorenko@devarda_a, and to @Nancy_Kanwisher for valuable conceptual input and feedback throughout. #MIT
Fate/Grand Order (English) celebrates its 9th anniversary! The new illustration by Namie (@nambarimasu) is a grand slam! Masters, look forward to more surprises and news coming soon!
#FateGOUSA