Transformers Can Reprogram Themselves. A New Paper Explains the Mind-Blowing Trick.
1/12
You know that magical feeling when an AI like ChatGPT learns a new skill instantly, just from a few examples in your prompt?
It's not magic. And it's not "learning" in the way you think.
New research shows the AI is performing a kind of "ghost fine-tuning" on itself, in real-time. 🤯
A thread on how AI really learns in-context. 🧵
2/12
The central mystery of modern AI is "In-Context Learning" (ICL).
A model is trained for months on massive datasets. Its weights are frozen. And yet, it can learn a new pattern at inference time, without a single permanent update.
How is this possible? It breaks the rules of classic machine learning.
3/1al/12
For years, we've debated this.
Is it really learning? Or just cleverly retrieving facts it already knows?
Most theories were stuck on "toy models" that were too simple, or they just waved their hands and called it an "emergent property." The real mechanism was a black box. Until now.
4/12
A brilliant paper from Google Research, "Learning without training," presents a stunningly simple explanation.
They didn't just look at the attention layer. They looked at the dance between two key parts of a Transformer:
The Self-Attention layer (the context reader)
The MLP layer (the "thinking" neural network that follows)
5/12
Here’s the core idea, using an analogy.
Think of the AI's base knowledge as a powerful, general-purpose computer motherboard (its permanent weights, W).
The Self-Attention layer reads all the examples in your prompt and distills them into a single, specialized "context vector." Think of this vector as a tiny, custom-built microchip.
6/12
And here's the mind-blowing part.
The model then mathematically COMBINES that new microchip (the context vector) with its main motherboard (W).
It creates a temporary, custom-built circuit (W + ∆W) designed specifically to solve the task you just gave it.
(I know, wild, right?)
7/12
This is the "ghost fine-tuning."
Your prompt isn't just passive context; it's an active blueprint for a temporary hardware upgrade.
But that's not even the craziest part.
8/12
The researchers showed that as the AI reads your examples one-by-one, it's like it's running a mini-training session on itself.
Each example adds another layer to this temporary "brain implant," refining the model's behavior step-by-step.
Mathematically, this process mirrors gradient descent—the very optimization algorithm used in training!
9/12
And they proved it.
They took a model's output on a task with a full prompt.
Then, they calculated the "ghost weights" (W + ∆W), gave the model NO prompt at all (just the final query), and got the EXACT same result.
The context was successfully transferred from the prompt and loaded into the model's weights.
10/12
So what does this mean for you?
From now on, when you write a prompt with examples (few-shot prompting), you're not just "showing" the AI what to do.
You are actively, temporarily, reprogramming its neural network. You're a co-pilot, not just a passenger.
11/12
This isn't just a theory about AI. It's a beautiful example of how simple, stacked components can create incredible, emergent abilities that feel like magic.
So the next time you prompt an AI, remember this:
You're not just talking to a machine. You're a temporary programmer, shaping its mind.
12/12
This insight fundamentally changes how I think about "learning." It's more fluid and dynamic than we ever imagined.
What other "magical" AI abilities do you think have simple explanations hiding in plain sight?
Full paper for the brave: arxiv. org/abs/2507.16003v1
Holy shit... Google just cracked the code on AI collaboration.
It's called TUMIX, and it might be the smartest thing they've published all year.
Here's the twist: instead of building one massive brain, they built a team of smaller ones that argue, debate, and improve each other's answers in real-time.
Each agent brings different skills. One codes. Another searches. Another reasons through logic. They tackle problems independently, then share solutions and refine them together until they reach consensus.
The numbers are absolutely wild.
Gemini-2.5 running TUMIX crushes every other reasoning system by up to +17.4%. And it does it at HALF the inference cost.
No retraining. No new data. Just coordination.
But here's where it gets crazy: diversity beats scale.
A team of 15 different agents destroyed 15 copies of the "best" single agent. When they let Gemini design its own new agents? Performance jumped even higher.
The system literally evolved better versions of itself.
This flips everything we thought about AI progress.
We've been obsessing over trillion-parameter models. Turns out, intelligence might come from organization, not just raw size.
The next breakthrough in reasoning won't be a bigger model.
It'll be smaller ones that learned how to think together.
Read the full paper: arxiv. org/abs/2510.01279
Excited to release new repo: nanochat!
(it's among the most unhinged I've written).
Unlike my earlier similar repo nanoGPT which only covered pretraining, nanochat is a minimal, from scratch, full-stack training/inference pipeline of a simple ChatGPT clone in a single, dependency-minimal codebase. You boot up a cloud GPU box, run a single script and in as little as 4 hours later you can talk to your own LLM in a ChatGPT-like web UI.
It weighs ~8,000 lines of imo quite clean code to:
- Train the tokenizer using a new Rust implementation
- Pretrain a Transformer LLM on FineWeb, evaluate CORE score across a number of metrics
- Midtrain on user-assistant conversations from SmolTalk, multiple choice questions, tool use.
- SFT, evaluate the chat model on world knowledge multiple choice (ARC-E/C, MMLU), math (GSM8K), code (HumanEval)
- RL the model optionally on GSM8K with "GRPO"
- Efficient inference the model in an Engine with KV cache, simple prefill/decode, tool use (Python interpreter in a lightweight sandbox), talk to it over CLI or ChatGPT-like WebUI.
- Write a single markdown report card, summarizing and gamifying the whole thing.
Even for as low as ~$100 in cost (~4 hours on an 8XH100 node), you can train a little ChatGPT clone that you can kind of talk to, and which can write stories/poems, answer simple questions. About ~12 hours surpasses GPT-2 CORE metric. As you further scale up towards ~$1000 (~41.6 hours of training), it quickly becomes a lot more coherent and can solve simple math/code problems and take multiple choice tests. E.g. a depth 30 model trained for 24 hours (this is about equal to FLOPs of GPT-3 Small 125M and 1/1000th of GPT-3) gets into 40s on MMLU and 70s on ARC-Easy, 20s on GSM8K, etc.
My goal is to get the full "strong baseline" stack into one cohesive, minimal, readable, hackable, maximally forkable repo. nanochat will be the capstone project of LLM101n (which is still being developed). I think it also has potential to grow into a research harness, or a benchmark, similar to nanoGPT before it. It is by no means finished, tuned or optimized (actually I think there's likely quite a bit of low-hanging fruit), but I think it's at a place where the overall skeleton is ok enough that it can go up on GitHub where all the parts of it can be improved.
Link to repo and a detailed walkthrough of the nanochat speedrun is in the reply.
Holy shit. MIT just built an AI that can rewrite its own code to get smarter 🤯
It’s called SEAL (Self-Adapting Language Models).
Instead of humans fine-tuning it, SEAL reads new info, rewrites it in its own words, and runs gradient updates on itself literally performing self-directed learning.
The results?
✅ +40% boost in factual recall
✅ Outperforms GPT-4.1 using data it generated *itself*
✅ Learns new tasks without any human in the loop
LLMs that finetune themselves are no longer sci-fi.
We just entered the age of self-evolving models.
Paper: jyopari. github. io/posts/seal
🚨 New paper!🚨
A few months ago, Meta released Segment Anything Model (SAM). It's already in photography apps, medicine, and video-generation. Now we’ve invented SAM’s little brother, EfficientSAM. It is small but mighty!
With 20x fewer parameters and 20x faster runtime, EfficientSAM is within 2 points (44.4 AP vs 46.5 AP) of the original SAM model.
How'd we do it? I have 2 words for you: Masked Autoencoders. Check out the following for more details!
Paper: https://t.co/2BufoGKDl4
Project details and demo: https://t.co/5Wuze3x0kD
with: @balakrishnan_vr@klightlm@mukosame@Fanyi_Xiao@dilin_wang@fiandola@raghuraman@vikasc, et al
Gemini: A Family of Highly Capable Multimodal Models
paper: https://t.co/h9cD6XTYvK
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultra model advances the state of the art in 30 of 32 of these benchmarks — notably being the first model to achieve human-expert performance on the well-studied exam benchmark MMLU, and improving the state of the art in every one of the 20 multimodal benchmarks we examined. We believe that the new capabilities of Gemini models in cross-modal reasoning and language understanding will enable a wide variety of use cases and we discuss our approach toward deploying them responsibly to users.
What is the simplest way to animate 3D Gaussians?
What if the mesh tracking is not perfect?
No canonical space, no deformation network, just triangles and Gaussians.
Amazing collaboration with Davide Davoli, @liamschoneveld, @TobiasKirschst1 ,@SGiebenhain, and @matthias!
IBM & Meta are launching the AI Alliance to advance *open* & reliable AI.
The list of over 50 founding members from industry, government, and academia include AMD, Anyscale, CERN, Hugging Face, the Linux Foundation, NASA....
https://t.co/NdcMb7GMXv
Many AI researchers today display signs of burn out. Companies are racing to build bigger models, individuals rush to publish more papers.
I miss the old days when things were slow. less noise, less hype. More time to savor.
Collectively as a community, are we happier?
@EnisField@vinceflibustier Short, t-shirt, basket voire parfois tongs.
Chercheur en IA & Robotique dans une très grande boîte…
Comme quoi l’habit ne fait pas le moine. Ouvrir son esprit comme on dit 😊
Voici le scientifique le plus prolifique d'Espagne en matière de publications scientifiques.
L'année passée, il a signé 176 articles.
1 article tous les 2 jours en gros.
Et c'est l'occasion de découvrir certaines pratiques, disons, intéressantes.
1/15
LERF: Language Embedded Radiance Fields
TL;DR: Grounding CLIP vectors volumetrically inside a NeRF allows flexible natural language queries in 3D
abs: https://t.co/Xtk6lwBs6x
project page: https://t.co/FkT0F9jqDB
Here’s our LLaMA-13B fine tuned with RLHF & SFT
This has only been trained on 3% of our total dataset size, and no NSFW yet.
It is better than GPT3.5
We’re open sourcing all weights and inference code in a few days after training
This critique of AI gets absolutely everything wrong.
Quote: "There have been no major breakthroughs in the academic discipline of artificial intelligence for a couple of decades"
What ???
Seriously ???