🚀MetaChat is out now in @ScienceAdvances!
Still feels surreal to see fabrication-ready design files produced in minutes instead of the usual days to weeks.
Looking forward to an agentic future for research and engineering in photonics🤖
"...push the boundaries of science using an infinite supply of ultrafast agentic intelligence."
-@JonFanLab
Excited to present MetaChat, a multi-agentic🤖 framework that automates metasurface design in near real-time (as opposed to weeks of manual specialized work)!
🧵(1/5)
Here’s a short summary of our recent #xenocortication article in @Nature, including the rationale, approach, key findings, initial applications to disease modeling, and current limitations.
A tremendous effort led by Konstantin Kaganovsky, Kevin Kelley, Tilo Gschwind, and Paul Harary.
Link to the article below
1/ I have long dreamed of systems that could causally model human neurobiology across scales, from molecular to behavioral.
In @Nature, delighted to share our efforts at creating xenocortical mice where human organoids grow to occupy most of mouse cortex.
Can language models learn useful priors without ever seeing language?
We pre-pre-train transformers on neural cellular automata — fully synthetic, zero language. This improves language modeling by up to 6%, speeds up convergence by 40%, and strengthens downstream reasoning.
Surprisingly, it even beats pre-pre-training on natural text!
Blog: https://t.co/Pni0RsIcxL
(1/n)
Trying to interpret how a neural-network does what it does? Activations tell you if a neuron responded. Contributions tell you if a neuron mattered!
New paper from myself, @Zaki_Alaoui1, @sunnyliu1220 , @SuryaGanguli, and Steve Baccus: https://t.co/mcGZuj56AI
@HcwXd Excellent work! A suggestion: AI-generated mini-quizzes on fundamentals before starting to better tailor the lesson plan to the user's knowledge level. (It's more robust than guessing based on self-reported background knowledge)
1/4 LLMs solve research grade math problems but struggle with basic calculations. We bridge this gap by turning them to computers.
We built a computer INSIDE a transformer that can run programs for millions of steps in seconds solving even the hardest Sudokus with 100% accuracy
What’s the point of a “helpful assistant” if you have to always tell it what to do next?
In a new paper, we introduce a reasoning model that predicts what you’ll do next over long contexts (LongNAP 💤).
We trained it on 1,800 hours of computer use from 20 users.
🧵
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project.
This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.:
- It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work.
- It found that the Value Embeddings really like regularization and I wasn't applying any (oops).
- It found that my banded attention was too conservative (i forgot to tune it).
- It found that AdamW betas were all messed up.
- It tuned the weight decay schedule.
- It tuned the network initialization.
This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism.
https://t.co/WAz8aIztKT
All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.
And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
I packaged up the "autoresearch" project into a new self-contained minimal repo if people would like to play over the weekend. It's basically nanochat LLM training core stripped down to a single-GPU, one file version of ~630 lines of code, then:
- the human iterates on the prompt (.md)
- the AI agent iterates on the training code (.py)
The goal is to engineer your agents to make the fastest research progress indefinitely and without any of your own involvement. In the image, every dot is a complete LLM training run that lasts exactly 5 minutes. The agent works in an autonomous loop on a git feature branch and accumulates git commits to the training script as it finds better settings (of lower validation loss by the end) of the neural network architecture, the optimizer, all the hyperparameters, etc. You can imagine comparing the research progress of different prompts, different agents, etc.
https://t.co/YCvOwwjOzF
Part code, part sci-fi, and a pinch of psychosis :)
Perch 2.0 was evaluated on a range of whale vocalization tasks: distinguishing different baleen whale species, and killer whale subpopulations.
Compared against pre-trained models, it landed consistently in the top or second-best performing model for each dataset and sample size.
Interested in the latest advances in neuroscience (neural dynamics and internal models) and how they can be leveraged to build smarter, adaptive AI?
➡️ My first real solo piece 🖤🫶 @NatureNeuro
https://t.co/oa0Ky1qDZN
Our new paper, “End-to-End Test-Time Training for Long Context,” is a step towards continual learning in language models.
We introduce a new method that blurs the boundary between training and inference. At test-time, our model continues learning from given context using the same next-token prediction objective as training.
With this end-to-end objective, our model can efficiently compress substantial context into its weights and still use it effectively, unlocking extremely long context windows for complex reasoning and applications in agents and robotics.
Paper: https://t.co/tqPYECjFpn
Code: https://t.co/tADD7wYDAL
Another great short essay, by Paul Nurse, 2021:
"Theorizing should be encouraged, and theories should be included in experimental papers to put data in context."
Dive into the intricate connections in the mouse brain! 🧠✨
This video shows reconstructed neurons projecting from the thalamus to various regions of the cortex. Each neuron was fluorescently labeled and imaged with our cutting-edge ExA-SPIM microscope.
Self-wiring neural networks that learn structure without a single gradient
Reservoir computing is elegant: feed signals through a random recurrent network, then train only a simple linear readout. It's fast, cheap, and works surprisingly well—until you hit a wall. That static, random wiring was never designed for your task, and performance suffers.
Tanguy Cazalets and Joni Dambre take inspiration from neuroscience to fix this. Their method, Hebbian Architecture Generation (HAG), starts from an almost empty reservoir and progressively grows connections between neurons that frequently fire together—embodying the classic maxim "neurons that fire together wire together." No backpropagation, no gradient descent. Just unsupervised structural plasticity guided by long-horizon correlation statistics.
The results are striking. On speech and audio classification benchmarks, HAG-grown reservoirs consistently beat traditional Echo State Networks and popular plasticity rules like Intrinsic Plasticity or Anti-Oja learning. On smaller datasets (Japanese Vowels, CatsDogs), HAG matches or exceeds gradient-trained LSTMs and GRUs—while training orders of magnitude faster. The method slashes neuron-to-neuron correlation (from ~0.99 to ~0.47 on Speech Commands), expands effective dimensionality, and dramatically improves class separability metrics like silhouette scores and inter/intra-class distance ratios.
What's conceptually satisfying is why it works: by wiring together co-active units over extended time windows, HAG carves out a task-relevant subspace rather than blindly inflating dimensionality. It learns just enough structure to make classes linearly separable—nothing more, nothing less.
The broader message: structural self-organization, long studied in biological neural circuits, is a practical route to adaptive machine learning. HAG bridges the efficiency of reservoir computing with the flexibility of learned architectures, all without propagating a single error signal through time. For neuromorphic hardware, physical reservoirs, and low-data regimes, that combination could matter a lot.
Paper: https://t.co/0RMyHdTtyK
I've been thinking about the "virtual cell" concept and wanted to write up a few thoughts. Specifically on how I think the prior experience in GWAS informs the most likely way these models will be useful.
https://t.co/2XgLmRxYB0
🧠🔄 The brain isn’t a one-way street.
We’ve long been taught that information flows in a fixed "bottom-up" hierarchy—from sensory to the executive areas. Our new preprint shows this isn't true.
The brain actually reverses its information flow when things get blurry. 🧵 (1/6)
Many people think of the genome as a string of "letters." The human genome, say, has 3.2 billion base pairs of DNA organized across 23 pairs of chromosomes.
But the genome is a 3D object. Genes located on entirely different chromosomes might be clustered together. Mutations in these "distant" genes can lead to disease in surprising ways.
For a new paper in @Nature, researchers released several "maps" of human genomes from two types of cells: embryonic stem cells and fibroblasts. They compared methods to see which ones are least biased, and found many long-range interactions between genes.
The article does a good job explaining how “the genome is organized at different scales”:
> On a single chromosome, histones control which parts of the DNA sequence are accessible and expressed.
> At the scale of hundreds of thousands of bases, “chromatin loops in a dynamic manner,” the authors write, bringing distant genes closer together. > Across chromosomes, sequences "cluster together in space to form subnuclear compartments."
Examples abound. Enhancers, for example, are short DNA sequences that regulate the expression of far away genes. They do this by *physically* touching the genes they control; a protein called cohesin grabs the DNA and tugs it into big loops.
Even promoters, which are thought of as being associated with one gene or operon, can cluster together across many genes! A protein, Ronin, grabs promoters and pulls them together. This is apparently done mostly for genes that tend to be "on," as it helps enzymes find genes faster/not have to diffuse far away to find targets. (This also happens with genes that tend to be "off;" so-called polycomb proteins grab onto promoters, cluster them up, and silence all of them at once. It's a way for the cell to conserve energy.)
One consequence of this spooky "action-at-a-distance" is that diseases might arise from mutations in unexpected locations. Editing these regulatory sequences, in other words, might in turn affect a gene located on an entirely different chromosome that *is* associated with that disease.
Genetic mutations linked to autism, for example, are known to disrupt the 3D organization of the genome. A single deletion at a gene, TAL1, also affects its ability to form long-range chromatin interactions with other genes, leading to leukemia. There are probably many other, as-yet-undiscovered, instances of this.