Today, we open sourced Mojo 🔥.
Announced just now during the ModCon keynote, effective immediately, Apache 2.0 License.
Thank you to our community for waiting patiently and building alongside us.
#ModCon2026
Full blog: https://t.co/y5cSUsohrs
Yes indeed. If you're confused by this, you should be. Those things are not all the same. I certainly still am confused. But, you know, natural language. And humans. Always getting in the way of understanding.
Computer Science decided to use Kernels for everything
>Kernel (GPU Programming): A parallel-executed function on a GPU device.
>Kernel (OS): core OS component managing hardware and processes.
>Kernel (ML): Similarity measure in transformed feature space.
>Kernel (Computer Vision): small matrix for image convolution filters.
>Kernel (Linear Algebra): Set of vectors mapping to zero.
>Kernel (Functional Analysis): function defining inner products in spaces.
>Kernel (Statistics): Weighting function for density estimation.
>Kernel (NLP): Similarity function for structured data comparison.
>Kernel (Cryptography): A keyspace/function concept in some cryptographic constructions.
>Kernel (Compilers): Minimal/core representation or execution component in some compiler architectures.
AI competence has always been very spiky, superhuman in some narrow domains and largely useless in others. The fundamental marketing trick of the AI industry is to make you believe the tallest spike is a floor.
Real brains follow Dale's principle: a neuron can either excite its neighbors or suppress them, but never both. Standard deep learning ignores this and uses backpropagation.
In our new paper, Diffusing Blame, we fix this disconnect. By introducing a routing method that broadcasts error signals directly to the hidden layers, we can train networks made of dedicated positive and negative neurons to strictly obey Dale's principle, all without relying on backprop!
This method works surprisingly well on image recognition tasks despite the strict biological constraints. We also achieved competitive, backprop-free reinforcement learning on complex locomotion tasks and the open-ended Craftax environment.
It is neat to see that representation learning remains possible even when we force deep learning to play by the rules of real neurons.
We scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot policies used to live their lives a few frames at a time (< 0.1 sec), instantly forgetting what just happened. We pushed to 3 orders of magnitude beyond SOTA.
Introducing RoboTTT. Test-Time Training (“TTT”) carries a tiny model *inside* the model. Every incoming sensor reading triggers one gradient step on that tiny core, so the history keeps getting compressed into its weights. The hidden state has a fixed size (literally a small neural net), so the robot can “grok” arbitrarily long experience with little overhead. Learning continues indefinitely after deployment.
We can then put an entire video in context as prompt! RoboTTT enables one-shot in-context learning from human video: in circuit board assembly, a human demonstrates a never-seen configuration once, and the robot imitates it faithfully.
Humans drop things all the time, but we pick them up so fast that we don’t even notice. That reflex to fix is half of our physical competence. RoboTTT shows self-improvement on the fly: the robot is skilled at recovering from its own errors mid-episode, and each fix enters its context to inform the next move. The TTT core distills a general-purpose, failure-to-correction mapping from the training data.
One more thing. What excites me the most is a new Context Scaling Curve: from 128 to 8K timesteps, closed-loop performance hill-climbs steadily with no sign of saturation. 8K-context pretraining beats 1K by 62%. What LLM enjoys, robotics should too. Soon, even 1M context is not a fantasy.
Deep dive in thread:
If your benchmark relies on a static dataset or sampling from a static distribution densely known at training time, then it is fundamentally measuring memorization/retrieval. Which might be fine if you're looking for a retrieval benchmark! But don't confuse it with intelligence.
The true measure of a software engineer isn't their ability to write clever code. It's their ability to ruthlessly protect the codebase from unnecessary cleverness.
This is a new paradigm for interacting with Claude that is significantly more "inline" with all the other human activity org-wide. Once you do all of the under the hood engineering work to make this "just work" (e.g. across tools, integrations, compute environments, memory, security, etc.), Claude basically joins the team in a seamless way - you can talk to it as you would talk to a person and it can help with a very large variety of workloads.
Imo this is the 3rd major redesign of LLM UIUX. The first paradigm was that the LLM is a website you go to, the second was that it is an app you download to your computer. This third one is that it is a self-contained, persistent, asynchronous entity with org-wide tools and context, working alongside teams of humans. It really takes a while to wrap your head around it, but it works and it is awesome.
Programming is not about code, just like music is not about notation. It is the art & science of managing complexity through layers of abstraction. AI is simply a part of it.
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees.
The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance.
Access to all other Claude models is not affected.
We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible.
Read our full statement: https://t.co/bwn0sximKZ
So-called “neural networks” in AI have hardly anything to do with actual neural networks.
Excellent new study shows how enormous the the gulf is between the simplified neurons in typical AI systems and actual neurons.
Some considerations that many folks seem not to get:
1. It can be a bubble even if the tech works. (For instance, if the tech doesn't have a high-demand use case.)
2. It can be a bubble even if the tech works and has strong product-market fit. (For instance, if the tech cannot be economically viable.)
3. It can be a bubble even if the tech works, has strong product-market fit, and has a path to eventual economic viability. (For instance, if profitability takes too long to achieve or makes margin/competition assumptions that fail to materialize.)
4. It can be a bubble even if the tech works, has strong product-market fit, and is currently highly profitable. (For instance, if demand has a hard ceiling and growth stops once the ceiling is reached.)
5. It can be a bubble even if the tech works, has strong product-market fit, is currently highly profitable, and has unlimited future demand.
Literally all it takes for something to be a bubble is for lots of people to over-enthusiastically bet their money on it, and subsequently get panicky.
Importantly, bubbles can be attached both to things that are completely hogwash, like the Metaverse, and to world-changing developments like the Internet or railways. Bubbles don't care. They're brought into existence by the thoughts and feelings of investors, not by actual tech or products.
"The bubble has burst" doesn't mean "the tech didn't work" or "people stopped using the tech." It only means that people got panicky, investor money dried up, and valuations collapsed. Internet adoption didn't stop in 2000.
wondering why I feel exhausted. maybe: the agents do all the easy stuff, and I have to work through the leftover hard bits, which means I'm perpetually locked in. and as the models get better, "my" work just gets harder and harder, until I'm basically underqualified to do the work (which... is better than the alternative, there's nothing left for me to do, and I'm paperclipped).
Local artist, Light Guerrilla, projected TAX THE RICH on Mark Zuckerberg’s $300M mega yacht. It is docked at Seattle’s Lake Union and the company is laying off 1395 employees in Washington state starting next month
https://t.co/EYmpueOCk8
Instead of forcing the entire model into DRAM, the full model is stored in flash memory (NAND). Because NAND-to-DRAM bandwidth is too slow to swap weights token by token, as standard MoE models require, AFM 3 Core Advanced makes routing decisions per prompt.
.....
The model periodically reselects and updates the activated experts during the token generation process.
And another open-weight release. Nemotron 3 Ultra has an ultra impressive capability:efficiency ratio!
Design-wise, it carries forward the Mamba-2-attention hybrid stack and LatentMoE introduced in the previous Super variant. But everything is a bit bigger.