Fable 5 probably running locally in about two years.
That is the projection in this r/LocalLLaMA chart. It tracks how long it takes for cloud-frontier capability to become broadly comparable in laptop-runnable open-weight models. The observed average lag: ~24.8 months.
GPT-3-class capability: 37 months.
GPT-3.5-class: 17 months.
GPT-4-class: ~24 months.
The projection puts Fable / Mythos 5-class capability on high-end consumer hardware around July 2028.
🚨 Holy shit... LeCun's team just cracked world models wide open.
Everyone's obsessing over the next Claude update.
Meanwhile Yann LeCun quietly dropped a paper that could matter way more long term.
It's called LeWorldModel.
And to understand why it's a big deal, you need to understand the difference between what LLM does and what this does.
LLMs predict the next word. That's it.
They're incredibly good at language. But they don't understand reality.
They can write about a ball bouncing off a wall. They can't predict where it lands.
World models predict what happens next in the physical world. Objects moving, colliding, falling.
That's the foundation for robots that plan, self-driving cars that simulate scenarios, any AI that needs to act in reality instead of just talk about it.
The problem? World models kept collapsing.
The model would cheat by mapping every input to the same output. Like a weather app that predicts "sunny" every single day.
Technically it's predicting. It's just useless. And fixing this required 6+ loss hyperparameters, frozen pre-trained encoders, stop-gradient hacks, exponential moving averages.
A house of cards just to keep the thing from breaking.
LeCun's team (Mila, NYU, Samsung SAIL, Brown) threw all of that out. LeWorldModel uses just 2 loss terms.
A prediction loss and a regularizer called SIGReg that forces representations to stay diverse instead of collapsing into garbage.
6 hyperparameters reduced to 1.
The simplicity IS the breakthrough.
The numbers: 15M parameters. Trains on a single GPU in a few hours. Plans up to 48x faster than foundation-model-based world models.
Uses roughly 200x fewer tokens than alternatives. Competitive across 2D and 3D control tasks.
This isn't a supercomputer experiment. You could run this on your own hardware.
LeCun has been pushing JEPA as the architecture for real AI since 2022.
The criticism was always the same: "sounds nice, doesn't train stably."
LeWorldModel just removed that objection. Small model. Stable training.
No hacks. No frozen encoders. No collapse.
Two AI futures are competing right now.
Path 1: bigger LLMs, more text, more compute.
Path 2: world models that learn physics from raw pixels and plan in real time.
LeWorldModel is the strongest signal yet that Path 2 is real, getting cheaper, and closing in fast.
@karpathy@Wikipedia shares this qualitative list of writing and formatting conventions that are indicative of AI writing, but not quantitative per se - https://t.co/GL0cUmXGmc
Readers responded with both surprise and agreement last week when I wrote that the single biggest predictor of how rapidly a team makes progress building an AI agent lay in their ability to drive a disciplined process for evals (measuring the system’s performance) and error analysis (identifying the causes of errors). It’s tempting to shortcut these processes and to quickly attempt fixes to mistakes rather than slowing down to identify the root causes. But evals and error analysis can lead to much faster progress. In this first of a two-part letter, I’ll share some best practices for finding and addressing issues in agentic systems.
Even though error analysis has long been an important part of building supervised learning systems, it is still underappreciated compared to, say, using the latest and buzziest tools. Identifying the root causes of particular kinds of errors might seem “boring,” but it pays off! If you are not yet persuaded that error analysis is important, permit me to point out:
- To master a composition on a musical instrument, you don’t only play the same piece from start to end. Instead, you identify where you’re stumbling and practice those parts more.
- To be healthy, you don’t just build your diet around the latest nutrition fads. You also ask your doctor about your bloodwork to see if anything is amiss. (I did this last month and am happy to report I’m in good health! 😃)
- To improve your sports team’s performance, you don’t just practice trick shots. Instead, you review game films to spot gaps and then address them.
To improve your agentic AI system, don’t just stack up the latest buzzy techniques that just went viral on social media (though I find it fun to experiment with buzzy AI techniques as much as the next person!). Instead, use error analysis to figure out where it’s falling short, and focus on that.
Before analyzing errors, we first have to decide what is an error. So the first step is to put in evals. I’ll focus on that for the remainder of this letter and discuss error analysis next week.
If you are using supervised learning to train a binary classifier, the number of ways the algorithm could make a mistake is limited. It could output 0 instead of 1, or vice versa. There is also a handful of standard metrics like accuracy, precision, recall, F1, ROC, etc. that apply to many problems. So as long as you know the test distribution, evals are relatively straightforward, and much of the work of error analysis lies in identifying what types of input an algorithm fails on, which also leads to data-centric AI techniques for acquiring more data to augment the algorithm in areas where it’s weak.
With generative AI, a lot of intuitions from evals and error analysis of supervised learning carry over — history doesn’t repeat itself, but it rhymes — and developers who are already familiar with machine learning and deep learning often adapt to generative AI faster than people who are starting from scratch. But one new challenge is that the space of outputs is much richer, so there are many more ways an algorithm’s output might be wrong.
Take the example of automated processing of financial invoices where we use an agentic workflow to populate a financial database with information from received invoices. Will the algorithm incorrectly extract the invoice due date? Or the final amount? Or mistake the payer address for the biller address? Or get the financial currency wrong? Or make the wrong API call so the verification process fails? Because the output space is much larger, the number of failure modes is also much larger.
Rather than defining an error metric ahead of time, it is therefore typically more effective to first quickly build a prototype, then manually examine a handful of agent outputs to see where it performs well and where it stumbles. This allows you to focus on building datasets and error metrics — sometimes objective metrics implemented in code, and sometimes subjective metrics using LLM-as-judge — to check the system’s performance in the dimensions you are most concerned about. In supervised learning, we sometimes tune the error metric to better reflect what humans care about. With agentic workflows, I find tuning evals to be even more iterative, with more frequent tweaks to the evals to capture the wider range of things that can go wrong.
I discuss this and other best practices in detail in Module 4 of the Agentic AI course on https://t.co/zGHUh1loPO that we announced last week. After building evals, you now have a measurement of your system’s performance, which provides a foundation for trying different modifications to your agent, as you can now measure what makes a difference. The next step is then to perform error analysis to pinpoint what changes to focus your development efforts on. I’ll discuss this further next week.
[Original text: https://t.co/hZyBupYIgz ]
Introducing Alterego: the world’s first near-telepathic wearable that enables silent communication at the speed of thought.
Alterego makes AI an extension of the human mind.
We’ve made several breakthroughs since our work started at MIT.
We’re announcing those today.
The UK is being hit by a wave of industrial closures from Port Talbot to Scunthorpe to Grangemouth where energy policy is a major factor.
I have two longreads out today (link in bio) on what's gone wrong. Quick thread:
I’m fundraising for the The Alzheimer Society of Ireland (@alzheimersocirl) charity by walking '100k in a day' starting tonight.
If you could donate any small amount using the below JustGiving page that would be much appreciated - https://t.co/9jjB23x3Ff
Have seen 2x RTAs at the junction of Cotswold/Monksdale due to how blind this junction is with the new parking. However the major concern is the blind "crossing" from the cycle path to the playground; unfortunately it seems inevitable that a child will eventually be knocked down.
Congratulations to @bathnes & @jessd4moorlands on the successful completion of the new Cotswold Road carpark! Such a lovely safe space for kids on the daily school run. Thanks for the careful consideration and thoughtfulness you've shown to your constituents! cc @bathlive
This was preceded by 2 men going head to head in the middle of the road for 5 minutes due to the chaotic bottleneck the local parking changes have introduced on Cotswold Road.
Amazing insight from @bernardjackman
Up to date. Insightful and doing the hardest bit, delivering it on tv live without, missing a beat. One of the very best pundits in the world. @RTEsport
@MindCharity Update 3/3: The fundraising page is also now updated with my write-up of the challenge - https://t.co/2Ro1HFA7Jd
(or if you’re on Strava go here to read it - https://t.co/VgXjj3OLwh)
I’m fundraising for the mental health charity @MindCharity by walking '100k in a day'. Below is my JustGiving page to learn more.
If you could donate a 5er that would be much appreciated and goes towards a fantastic cause: https://t.co/anbCr3CmP3
@MindCharity Update 2/3: So far I’ve raised 706% of my original fundraising target - but there’s still time to donate to @MindCharity - the fantastic mental health charity to whom your donation will go.
@MindCharity Update 1/3: I completed the challenge to walk ‘100k in a day’ on Saturday afternoon after 19.5 hours. The Strava activity is here: https://t.co/uV1IrpFIoX