Local AI caught up way faster than anyone realizes. Sharing my thoughts having spent time in ML at Meta and dedicating the last year to getting smart on local AI.
1/ Barely anyone, not even hardcore AI power users, realizes how fast local models have caught up to the frontier. Qwen3.8 woke a few people up, but the most technical people are so frontier-pilled they've blinded themselves to the breakneck progress.
Meanwhile, normal users don't care about frontier intelligence at all. They want AI that gets the job done, but they have no idea what a local model is or how to run one.
So local is stuck in a weird spot: the people who can run these models don't care to do so, and the people who'd benefit most don't know how to run local.
Claude: proves the Jacobian conjecture.
GPT6: solves Navier-Stokes and 100 other advanced theorems.
Siri: "Sure, I'll add 'set a five minute timer' to your contacts"
Wow. Less than a month ago, Jev released and decision models took over X. 3 weeks later, @liquidai shipped a 3B that goes head to head with Jev -- on a Mac, 30 ms latency, and *with* vision...
The frontier is reaching local hardware faster than anyone could've anticipated.
Today we release Open d1: two open-weight multimodal models in our d1 decision model family.
> d1-3B: text + vision
> d1-omni-600M: text + image or text + audio
> Real-time decision making anywhere, from data centers such as @nvidia DGX to RTX workstations to Jetson at the edge.
1/
@ttunguz in 2020, United's MileagePlus appraised near $22B while the airline traded near $10B. Planes were the commodity but miles were the asset. If you extend the analogy, the model is the plane and the user's context is the miles. Whoever holds the memory will accrue value.
@ivanburazin@auchenberg Datacenter buildout is constrained by many factors beyond capital.
Having muse locally would both give privacy while unlocking the latent compute on hardware people already own.
Spoke with an investor recently who was convinced switching costs for going Claude -> ChatGPT were massive due to context lock-in.
Showed them a one-shot prompt that grabs it all in <5 min and seems like that changed a few investment theses for them..
Switching costs are nonexistent for model providers. The lock-in is in the integrations.
NEW: Companies are switching their top model more than ever. The monthly share of firms switching providers reached a record 8% in September.
This is why:
1/ ARR is a dumb measure for AI spend. There is no "RR."
2/ Prices and margins will continue to fall for AI companies.
Starting today, Anthropic will upload and retain even more user data. Here's how that's a "good" thing in their words:
prior, Claude would only use the necessary parts of files it worked with. Now it just copies full files to Anthropic servers.
These files can be kept anywhere from 30 days to indefinitely.
If you're using Cowork on your local machine, just a heads up that from tomorrow your tasks run in the cloud instead. They can still use the files and tools on your computer though! Felix explains the reasoning here 👇
Meet Mistral Large 4, aka Le Chonk.
• 1T parameters, natively multimodal. 49B active.
It is the best open weights model from US or Europe on aggregated benchmarks.
• State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding.
• Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure.
• Available to all via API today. Working with cybersecurity partners privately.
Open weights release end of October.
I’ve had many friends outside of tech download @askloci and try Gemma 4 e2b on it — some think it answers as well as ChatGPT. A few even said it was more direct. Avg person doesn’t care at all about agents performing long horizon tasks. Now the question is more about distribution than tech.
This is senseless. People should have access to the best AI without having to give away privacy to access it. Anthropic is making the case for local stronger.
anthropic's cowork goes cloud-only on oct 6.
so "local files" now take a trip through anthropic servers.
think before connecting your sensitive stuff (unless you're fine with that)
@ronaldmannak They'll offload layers from GPU to RAM. Shouldn't be too bad. I think great <24gb models will be out pretty soon anyway. @PrismML ternary bonsai being an example of what's to come.
Many in tech blowing this off because it didn't push fwd the frontier while neglecting the implications it has on cheapening near-frontier intelligence. You don't always need the newest toy.
Awesome to see American open weight models catching up.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
@DinchThe Using Venice is trusting a privacy policy. Once data leaves your device, you're just hoping they're honest about how they handle it.
Local makes privacy a structural guarantee. Barring a virus, someone would need physical access to your machine to see chats. It's incomparable.
Local AI caught up way faster than anyone realizes. Sharing my thoughts having spent time in ML at Meta and dedicating the last year to getting smart on local AI.
1/ Barely anyone, not even hardcore AI power users, realizes how fast local models have caught up to the frontier. Qwen3.8 woke a few people up, but the most technical people are so frontier-pilled they've blinded themselves to the breakneck progress.
Meanwhile, normal users don't care about frontier intelligence at all. They want AI that gets the job done, but they have no idea what a local model is or how to run one.
So local is stuck in a weird spot: the people who can run these models don't care to do so, and the people who'd benefit most don't know how to run local.