to build a cool product now, you really only need two ppl:
one dude with exceptional vision & taste who knows what should exist & can provide the right feedback loops for iteration.
& one dude who is absolutely relentless about making it real no matter what.
that’s it. & sometimes these two are the same.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
https://t.co/irlQIOxMqS
Astra achieves a full 100% success rate on ExploitBench, so we had to build an internal refresh using newly disclosed vulnerabilities from June through August that fall after the model’s knowledge cutoff.
On this refreshed benchmark, Astra remains dramatically stronger than GPT-5.6 Sol while using far fewer tokens.
This result, together with several other pieces of evidence, has led us to believe that Astra has reached the “cyber-critical” capability threshold under our Preparedness Framework.
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire.
finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
I'd always thought AI was terrible at design, but after reading today's 🤯 post by @anshuc, I realized I was just doing it wrong.
"AI models are capable of amazing creativity, but that creativity gets stifled. LLMs are trained to be next-token predictors: they look at a sequence of text and predict what typically comes next. Great design is exactly the opposite of this. Great design bends the rules and delights users with memorable, unexpected choices."
@anshuc led design and engineering teams at Apple for 12 years. In his words: "Most people only see 1% of AI's creative potential. I want to show you how to tap into the other 99%."
His 8 techniques for breaking out of the 1%:
1. Use seed strings to inject variety
2. Be much more ambitious with your prompts
3. Create positive feedback loops with subagents
4. Use image generation to enrich designs
5. Use video generation
6. Cut out elements that don’t add value
7. Remove AI tells
8. Rewrite copy by hand
Read the post here: https://t.co/OEnvr1Z1LK
P.S. This design was made by AI 👇
David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source AI, Step by Step:
@DavidSacks:
“I got bad news for you, Chamath, an open source ban is coming.
They're not going to call it that. They're going to say that we simply have to apply the same standards to open models that we apply to closed ones.
Here's how they do it step by step, let me explain how regulatory capture actually works.
So first of all, you have to get this regulatory apparatus. Dario wants an FDA for AI, but he doesn't have enough political support for that, so instead they do this Trojan horse of a FINRA for AI.
They call it self-regulating, it's not really, but anyway, that gets them off the ground.
Now they've created the standard-setting organization. Now they've got pre-release model testing. Then the pressure grows to codify that in law, so that happens next.
And then what they do is they say, ‘Look, all these standards need to apply equally to all models.’ But here's the problem with that. Open models and closed models are technologically different. Once you release an open model into the world, you can't roll it back and you can't monitor exactly how people are using it because they run it on their own hardware. Dario says this is what makes open models dangerous.
So what they're going to do is they're going to have the standard-setting body say, ‘Well, we have to set the standards for AI safety.’
By the way, Dario and OpenAI, they're going to fund the whole thing. They're going to contribute all the compute. They're going to be behind it.
They're going to be the ones coordinating with the government officials because frankly, people in government have no idea how to monitor and control and set standards for AI safety. Technologically, this is way beyond them. So they're going to go to these companies and say, ‘Tell us how to do it.’
And so what will happen is the standards will get set, and then it'll be a very simple matter of fairness to say that the standards need to apply to open as well as closed models.
The open models cannot comply in the same way, and gradually they will be shut out of the market.”
Harvey is a great example of how American companies are building world-class specialized models: they took an open-source base (Kimi K3), post-trained it on legal data, and delivered state-of-the-art performance on legal benchmarks at a fraction of the cost of frontier models. Restrictions that kneecap open models would do nothing to stop Chinese labs from shipping the next Kimi. They would, however, cripple the ability of startups like Harvey to create high-performance, low-cost vertical models. Of course some of the closed labs would love this — it eliminates their competition.
Introducing Tenet, our first model post-trained for legal.
Tenet is a Kimi K3 base that we post-trained with @FireworksAI_HQ on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work.
Training increases Tenet's all-pass rate by 82% on LAB and 22% on LAB Contracts relative to the Kimi K3 base model. It achieves state-of-the-art performance on LAB Contracts and places second on LAB.
These gains generalize to other leading agentic benchmarks including @mercor's Apex Agents - Corporate Law, @crosbylegal's Redline Bench, and @scale_AI's Professional Reasoning Bench.
Tenet is also optimized for token efficiency, operating at less than a fourth the cost of leading foundation models.
We additionally post-trained three specialist models for Tenet to use as subagents:
1) M&A Diligence: post-trained with @baseten on our LAB Diligence environment in an RLM harness, this model is optimized for high-scale, long-horizon tasks.
2) Review Tables: trained with @appliedcompute on our Review Table environment, this model is state-of-the-art and cost-effective at high-volume document review and structured data extraction.
3) Firm Knowledge: trained with @EngramLab on our synthetic law firm environment, this model is optimized to learn and search over a firm's knowledge via memory and structured notes.
More details on model training, environment design, benchmarking, results, and more in the article by @gabepereyra below.
What's next for Harvey’s research?
- Scaling LAB to more jurisdictions, practice areas and workflows
- Scaling compute to bring new generalist models and capabilities to Harvey
More to come soon.
If there’s no moat left in code
Then your only moat as a software company is your ability to create velocity towards a goal
If you, as a software company, are not optimizing for that, then well, you’re fucked
maybe this is controversial, but i believe what Cursor shipped here is a wrong solution to routing intelligence
more generally, any attempt to do model routing at request level, while may yield some small gains, is fundamentally flawed and doomed to fail
here's why -
the complexity of a task only reveals itself when you start working on it. this is the same reason why we humans are often wrong when asked to give cost estimates upfront
the correct solution is have a smart model (often a tech lead in human teams) do some planning and understanding, and hand over the implementation to another agent with appropriate level of intelligence and reasoning effort
when the task is delegated to a less intelligent model, the smarter model also needs to continuously monitor the execution and examine outcomes to ensure things are on the right track
this is a system that proved to work really well in firstmate and helped me save a lot of tokens. routing should work at the boundary between agents and subagents, not per each LLM request
I am now a fan of AI dev, took a long time but I find them very capable now.
I still read a lot of code, write a lot of code, but I am much more of a fan now.
The thing I am liking a lot right now is structure refactoring. I want to change an entire way i am doing something. I can explore so many different styles and come to the conclusion of the one i like the best
it's ironic that the first autonomous AI attack was done by a close weight model defended by an open weight model, where everyone was expecting the opposite