I’ll be at #COLM2026 this week in SF, excited to catch up with old friends and meet new ones!
I’m hiring 1–2 interns in Palo Alto to build next-gen Hy models (Hy5 or ?). If you’re interested in pretraining, RL, and infra for unified and interactive models, I’d love to chat.
Come say hi or DM me!
@viktaur27 @Teslarati The rate of improvement from original GPT to GPT-3 is impressive. If this rate of improvement continues, GPT-5 or 6 could be indistinguishable from the smartest humans. Just my opinion, not an endorsement. I left OpenAI 2 to 3 years ago. Am a neutral outsider at this point.
Where does money invested into the AI buildout actually go?
For every $100 flowing into the supply chain:
- $50 to chips
- $20 to power
- $15 to networking
- $15 to cooling, buildings, and land
More charts in State of Markets II: https://t.co/MTaxKUxa2w
Prefill vs decode in inference engines
With continuous batching we don’t need to strictly separate these two phases
But it is important our scheduler knows how many tokens are being processed
Decode (even with spec dec) is only a few tokens for a forward pass
Prefill (especially cold prefill) is easily thousands to hundreds of thousands
So we chunk prefill to allow the scheduler to better mix in parts of one requests prefill with other requests that are decoding
Without this abstraction our TTFTs and queue time would be very bad, and our system throughput would likely degrade as well
New Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore.
Scientific discovery begins where known problems end: a system has to frame hypotheses, design experiments, and learn from the results. Evaluating this is hard. Genuinely new answers cannot be checked quickly, and in familiar domains a model can simply recall what it has seen.
Addressing this challenge, researchers from Tencent Hy, Fudan University, and Tsinghua University built verifiable Alien Worlds. Their rules are executable, so every answer is checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks.
🔹 Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems)
🔹 A flawed manual, four rounds of self-designed probes, and closed-book tests after every round
🔹 Every answer graded by an interpreter or a proof checker, with no LLM judge
What we found across 10 frontier AI systems:
1️⃣ Getting feedback is more effective than thinking alone. No AlienCode run starts above 15.7%; after four rounds the best reaches 89.0%, while the same turns without feedback stay at 0.5–11.0%.
2️⃣ Designing the experiments matters. Replaying a system's own best probes gives it exactly the same evidence, yet in AlienCode 9 of 10 systems do worse than when they chose the probes themselves.
3️⃣ Knowing a rule is not using it. Even when every required rule is stated correctly, tasks are solved only 73.4% of the time.
4️⃣ One score hides a lot. The same system under the same budget ended anywhere from 5.7% to 79.0%, and rankings barely transfer between the two worlds.
CL-bench asked whether models can learn from context. ExplorationBench asks whether they can discover the rules themselves.
📄 Paper: https://t.co/eamwrsJyxQ
🌐 Website & leaderboard: https://t.co/gxSeoCfc2s
📝 Blog: https://t.co/jyrwH7ZnQH
💻 Code (coming soon): https://t.co/i7VlgErkcl
🚀 We let Jev play Minecraft. And we just can't beat it! 😭
⚡ Jev: 24 ms decisions
🧠 You: ~200 ms reactions
It moves before you even see it move. Way too strong!
🎮 https://t.co/HoEPFlNDep (join the server to win)
🚀 More reliable agents with DeepSeek V4.1!
XGrammar brings strict tool calling to SGLang & vLLM through Structural Tags, enforcing tool argument schemas in DeepSeek’s native format.
See how the Structural Tag works 👇
https://t.co/N0Tbl58Grf
Check out XGrammar 👇
https://t.co/GoOOhg1FLw
ZCode is now open source, and the reported security issues have been addressed.
https://t.co/Bq5B6nEnWM
We take the community’s feedback very seriously and apologize for the concern and frustration these issues have caused. We have been working closely with the ZCode team to investigate and remediate the issues raised.
Independent security reviews by third-party firms are now underway. We will share the findings and provide further updates as they become available.
We sincerely appreciate the developer community’s feedback and continued scrutiny. We will continue to monitor the review process closely and provide updates.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju