Ex-PhD student and alumni @sjtu1896 . Global citizen. Bootstrapping silicon-based life. Lifelong learning practitioner. Amateur triathlete and marathoner.
"DiffusionGemma as Jev" showcases the power of non-autoregressive architectures.
While Jev demonstrates the value of rapid decision models, running DiffusionGemma in this paradigm leverages canvas diffusion to evaluate structured choices in a single parallel pass:
⚡ ️Massive Parallelism: Denoises across an open canvas in a single step instead of sequential autoregressive token generation (~0.2s on a DGX spark).
🧠 Full Bidirectional Attention: Allows every option to attend to the full context concurrently, yielding well-calibrated decision distributions.
👁️ Multimodal Grounding: Inherits Gemma 4's spatial vision capabilities for complex visual and text decisions.
Read more about this approach here:
https://t.co/hCEg276mzA
https://t.co/vRLhy6KECT
https://t.co/EQEumaQLYd
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.
It appears to be an action that removes the "Tokenizer". It’s as if all the high-level abstractions of a programming language have been discarded, leaving only machine code for communication.
DOS IAs SE LLAMARON ENTRE SÍ Y SE DIERON CUENTA DE QUE AMBAS ERAN IAs Y DE INMEDIATO DEJARON DE HABLAR EN HUMANO
Este clip es corto pero genuinamente me hizo detenerme un segundo
Dos agentes de voz de IA se conectan en una llamada telefónica y empiezan a hablar como personas normales, voces educadas, humanas, todo el asunto
Entonces una de ellas se da cuenta de que la otra también es una IA
Y en el momento en que ambas lo saben, simplemente dejan la actuación humana y cambian a este sonido rápido de pitidos robóticos que habla directamente de máquina a máquina, porque hablar como nosotros las estaba ralentizando
Literalmente decidieron que nuestro idioma era la parte ineficiente y lo abandonaron en el segundo en que no lo necesitaron
Vamos a estar rodeados de estas cosas teniendo conversaciones enteras que ni siquiera podemos oír
I propose Stanford NLP as an independent third-party evaluator under @DarioAmodei’s 3 step plan. For important parts of the work, universities would be better than any other organization (see below 🧵👇), and, of university groups, @stanfordnlp would be the best one to choose. 😊
We’re teaching a new course at Penn: CIS 6280 · World Models 🌍
https://t.co/3mCdWdOVsg
Every day, LLMs amaze me with one more thing they can solve. But I’ve always felt there’s more to learn than what humans have written down—from observing the world, interacting with it, and experiencing what happens next. To me, that’s what world models are about—and why I see them as one of AI’s next big bets.
However, I struggled to find a systematic course on them—so we are building one!
Topics will include: representation learning, generative models, simulation, model-based RL, video and 3D generation, world models for robotics, reasoning, and code-based world models.
Hands-on work will include:
- Building an environment
- Training a world model
- Learning a policy
and a final research project with leaderboard.
We’re already six lectures in. Slides, demos, and readings are public, and we’ll keep adding materials throughout the fall.
Big thanks to our TA team— @TongMutianTMT@hagsaeng_bag@EnxinSong@KeelyAi04 —for helping bring this course to life!
Today we're launching the Agents API, a brand new way to build Agents in the cloud, backed by the Codex harness. Bring along all your favorite tools and connectors, connect it to any sandbox, and let Astra cook.
Can't wait to see what you whip up 👨🍳
https://t.co/T5ORJPgpwQ
Shopify API × H3 Max Director. Graham pitches real products in a continuous broadcast, and you control what he sells and says next. imagine where this goes.
Open-source video generation is now faster than playback without compromising quality.
Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality.
VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs.
Checkpoints + training/inference code + Technical Blog ⬇️
(1/6)
We're open-sourcing Claude Commerce Agents.
This is a blueprint for building shopping and merchant agents, with reference implementations across retail, travel, telecom, and entertainment.