Proud of this one. We spent weeks chasing a question that sounds simple: why is the trainer waiting when inference is fast? The answer touched routing, sandboxes, scheduling and weight publication — all at once. Huge credit to @MingyeGao for the technical heavy lifting.
Training AI agents takes more than fast token generation.
Tools run. Tests stall. Training batches wait.
We cut fixed-weight batch collection time by 66.9% on a Terminal-Bench-based workload by improving routing, sandbox execution and scheduling.
Here’s how: https://t.co/Dv0UzCvzm3
Training AI agents takes more than fast token generation.
Tools run. Tests stall. Training batches wait.
We cut fixed-weight batch collection time by 66.9% on a Terminal-Bench-based workload by improving routing, sandbox execution and scheduling.
Here’s how: https://t.co/Dv0UzCvzm3
Today, we're announcing that Eigen AI is joining Nebius (NASDAQ: NBIS).
From day one, our mission has been Artificial Efficient Intelligence — building the world's most efficient engines for generating intelligence. Together with Nebius, we're working toward the best AI cloud, uniting Eigen's full-stack model and inference software, ranked #1 on Artificial Analysis for inference speed, with Nebius's global hardware and infrastructure footprint, so any developer or enterprise can run the best models at the best price, with no capacity ceiling.
After close, Eigen's optimization stack will be integrated directly into Nebius Token Factory. The entire Eigen AI team is joining Nebius in full, establishing Nebius's engineering and research presence in the San Francisco Bay Area.
To our customers, our team, our investors at Tectonic Ventures, E14 Fund, Uncorrelated Ventures, and AGI House Ventures, our angel investors, advisors, mentors, and supporters — and to the Nebius team for the conviction and partnership — thank you.
The mission doesn't change. The leverage behind it does.
Ryan Hanrui Wang, co-founder and CEO of Eigen AI, said:
“We’re proud to join Nebius and work alongside the Token Factory team to push the boundaries of inference performance. Nebius has built a world-class AI cloud with a deep engineering culture that perfectly aligns with our own. Together, we are removing the friction of AI model customization and deployment so developers can run models reliably in production without managing the underlying infrastructure.”
Full announcement at:
https://t.co/XEeMZH4QKv
Run multimodal agents faster on Eigen AI Model Platform with @nvidia Nemotron™ 3 Nano Omni—a single model for text, video and audio—NVFP4 optimized on NVIDIA Blackwell → 500+ tok/s/user, zero quality loss across multiple benchmarks vs. BF16 🔥🥇
Explore how Eigeninference brings production-ready performance from day one 👉 https://t.co/zEW3TXSDDD
Try it out at the Eigen AI Model Studio 👉 https://t.co/iMQ2KMOYCw
🚀🚀🚀We're excited to introduce Galileo 0 (https://t.co/rWeqZzMTx1) — our first research preview of a world critic for AI-generated video, which already outperforms Qwen 3.5-Plus, Gemini 3.1 Pro, Pegasus 1.2, and GPT 5.4 on physical consistency reasoning 🚀🚀🚀
Galileo doesn't just score outputs. It diagnoses them — identifying what failed, when it failed, where it happened, and why it broke the rules of the world.
This is a step toward a new paradigm: generate → critique → refine → repeat — where models don't just produce worlds, but learn to keep them consistent over time.
𝐖𝐡𝐚𝐭 𝐦𝐚𝐤𝐞𝐬 𝐭𝐡𝐢𝐬 𝐦𝐢𝐥𝐞𝐬𝐭𝐨𝐧𝐞 𝐞𝐯𝐞𝐧 𝐦𝐨��𝐞 𝐦𝐞𝐚𝐧𝐢𝐧𝐠𝐟𝐮𝐥:
We built Galileo 0 — along with our datasets (including our public Physion-Eval benchmark), evaluation pipeline, and early pilots — with less than $200K total spend in 3 months.
No massive training clusters. No billion-dollar budgets. Just a small, relentless team, strong conviction — and yes, at one point, a five-day stretch of not showering to get this model out.
Because we believe reliability will become core infrastructure for world models.
In a world where billions are being poured into generation, the missing piece isn't more pixels — it's better critics 😊
#PhysionLabs #Galileo0 #WorldModels
🚨 “𝐖𝐡𝐢𝐜𝐡 𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐞𝐝 𝐯𝐢𝐝𝐞𝐨 𝐥𝐨���𝐤𝐬 𝐛𝐞𝐭𝐭𝐞𝐫?” is the wrong question. And yet, that’s exactly what arena-style evaluation asks — and much of multimodal AI is still judged this way. The problem is that it captures visual preference, but fails to measure whether the scene is actually coherent — whether objects behave consistently, interactions make sense, or events follow causal structure.
The real challenge isn’t visual quality. It’s whether a model can produce outputs that are correctly grounded across space, time, objects, and interactions — in other words, 𝐝𝐞𝐭𝐚𝐢𝐥𝐞𝐝 𝐦𝐮𝐥𝐭𝐢𝐦𝐨𝐝𝐚𝐥 𝐠𝐫𝐨𝐮𝐧𝐝𝐢𝐧𝐠 𝐚𝐧𝐝 𝐫𝐞𝐚𝐬𝐨𝐧𝐢𝐧𝐠.
At 🚀 𝐏𝐡𝐲𝐬𝐢𝐨𝐧 𝐋𝐚𝐛𝐬 🚀, in collaboration with researchers from 𝐒𝐭𝐚𝐧𝐟𝐨𝐫𝐝, 𝐌𝐈𝐓, and 𝐇𝐚𝐫𝐯𝐚𝐫𝐝 -- including Peiyu Jing, Hong-Xing "Koven" Yu, Fangqiang Ding, Fan Nie, Weimin Wang, Yilun Du, James Zou, Jiajun Wu, and Bing Shuai -- we analyzed state-of-the-art video generation models. What we found is hard to ignore: 𝐚𝐜𝐫𝐨𝐬𝐬 𝐥𝐞𝐚𝐝𝐢𝐧𝐠 𝐯𝐢𝐝𝐞𝐨 𝐠𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐦𝐨𝐝𝐞𝐥𝐬, 𝟴𝟯.𝟯% 𝐨𝐟 𝐞𝐱𝐨𝐜𝐞𝐧𝐭𝐫𝐢𝐜 𝐯𝐢𝐝𝐞𝐨𝐬 𝐚𝐧𝐝 𝟵𝟯.𝟱% 𝐨𝐟 𝐞𝐠𝐨𝐜𝐞𝐧𝐭𝐫𝐢𝐜 𝐯𝐢𝐝𝐞𝐨𝐬 𝐜𝐨𝐧𝐭𝐚𝐢𝐧 𝐩𝐡𝐲𝐬𝐢𝐜𝐚𝐥 𝐢𝐧𝐜𝐨𝐧𝐬𝐢𝐬𝐭𝐞𝐧𝐜𝐢𝐞𝐬. These are not just visual artifacts, but failures in object interactions, temporal continuity, and causal structure. Many are subtle, but fundamentally wrong.
This reveals a critical gap. We’ve made massive progress in making videos look better, but far less progress in making them actually grounded and consistent. The uncomfortable truth is that “looks right” does not mean “is right,” and preference does not imply understanding.
We’re releasing 🎬 𝐏𝐇𝐘𝐒𝐈𝐎𝐍-𝐄𝐕𝐀𝐋, the first human-centered benchmark for physical realism in AI-generated video. It includes over 10,000 expert reasoning traces, spans 22 fine-grained physical phenomena, provides temporally grounded annotations, and enables direct comparison between human and model reasoning.
📄 Paper: https://t.co/uIwSYBHva1
🤗 Dataset: https://t.co/QDeWxi26gU
🖼️ Preview: https://t.co/fUUrWj5XZD
𝐈𝐟 𝐰𝐞 𝐤𝐞𝐞𝐩 𝐨𝐩𝐭𝐢𝐦𝐢𝐳𝐢𝐧𝐠 𝐟𝐨𝐫 𝐚𝐩𝐩𝐞𝐚𝐫𝐚𝐧𝐜𝐞, 𝐰𝐞’𝐥𝐥 𝐠𝐞𝐭 𝐦𝐨𝐫𝐞 𝐜𝐨𝐧𝐯𝐢𝐧𝐜𝐢𝐧𝐠 𝐢𝐥𝐥𝐮𝐬𝐢𝐨𝐧𝐬 — 𝐧𝐨𝐭 𝐦𝐨𝐫𝐞 𝐫𝐞𝐥𝐢𝐚𝐛𝐥𝐞 𝐬𝐲𝐬𝐭𝐞𝐦𝐬. And for world models, robotics, and real-world deployment, that’s a fundamental failure.
We’re open-sourcing the dataset and releasing the paper today. This is a step toward a new standard: not just generating what looks good, but generating what is actually 𝐠𝐫𝐨𝐮𝐧𝐝𝐞𝐝, 𝐜𝐨𝐧𝐬𝐢𝐬𝐭𝐞𝐧𝐭, 𝐚𝐧𝐝 𝐜𝐨𝐫𝐫𝐞𝐜𝐭.
#AI #VideoGeneration #MultimodalAI #AIEvaluation #AIBenchmark #WorldModels #DeepLearning 🚀🐶
The future of AI is open -- but it also needs to be fast, efficient, reliable, and production-ready.
Excited to partner with @NebiusAI to bring optimized frontier open models to Token Factory. 🚀
Together, we’re helping developers and enterprises run leading open-source models in production with greater speed, reliability, and scale by combining Eigen AI’s deep inference optimization with Nebius’s production-grade infrastructure. ⚡
Read more below. 🤝
https://t.co/eCOpGyRdnM
#AI #OpenSourceAI #Inference #LLM #GenAI #AIInfrastructure #Nebius #MLSys #EigenAI
Every AI product running in production depends on an inference system someone had to engineer, optimize and more.
Scheduling, batching, routing, cost per token.
This is a craft.
The Inference Frontier Program spotlights the builders behind that work. 💡
Watch the video and nominate a team: https://t.co/Aae6WAvPkx
𝐇𝐚𝐩𝐩𝐲 𝐍𝐞𝐰 𝐘𝐞𝐚𝐫! 🎆
2025 was our 0→1 year at Eigen AI.
We set out to build the foundation for 𝐀𝐫𝐭𝐢𝐟𝐢𝐜𝐢𝐚𝐥 𝐄��𝐨𝐥𝐯𝐢𝐧𝐠 𝐚𝐧𝐝 𝐄𝐟𝐟𝐢𝐜𝐢𝐞𝐧𝐭 𝐈𝐧𝐭𝐞𝐥𝐥𝐢𝐠𝐞𝐧𝐜𝐞 (𝐀𝐄𝐈): high-performance AI that’s accuracy, fast, reliable, and cost-effective enough to ship into real production.
In the past 2025, we:
• Earned the trust of teams at multiple leading enterprises
• Reached #1 output throughput on Artificial Analysis across multiple benchmarks
• Shipped day‑0 serving support for OpenAI GPT‑OSS (120B/20B) on Hopper + Blackwell GPUs and launched a free playground
• Advanced open-source inference performance with SGLang (Multiple Token Prediction), showing up to ~60% throughput gains in our reported benchmarks
• Turned research into real systems: Eigen‑1, VISTA‑R1, WorkForceAgent‑R1, Eigen‑Banana and more
Most importantly: we built a team and a culture that can execute at the frontier — research, systems, and product, with speed.
2026 is about scaling this foundation:
• Expanding EigenLoop post-training platform
• Pushing the speed‑quality frontier across more modalities
• Helping more top teams go from prototype to SLA‑backed production, faster than ever
If you’re building AI products where accuracy, latency, throughput, reliability, and unit economics matter, let’s talk.
AGI tomorrow, AEI today. 🚀#AI #LLM #Inference #MLOps #OpenSource #GenAI #infrastructure #MLsys
🚀 Introducing Flash-ColReduce: a CUDA kernel for fast, memory-efficient attention statistics.
Exact column-wise softmax reductions, no QKᵀ materialization, >5X faster than PyTorch.
🚀 Eigen AI is heading to NeurIPS 2025! If you're joining us in San Diego (Dec 2–5), swing by Booth #943 to see what we're building.
Bonus: On Dec 3, 6–10 PM, we’re hosting a 𝐖𝐨𝐫𝐤𝐬𝐡𝐨𝐩 & 𝐏𝐚𝐫𝐭𝐲 on Efficient AI Computing — a casual evening of lightning-fast talks, deep dives into efficient LLM training/serving, and good food + great conversations. 🍽️✨
Whether you’re into accurate and high-performance model post-training, RL, compression, inference, deployment, or just want to chat about scalable AI — we’d love to meet.
➡️ Register now & catch you in San Diego!
https://t.co/2iClFWYOW0
🚀 #Eigen AI ranks #1 on Artificial Analysis @ArtificialAnlys for DeepSeek-V3.1-Terminus, Qwen3-VL-235B-A22B (BF16), and GPT-OSS-120B GPT-OSS-120B — hitting 791 tokens/sec, faster than any other provider with our EigenInference and EigenDeploy frameworks.
Pure optimization through model & system design, no hardware tricks.
Deploy anywhere: cloud, on-prem, or hybrid GPU clusters.
Built on QAT, sparsity, and kernel-level innovation.
Check out our blog for more details: https://t.co/FjCzgBB3AI
Try out our playground and lightning fast API at: https://t.co/FQJYZDirw2
Special thanks to the @ArtificialAnlys team for their great work on onboarding Eigen!
#LLM #Inference #AIInfra #eigen #eigenai #speedup #llm
🚀Founded by four dedicated MIT graduates, Eigen AI is the world's first company focusing on AEI – Artificial Efficient Intelligence, making AI accessible for all.
Today OpenAI dropped GPT-OSS. We teamed up with our partners SGLang @lmsysorg and @NVIDIA to deliver open-source support of the model with blazing-fast performance on Hopper and Blackwell GPUs just within 4 hours of the release. 🔥
With @YottaLabs, we're stoked to launch a free GPT-OSS-120B playground chatbot & API at https://t.co/BQfsnXIGFo 🚀 Easy-to-use, high-performance, and ready for your projects. Share with us what you are building with it! 🌟
Join us to unlock AI’s potential. Let’s democratize efficient AI for everyone! 💪 #AI #Innovation #EfficientAI #Chatgpt #GPT #performance #LLM #openai #eigenai
Tired intricate system code for RL training? 🤯
We release AReaL-lite – A lightweight AReaL version for AI researchers! 🚀#opensource
✨ Algorithm-first design & APIs🎉 ✨ 80% less code w. 90% AReaL's full efficiency 🎉 ✨ Customizable agentic RL🎉
🔗 https://t.co/YUa03pp9LR
I often wonder how Meta did such a good job post training the Llama series of models.
They just released a paper that gives us a good idea.
The big challenge is that using a single reward model to align an LLM on multiple tasks fails due to reward hacking, multi-objective issues, contradictory goals and so on.
They introduce "Constrained Generative Policy Optimization" that uses a Mixture of Judge models to identify a perfect blend of RLHF goodness.
They have the False refusal judge, precise instruction following judge (which is why 1B is so good), regex math/code reasoning judge, factuality judge, and safety judge.
The idea is pretty simple, and the graphic gives a good explanation. Amazingly, it improves everything: MATH, Human Eval, ARC, AlpacaEval, etc.
Bring on the judges!