π New #1 on the @ArtificialAnlys Kimi K2.6 leaderboard: Eigen AI at 265 tok/s β in collaboration with @Nebius Token Factory. @nebiustf
On B200. Not GB300.
Topping the chart without Blackwell Ultra isn't a silicon story β it's a serving stack story. Every layer of EigenInference is co-designed for trillion-param MoE.
More coming π₯
https://t.co/YiUoasxFko
Today, we're announcing that Eigen AI is joining Nebius (NASDAQ: NBIS).
From day one, our mission has been Artificial Efficient Intelligence β building the world's most efficient engines for generating intelligence. Together with Nebius, we're working toward the best AI cloud, uniting Eigen's full-stack model and inference software, ranked #1 on Artificial Analysis for inference speed, with Nebius's global hardware and infrastructure footprint, so any developer or enterprise can run the best models at the best price, with no capacity ceiling.
After close, Eigen's optimization stack will be integrated directly into Nebius Token Factory. The entire Eigen AI team is joining Nebius in full, establishing Nebius's engineering and research presence in the San Francisco Bay Area.
To our customers, our team, our investors at Tectonic Ventures, E14 Fund, Uncorrelated Ventures, and AGI House Ventures, our angel investors, advisors, mentors, and supporters β and to the Nebius team for the conviction and partnership β thank you.
The mission doesn't change. The leverage behind it does.
Ryan Hanrui Wang, co-founder and CEO of Eigen AI, said:
βWeβre proud to join Nebius and work alongside the Token Factory team to push the boundaries of inference performance. Nebius has built a world-class AI cloud with a deep engineering culture that perfectly aligns with our own. Together, we are removing the friction of AI model customization and deployment so developers can run models reliably in production without managing the underlying infrastructure.β
Full announcement at:
https://t.co/XEeMZH4QKv
Run multimodal agents faster on Eigen AI Model Platform with @nvidia Nemotronβ’ 3 Nano Omniβa single model for text, video and audioβNVFP4 optimized on NVIDIA Blackwell β 500+ tok/s/user, zero quality loss across multiple benchmarks vs. BF16 π₯π₯
Explore how Eigeninference brings production-ready performance from day one π https://t.co/zEW3TXSDDD
Try it out at the Eigen AI Model Studio π https://t.co/iMQ2KMOYCw
The future of AI is open -- but it also needs to be fast, efficient, reliable, and production-ready.
Excited to partner with @NebiusAI to bring optimized frontier open models to Token Factory. π
Together, weβre helping developers and enterprises run leading open-source models in production with greater speed, reliability, and scale by combining Eigen AIβs deep inference optimization with Nebiusβs production-grade infrastructure. β‘
Read more below. π€
https://t.co/eCOpGyRdnM
#AI #OpenSourceAI #Inference #LLM #GenAI #AIInfrastructure #Nebius #MLSys #EigenAI
Open models are improving fast.
Running them efficiently in production is still hard.
@nebiustf Γ @Eigen_AI_Labs are partnering to bring optimized frontier open models to Token Factory.
DeepSeek, GPT-OSS, Kimi, Qwen, Llama, GLM and more, optimized for speed and efficiency at scale.
High-performance open model inference without building the optimization stack yourself.
Read more: https://t.co/PiavQ3dLBU
(1/7)π Eigen AI inference milestone.
Weβve reached industry-leading speed on Artificial Analysis across 11 major modelsβ including DeepSeek-V3, Qwen3, Qwen3-VL, and Llama 4.
This wasnβt achieved via per-model hacks, but by building a production-grade inference stack that scales across architectures and workloads.
#artificial_intelligence #DeepLearning #LLMs
π Introducing Flash-ColReduce: a CUDA kernel for fast, memory-efficient attention statistics.
Exact column-wise softmax reductions, no QKα΅ materialization, >5X faster than PyTorch.
We'll be at @NeurIPSConf next week β swing by the @Eigen_AI_Labs booth (#943) and say hi!
There's also a Workshop & Party on Dec 3 for anyone interested in efficient AI. Come for the ideas and discussions.
π Grab a spot here: https://t.co/zrVgFuR13W
See you in San Diego! β±οΈ
π #Eigen AI ranks #1 on Artificial Analysis @ArtificialAnlys for DeepSeek-V3.1-Terminus, Qwen3-VL-235B-A22B (BF16), and GPT-OSS-120B GPT-OSS-120B β hitting 791 tokens/sec, faster than any other provider with our EigenInference and EigenDeploy frameworks.
Pure optimization through model & system design, no hardware tricks.
Deploy anywhere: cloud, on-prem, or hybrid GPU clusters.
Built on QAT, sparsity, and kernel-level innovation.
Check out our blog for more details: https://t.co/FjCzgBB3AI
Try out our playground and lightning fast API at: https://t.co/FQJYZDirw2
Special thanks to the @ArtificialAnlys team for their great work on onboarding Eigen!
#LLM #Inference #AIInfra #eigen #eigenai #speedup #llm
Thrilled to share that our team just released the open-source ππ’π ππ§ πππ§ππ§π π. Feel free to explore the complete Eigen stack, including EigenTrain, EigenInference, and EigenDeploy. Experience lightning-fast model serving without compromising on quality! π₯
Excited to announce that @Eigen_AI_Labs teams up with @sgl_project, @nvidia and @YottaLabs to deliver free playground for @OpenAI latest GPT-OSS-120B model! Check it out at https://t.co/4PTUK3ORruπ
Eigen AI aims to democratize AIβs transformative power for everyone, everywhere.
πFounded by four dedicated MIT graduates, Eigen AI is the world's first company focusing on AEI β Artificial Efficient Intelligence, making AI accessible for all.
Today OpenAI dropped GPT-OSS. We teamed up with our partners SGLang @lmsysorg and @NVIDIA to deliver open-source support of the model with blazing-fast performance on Hopper and Blackwell GPUs just within 4 hours of the release. π₯
With @YottaLabs, we're stoked to launch a free GPT-OSS-120B playground chatbot & API at https://t.co/BQfsnXIGFo π Easy-to-use, high-performance, and ready for your projects. Share with us what you are building with it! π
Join us to unlock AIβs potential. Letβs democratize efficient AI for everyone! πͺ #AI #Innovation #EfficientAI #Chatgpt #GPT #performance #LLM #openai #eigenai
We are excited to open source TinyEngine, a memory-efficient and high-performance neural network library for Microcontrollers:
https://t.co/4ZxxQtpQDp
Tutorial of visual wake words on MCUs: https://t.co/SYhdlK4bae
Project page: https://t.co/ouy7niEilX
Welcome to try TinyEngine!