Kimi K3 is live on Modal.
Moonshot has shipped the world's first open 3T-class model, and we're a day zero launch partner.
We trained a custom DFlash speculator for K3's novel architecture so you can run it faster, losslessly. The most capable open model we've worked with by far.
Kimi K3 is live on Modal.
Moonshot has shipped the world's first open 3T-class model, and we're a day zero launch partner.
We trained a custom DFlash speculator for K3's novel architecture so you can run it faster, losslessly. The most capable open model we've worked with by far.
Inkling by @thinkymachines is now available on Modal, backed by a custom DFlash speculator for 67% higher throughput and interactivity.
Running on Modal Auto Endpoints with SGLang today.
Our new Auto Endpoints feature is powered by a new Modal primitive: Modal Servers.
In this blogpost, we walk through design principles and detailed architecture: @EnvoyProxy, @googlecloud Spanner config store, and a @Cloudflare Pingora-based custom proxy.
Modal Auto Endpoints provide state-of-the-art open source inference perf with a click.
Learn how we developed our low latency inference playbook with @DecagonAI, delivering responses 60ms faster than the best proprietary provider.
https://t.co/7OMcB4mD7N
Speculation Is All You Need.
In this blog post, we announce the co-release (w/ Z Lab) of six more state-of-the-art DFlash speculators for @Alibaba_Qwen 3.x.
Over 1k output tps for 3.5 122B-A10B on a B200.
Read the blog for why we're all-in on spec dec.
https://t.co/Bv3Zc95Xgh
We're bringing together our friends and community to celebrate our Series C.
Join us at Noguchi's Sunken Garden in NYC on June 16th or at the Legion of Honor in SF on June 25th.
Invites are limited, apply here: https://t.co/B4h0C7Wq0Q
Today we're announcing our Series C funding: $355M at a $4.65B valuation, led by some great investors @generalcatalyst and @Redpoint.
We've had insane growth in the last year, but we're still very early. So proud of the team and what we have built so far!
FormulaCode tests AI's ability to (super-)optimize at the system level: can LLMs optimize multi-load benchmarks of entire codebases? Turns out, human experts are still better (for now).
(appearing at #ICML2026)
Check out FormulaCode๐๏ธ! We study how coding agents tackle codebase optimization, where gains require coordinated tradeoffs and changes across the repo (just like tuning a car) - not just making one workload go brrr. Awesome working with @atharva_sehgal, @yisongyue, and the team!
Excited to share FormulaCode, a continually updating benchmark for evaluating the holistic ability of LLM agents to optimize codebases. Our current dataset consists of 957 tasks curated from 245477 pull requests in 70+ repositories (and growing!).
๐ https://t.co/bmUgfZkvsw
๐งต๐
This is @_gongy. He leads our LLM inference team. He has a car that goes vroom!
You can work with us and have a car that goes vroom!
https://t.co/tXDNmmf5Kw
We release ForeAct (accepted to CVPRโ26๐), a world model planner powered by visual foresight for VLAs - efficiently, modularly, and at scale.
โจ Seamlessly integrates with VLAs by visual augmentation โ no architectural changes required
โก Generates high-fidelity 640ร480 subgoal images in just 0.33s
๐ง Significantly boosts generalization capability and data efficiency
๐Paper: https://t.co/aOadlsgVqu
๐Code: https://t.co/aPtnZ0t8Q4
๐งต๐