lithos-metal runs fast... like Qwen3.8-27B at 200 tok/s peak on an M5 Max chip fast 🚀 And now it's open-source.
The LithosAI team combined megakernel and DSpark optimizations for @Apple silicon: lithos-metal uses one Metal megakernel per layer mixer.
That's how it runs Qwen3.8-27B 2.1x faster than MLX, 2.4x faster than Ollama, 4.5x faster than vLLM-Metal at 32K context.
Explore the code, try it on your Mac, and share what you learn. We'll keep working to speed up local inference across more models and devices.
We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.
Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: https://t.co/A396OYmw1S
Tech blog: https://t.co/bJIlwz9462
Having the fastest inference for open models is great, but what we really care about at LithosAI is the end-to-end response time.
Our team is looking at AI serving from the agent perspective, so what matters most to us (and our users) is how well and how quickly we help agents execute the entire task you’ve assigned them.
Maximizing output speed is not the same thing as maxing agentic speed. They’re inherently related, but you won’t get the agent speed you need until you put end-to-end latency first.
Three models, three #1 spots on @ArtificialAnlys: Kimi K3, DeepSeek V4.1 Flash, and GLM-5.3-Flash. 🚀
Together with LithosBox, we’re accelerating the full agent loop with the fastest inference and sandbox execution. Try @lithos_ai on your workloads!
Another day, another LithosAI launch, and another #1 spot on @ArtificialAnlys benchmarks for inference output speed and end-to-end latency.
That’s 3 for 3: LithosAI has the fastest inference for Kimi K3, DeepSeek V4.1-Flash, and now GLM 5.3 Flash.
Combine that with LithosBox, the ultra-fast agent sandboxes we just launched, and our users can significantly speed up the agent loop. Get started with our public API.
Super excited to release LithosBox, our ultra-fast sandboxes for AI agents! 🚀
17× faster startup, 43× faster snapshots, 16× faster branching, and 6× faster restores than existing sandbox providers.
🎉 FREE for the next two weeks during public preview. Give it a try!
Excited to see LithosBox out in public preview 🚀
Sandbox is another crucial piece of RL rollout, coding agent environment. We’ve been pushing hard on ultra-low-latency inference, and now we’re bringing the same focus to agent execution environments.
Really excited for what this unlocks for agentic workloads!
Today, LithosAI is launching a public preview of LithosBox: ultra fast versioned sandboxes that your agents can snapshot, fork, and restore in milliseconds.
So far, we've been focused on providing ultra-low latency inference for open models like Kimi K3, DeepSeek V4.1, and GLM 5.3 Flash (stay tuned!). Now, we're giving agents runtime and post-training environments with Git-like control.
With LithosBox, agents can:
→ spin up, branch, and snapshot thousands of environments as they explore multiple paths forward
→ decide on a winner and rebuild the winning state
→ then throw away the environments they don't need, all in milliseconds
@dimitriosSkar, @JiaZhihao, and the rest of our team put together this agent vs. agent showdown with Doom so you can see how LithosAI helps your agents run 5x faster with AI serving infrastructure co-designed for ultra-low latency across inference and execution environments.
In our demo, the agent completes its entire run 5.29x faster with LithosBox and Lithos DeepSeek V4.1 Flash inference than it does with leading sandbox and inference providers.
Try LithosBox for free for up to two weeks via our public API.
We will be teaching Agentic GPU Programming for MLSys as part of the CMU ML Systems course this upcoming spring at @SCSatCMU, touching on knowledge curation, compiler analysis, and feedback loops. We also curated an online book on the topic, check it out https://t.co/S3TmfiAJ7h
AI agents are becoming GPU kernel engineers. We're rethinking the compiler stack around them.
Meet TIRx Harness: an open compiler harness that gives agents an environment to explore, debug, and optimize GPU kernels.
🧱 Minimal stable low-level compiler foundation for predictable GPU programming 🔍 Domain-specific compiler analysis to guide debugging and optimization 📚 Kernel zoo of 60+ kernels and hardware specs to reuse strategies and discover new ones 📊 Remote GPU evaluation for reliable benchmarks and profiling
Agent-built Kimi Delta Attention kernels achieved geometric-mean speedups of 2.94x over FlashKDA (forward) and 6.84x over FLA (backward).
Our bet: the next leap in agentic GPU programming will come from engineering the environment agents optimize in.
Check out our blog: https://t.co/aZo91bGTgz
A counterintuitive lesson we have learned: more serving tiers can mean *better* GPU utilization.⚡
Most providers offer only one API and one latency choice, because adding more can fragment GPU capacity. We’ve found the opposite: new inference and GPU virtualization techniques let workloads across tiers share the same GPUs while meeting distinct performance targets.
More choices for developers. Better utilization underneath.
Excited to keep pushing the Pareto frontier of speed, latency, and cost, and to see what that makes possible for agents.🚀
LithosAI is pushing the Pareto frontier of agent inference speed, latency, and cost, and our latest results on Artificial Analysis proves it.
Unlike most providers, offering more tiers improves our utilization.
Read the full story: https://t.co/TxBoGIeU2G
🚀 More reliable agents with DeepSeek V4.1!
XGrammar brings strict tool calling to SGLang & vLLM through Structural Tags, enforcing tool argument schemas in DeepSeek’s native format.
See how the Structural Tag works 👇
https://t.co/N0Tbl58Grf
Check out XGrammar 👇
https://t.co/GoOOhg1FLw
If you’re using AI benchmarks as absolute proof of what’s the best model, inference, or task-specific agent, you’re doing it wrong.
For example, our inference engine users care most about end-to-end latency, tokens per second per user, time-to-first token, and aggregate throughput and cost per token.
Our @ArtificialAnlys results for Kimi 3 inference provide a lot of this, but not all, and not as a 1:1 match to how the Lithos Engine will perform for your specific workloads and infrastructure.
Instead, engineering teams should be:
-Deciding what metrics matter and how to evaluate them based on actual use cases and traffic BEFORE seriously looking at benchmark results
-Evaluating existing benchmark methods against real production conditions including operational constraints and metrics that matter (e.g., different agents need different optimizations from TTFT and E2E latency to throughput, as well as their combined tradeoffs and costs)
-Filtering out the noise, aka pay attention to new benchmark results that fit your team’s product requirements and methodology standards
-Identifying promising contenders then rigorously testing them against production workloads
Once you treat benchmarking as a directional sign and you know what deserves your attention, all the results out there become a lot less distracting.
That’s why what happens after benchmark testing matters the most.
LithosAI has Day-0 inference support for DeepSeek-V4.1-Flash, accessible via public API.
Yesterday, our public API went live.
<12 hours later, @deepseek_ai launched its smallest, fastest, model yet.
<12 hours later, our team’s got you the ultra low-latency inference you need to max out agent speed with DeepSeek-V4.1-Flash.
Get started with Standard speed at 250+ tokens/s/user. Fast and Ultra are coming soon.
Been testing DeepSeek-V4.1-Flash all day. Truly surprising: a **relatively** small model matching GPT5.6-Sol quality at ~30x lower cost.
Try it on your own workloads — 250+ tokens/s/user, $0.6 per 1M output tokens ⚡
LithosAI now has the highest speed, lowest latency, and lowest price inference for Kimi K3. And now you can test it yourself with our public API.
Building coding agents, SRE workflows, or voice AI? Put LithosAI to the test: https://t.co/CPNXTdk2H4
Agentic SREs are a great addition to human-in-the-loop workflows, a lot of teams are using them to bring MTTR down to <20 minutes. When a cluster or production service is down, 20 or even 10 min of downtime is huge.
But building these agents forces a decision: do reasoning capabilities matter more than your SLA? Or should you prioritize resolution time over output quality?
The saying, "Good, fast, and cheap – pick two," applies here.
Kimi K3 significantly narrows the gap between closed-source and open-source LLMs. At LithosAI, we’ve been pushing hard on the LLM serving stack to make inference faster, more efficient, and ready for latency-sensitive agent workloads. Kimi K3 is now running at 800+ tokens/sec/user. Looking forward to seeing what people build with it 🚀
🚀We’ve been pushing agentic inference toward the physical limits of the hardware.
Announcing LithosAI’s first pricing tiers, with early-access pricing ahead of the September 1 API launch.
Kimi K3 is live now at 800+ tokens/sec/user on standard GPUs, with full model quality.
Try the live demo at https://t.co/eLo8WWPpIX and sign up for early access.
Open-weight models are no longer just catching up.
Kimi K3 brings frontier-level agentic performance with open weights. As we discuss in our latest blog, with frontier quality on open models, agentic inference performance becomes the differentiator.
Kimi K3 is coming to LithosAI when the weights drop.
Read more: https://t.co/p51MKknWol