How do you design LLM inference systems for the lowest possible token cost?
For the answers, check out @MeryemArik9's recent talk at @QCon SF here https://t.co/PwlayqTxAE
We have seen this same phenomenon but to an even more extreme level.
Deepseek-V4-flash (the original one) with 64 attempts beat opus 4.8 & GPT-sol on SWE bench pro for a lower cost.
Parallel agents are currently deeply unappreciated way to build frontier intelligence.
Note with the new deepseek - the tradeoff would be even more stark.
Yesterday, the Doubleword team took a trip to Bristol to visit Isambard-AI 🇬🇧
We got a first-hand look at the infrastructure powering some of the UK’s most ambitious AI work - and the scale of compute being built here.
Huge thanks to BriCS and Isambard-AI for having us.
Now, back to inference 🚀
Give your agents memory that stays up to date at a fraction of the cost.
Say an agent learns you live in San Francisco, then later that you moved to London. Consolidation updates the fact, so the agent doesn't hold onto both.
With the Doubleword × @betterdbhq integration, background work like fact extraction and nightly memory consolidation runs through Doubleword's async Flex tier, while recall stays fast and local.
Flex cuts the cost by 25% versus real-time, with prompt caching saving up to 90% more on input tokens, and no rate-limit handling to write yourself.
Full setup guide at https://t.co/ZrJ7m163EH
So if ChatGPT isn’t hallucinating, the only European inference provider with ZDR which offers GLM 5.2 and is price competitive with OpenRouter is @Doubleword_
FAQs answered a 🧵
Q1: what’s the catch?
A1: you need to deliver value along the way - so should be a cracked engineer in a related field eg distributed systems.
Q2: wait? Teaching people on the job is still a thing?
At Doubleword, we use CRIU to cold start large language models in seconds.
In the latest Inference Lab blog, Radostin Stoyanov explains how CRIU’s new on-the-fly snapshot compression helps us restore models from checkpoints even faster.
Across five models, it reduced checkpoint sizes by up to 45.5% and cut restore-to-first-token time by up to 44.7%, with the same validated output.
https://t.co/CIs4sLIEjn
Run Qwen3.8-27B on Doubleword - the most intelligent small model to date.
Built for workloads where you want strong reasoning and agentic performance without paying frontier-model prices.
Run it Realtime, or save on token costs by moving suitable workloads to Async or Batch.
Available now on Doubleword: https://t.co/XNFOeea250
New post: what happens when a GPU reads memory
We follow an LDG instruction through the memory hierarchy on an RTX 4090, from the warps down to the DRAM and back, reverse-engineering the undocumented parts along the way.
https://t.co/NxRd2XZMmh
New post on the Doubleword blog: the case for disaggregated prefill.
Disaggregated prefill is sometimes treated as a last resort to manage TPOT and TTFT SLOs. We argue that its much stronger than that, and all sufficiently large scale deployments ought to run disaggregated
https://t.co/VrqtLjnCWv
Your ai agent needs long-term memory that stays on your own infrastructure.
betterdb/agent-memory keeps recall local: a vector query on your own boxes, fast and free to call. The only paid part is the cold path, embedding and reconciling new facts. That now runs on @Doubleword_ async flex tier. Reads and storage stay local.
Full walkthrough: https://t.co/Fr58rooZ5f
Heading to MLSys in Bellevue? Come hang with LMSYS and RadixArk people!
RadixArk CTO @BanghuaZ will give the MLSys opening remarks on Day 1 and co-host the Young Professional Symposium talks and panels.
We're also co-hosting two happy hours (event links in the comments):
Mon, May 18
Happy hour with @allen_ai, sponsored by @CrusoeAI and @Doubleword_. Come meet @natolambert, @GenAI_is_real, @BanghuaZ, Connor Guerrero, Jamie Dborin, Carlo Mussolini, and the SGLang core members.
Tue, May 19
MLSys Happy Hour with @RadixArk, @EssenceVenture, and @DeltaInstitutes.
One more thing! If you'd like to sit down with someone from the team for research, partnerships, open source, or hiring, drop your details and we'll set up a 15 or 30 min chat: https://t.co/VMeiMDzjYT
See you in Bellevue!
🍻 MLSys 2026 Happy Hour is coming up!
SGLang & Ai2 (@allen_ai) are co-hosting a happy hour for the open AI community during MLSys 2026 in Bellevue, sponsored by @CrusoeAI & @Doubleword_!
Come hang with:
-@natolambert, Sr. Research Scientist at @allen_ai
-@GenAI_is_real, SGLang Core Dev
-@BanghuaZ, Co-founder of @radixark
-Connor Guerrero, Sr. DevRel Manager at @CrusoeAI
-Jamie Dborin, Co-founder & Head of Research at @Doubleword_ Inference Lab
-Carlo Mussolini, Member of Technical Staff at Fractile
Spots are limited, RSVP now and we'll see you there!
🕐 Mon, May 18 in Bellevue downtown
👉 RSVP required: https://t.co/ivx2zO3Ivo
We’ve partnered with @Doubleword_ to bring our structured generation engine directly to their inference platform.
No more bad outputs!
👉 https://t.co/EsqlHkZ8h0