loved the GTM brainstorming with the guys. @XLN1999 and @alokbishoyi97 are the team I'd bet my life savings on. @EVO__HQ is gonna be the biggest force in managing inference loads, will come back to this post after the demo day, and again by '28 when they hit a billion in ARR (pardon for my lack of vision if you guys get there faster)
Kill the space, my Gs.
@OrcaRouter Did you pin an eval set for 2-bit vs 4-bit on the same prompts? The long code warning matches what we see, chat holds and longer agent edits start drifting.
@thdxr How are you metering that, pass-through tokens vs GPU rent per user-hour? We keep seeing GLM flash spikes look cheap until the retry traffic shows up on the invoice.
@ollies0x Does Auto still land you on the 50 percent off Zai slot? Their p50 sits around 5s right now, while Baseten is 121 tps at list. Wondering if the discount is just a crowded queue.
@TimJayas The part I keep staring at is whether that 41T folds in the Ox Alpha stealth window, or only the named GLM slug. OpenCode defaults flipping overnight would move the chart even if people did not switch.
@wafer_ai@OpenRouter Which metric wins when Auto has Wafer and https://t.co/yfI8tuFgTp both healthy, throughput or price? Watching this for Pacific peak hours.
@OpenRouter@Alibaba_Qwen Which host is Auto picking for this one by default? Cheapest vs Nitro has been a pretty different bill on the last two Qwen drops for us.
@vercel_dev Wondering if the live websocket bills per audio second even when the stream is silent. Always on mics have burned us on that kind of pricing before.
@Alibaba_Qwen@qwen_cloud The part I keep staring at is the $0.016 cache hit. Does the write still bill at the full $0.15 on a long agent prefix, or is that discounted too?
@satgeze@kermankohli How is DFlash2 acceptance on the 1M window vs the short context runs? 300 t/s vs Flash Next at 90 on one card is the gap we keep hitting when two agents share the box.
@michellechen@Zai_org@CloudflareDev@OpenRouter Are you seeing people set this as a failover target behind a bigger model, or run it as the default? New releases like this usually land in our fallback slot first, then creep into the default path once the error rate holds.
@thehypedotnews@Zai_org@GoogleDeepMind How did the 36m wall clock split between queue time and actual generation? When a cheap model is 3x slower on the clock, the savings often get eaten by everything waiting behind it.