DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro.
⚡ Up to 4.6× the speed of autoregressive decoding, with the same output.
This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!
https://t.co/We0lwYPSBl
Just pushed DFlash2 (@inco_ai) recipes to the Qwen3.8 27B cookbook⚡️
https://t.co/eAe1dFO5Sr
The community has been seeing great results with NVFP4 + DFlash2, and these recipes should be some very good starting points to play with.
More Qwen3.8 27B updates on the way 🫡
DFlash runs by default on our Kimi K2.7 Code endpoint, where we lead every provider in output speed on @ArtificialAnlys.
Congratulations to @inco_ai on DFlash 2. Thanks for the shoutout.
We look forward to even higher speeds ⚡
We’ve been hard at work these past few months, and this is just the beginning 🚀
Today, we’re announcing DFlash 2, one of many innovations powering the inference stack at @inco_ai.
DFlash 2 delivers up to a 4.6× speedup over standard decoding, unlocking major performance gains for both frontier-scale models and smaller models running on your personal AI computer.
DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro.
⚡ Up to 4.6× the speed of autoregressive decoding, with the same output.
This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!
https://t.co/We0lwYPSBl