Kernel Engineer at @Baseten | Ex CentML (Acq. by NVIDIA) | Ex Neural Magic (Acq. by Red Hat) | Research at UofT, EPFL, Uni. of Tehran, IPM Institute of Research
We're building in Canada!
We opened two new offices in Toronto and Montreal, and we're growing fast. If you’re excited about building high-performance AI infrastructure for the world's most ambitious companies, we'd love to meet you!
Fast image generation has become critical for production AI products, where latency directly affects user experience, throughput, and cost.
Proud to share we at Baseten reached a new milestone in optimizing image generation serving for two frontier models, FLUX.2-dev and Qwen-Image, on NVIDIA Blackwell and Hopper GPUs! 🚀
Key results:
1️⃣ FLUX.2-dev:
* 2.3× faster on B200
* 1.9× faster on H100
2️⃣ Qwen-Image:
* 1.57× faster on B200 with FP4
* 1.18× faster on B200 with FP8
* 1.08× faster on H100 with FP8
These gains come from deep-stack optimizations. Read the article for more details: