👀 Seeing is just the beginning.
With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native — read images, videos & documents, edit videos, work with 3D/CAD, and more.
From multimodal models → multimodal agents. 🚀
Watch it in action:
https://t.co/C2jQObXryM
DeepSeek V4 Flash is now #1 on OpenRouter.
It processed 7.22T tokens this week, up 13%. Both V4 Flash versions combined processed 670B tokens.
Chinese models dominate the top 10: DeepSeek, MiMo-V2.5, Hy3, and GLM-5.2. All are free or close to free.
Total platform volume: 3.29T tokens. Weekly pace: 61.3T.
GPT-5.6 Luna ranks #8 and is the only OpenAI model in the top 10. Its volume rose 465%, but DeepSeek still has 3.7x more.
This episode dives systematically into the co-design of models and infrastructure.
Kaichao You shared a vivid analogy: comparing tokens to electricity.
Hardware represents natural resources such as wind and water power.
Models are power generators, including wind turbines and hydroelectric generators.
The inference engine acts as the power grid system.
The depth of co-design dictates power generation efficiency. For areas with rapid water currents, what generator design can best harness this energy?
You Kaichao: vLLM, Open-Source Infra, Model Co-Design & Journey from Com... https://t.co/V5SYBtNeYj 来自 @YouTube