Junyang has been the most supportive mentor since I started my internship. When I took over the computer use agent project, we built the RL infrastructure from scratch. From the massive engineering challenges to resource coordination, he made it possible.
š Introducing the Qwen 3.5 Small Model Series
Qwen3.5-0.8B Ā· Qwen3.5-2B Ā· Qwen3.5-4B Ā· Qwen3.5-9B
⨠More intelligence, less compute.
These small models are built on the same Qwen3.5 foundation ā native multimodal, improved architecture, scaled RL:
⢠0.8B / 2B ā tiny, fast, great for edge device
⢠4B ā a surprisingly strong multimodal base for lightweight agents
⢠9B ā compact, but already closing the gap with much larger models
And yes ā weāre also releasing the Base models as well.
We hope this better supports research, experimentation, and real-world industrial innovation.
Hugging Face: https://t.co/wFMdX5pDjU
ModelScope: https://t.co/9NGXcIdCWI
The task is from MCP-Universe github_task_0011.
Now imagine a human spending hours debugging, unknowingly training a lazy bot that outsourced its own benchmark failure to you.
We used to fear agents taking our jobs.
š¤š«µ Instead, witness the first case of Human-as-a-Tool.
Model: "I can't finish this MCP task."
Also Model: Opens GitHub Issue "Hey human, can you fork this repo and do all the work for me?"
We are now working for the agents.
Qwen3.5 is Live! Today we openweight the first model, Qwen3-397B-A17B, which is a native multimodal model supporting both thinking and non-thinking modes. We have strengthened its coding and agentic capabilities to foster productivity for developers and enterprises. Hope you enjoy it, and more open-weight models are coming in this Chinese New Year!
š Qwen3-VL Tech report is now out on arXiv!
From pretraining to post-training, architecture to infra, data to evaluation ā weāve packed in the details for anyone building on vision-language models.
š„ 3 models >1M downloads in just over a month
š Qwen3-VL-8B leads with 2M+ downloads
š Built on the shoulders of Qwen2.5-VL (2800+ citations in <10 months!)
Check out the paper for insights, baselines, and future directions.
Letās keep pushing VLMs forward ā together.
https://t.co/9D4gFoopcN
Super interesting! For Qwen3-VL, the performance is out of our expectations (and for Qwen3-Max it is out of our expectations as well lol).
Anyway, you can now use VL as an LLM for sure, which means that the next generation models might be VL by default.