We just released Qwen-Drive-1.0.
Our goal: build a Vision-Language Foundation Model for autonomous driving that unifies 3D perception, driving understanding, and planning—while preserving the general capabilities of the pretrained VLM. 🧵
#AutonomousDriving#VLM#Qwen
@superalesha Camera calibration could be a problem. But I should point out that our planning section is merely a toy example that uses a very small amount of data (compared to real-world models), and it serves primarily as supplementary evidence of VLM capabilities.
Qwen just stepped into autonomous driving! 🚗
Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning in one framework, with the Qwen3.5-4B backbone left fully unmodified. Apache 2.0. 🤖 https://t.co/dWXStVVlQA
⚙️ Two plug-in modules do the driving: a BEV head for 3D perception, and a flow matching Planning Expert for trajectories.
📊 Leads driving VQA across the board: 77.8 on LingoQA, lowest Ego3D distance error, and 41.3 on causal reasoning where others score under 5.
🏁 The RL planner hits 90.7 PDMS on NAVSIM, ahead of AutoVLA and SpanVLA.
🧠 No catastrophic forgetting: general benchmarks stay on par with the base model.
Two open questions remain:
• Rare/extreme driving events are underrepresented in public data, so OOD generalization may be limited.
• We use no simulator-generated training data, so adaptation to simulator distributions and interactive dynamics remains open.
We just released Qwen-Drive-1.0.
Our goal: build a Vision-Language Foundation Model for autonomous driving that unifies 3D perception, driving understanding, and planning—while preserving the general capabilities of the pretrained VLM. 🧵
#AutonomousDriving#VLM#Qwen
One open problem is reasoning–planning consistency.
Driving rationales can be under-specified: multiple causes may coexist at different time scales. We therefore avoid an explicit consistency loss for now—an incomplete rationale should not rigidly constrain planning.
Excited to share Qwen-VLA paper, our exploration of generalist Vision-Language-Action models.
It extends Qwen’s multimodal backbone from visual understanding and reasoning to continuous action generation and trajectory prediction.
Paper:
https://t.co/9jvRW0nI8B
Excited to announce that https://t.co/gxEWEgS0dk now supports the latest Wan2.2, while continuing compatibility with Wan2.1 and HunyuanVideo! 🚀
If you find our project helpful, please give us a ⭐️ on GitHub. Your support means a lot! #AIGC#AI#wan#videogeneration#Diffusion
Thanks for sharing!
The code for HunyuanVideo and Wan2.1 is now publicly available on GitHub: https://t.co/gxEWEgS0dk.
Please consider giving the repository a star and joining the discussion. 🥳🥳
🚨Paper Alert 🚨
➡️Paper Title: Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
🌟Few pointers from the paper
🎯Video generation models have demonstrated remarkable performance, yet their broader adoption remains constrained by slow inference speeds and substantial computational costs, primarily due to the iterative nature of the denoising process.
🎯Addressing this bottleneck is essential for democratizing advanced video synthesis technologies and enabling their integration into real-world applications.
🎯Authors of this paper proposes “EasyCache”, a training-free acceleration framework for video diffusion models.
🎯EasyCache introduces a lightweight, runtime-adaptive caching mechanism that dynamically reuses previously computed transformation vectors, avoiding redundant computations during inference.
🎯Unlike prior approaches, EasyCache requires no offline profiling, pre-computation, or extensive parameter tuning.
🎯They conducted comprehensive studies on various large-scale video generation models, including OpenSora, Wan2.1, and HunyuanVideo.
🎯Their method achieves leading acceleration performance, reducing inference time by up to 2.1-3.3 compared to the original baselines while maintaining high visual fidelity with a significant up to 36% PSNR improvement compared to the previous SOTA method.
🎯This improvement makes their EasyCache a efficient and highly accessible solution for high-quality video generation in both research and practical applications.
🏢Organization: @2024_HUST , MEGVII Technology, @HKUniversity
🧙Paper Authors: @THELMDOFZHOUXIN , Dingkang Liang, Kaijin Chen, Tianrui Feng, Xiwu Chen, Hongkai Lin, Yikang Ding, Feiyang Tan, @HengshuangZhao , Xiang Bai
📝 Read the Full Paper here: https://t.co/fAWNwb1TcU
🗂️ Project Page: https://t.co/gu1C89Rb5V
🧑💻 Code: https://t.co/77Rf1VPv3e
🎥 Be sure to watch the attached Demo Video - Sound on 🔊🔊
Find this Valuable 💎 ?
♻️QT and teach your network something new
Follow me 👣, @NaveenManwani17 , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements.