@shamaz89720417 Would love a llama.cpp path later, but not via a simple quantize patch. Hard parts were fused MoE NVFP4, vLLM scales, GB10 vision kernels - patches + quantize script on the model page (Apache-2.0). Happy to review a port:
https://t.co/Jj3SXxlmbT
I shipped my first model release 🚀
A vision-safe NVFP4 quant of Qwen3.6-35B-A3B for the NVIDIA DGX Spark (GB10).
67 GB → 22 GB. Native vision fully intact. 512K context. Built for running agents + sub-agents concurrently.
No existing tool could quantize this arch — so I wrote the tooling. 🧵
Model + quantize script + vLLM patch, all Apache-2.0:
🤗 https://t.co/Jj3SXxlmbT
Running a DGX Spark and want Qwen3.6 intelligence + native vision without a second model eating your VRAM? This is for you. Feedback welcome 🙏
So this build ships MTP and it's a lossless ~1.26–1.44× decode speedup that keeps fp8 KV (no concurrency cost).
Best part: it holds ~73% draft acceptance at 64k context — where the external DFlash drafter collapses to 0%. It wins exactly where long-context agents live.
I opened X today to learn about the DeepSeek Innovation in AI and got his post recommend as first entry when I opened the app. A great and succinct summary of what all the optimizations are. Very promising!
@OpenAI I am paying for #chatgpt, my prompts and conversations are either super slow or I have to login again which then fails due to busy servers for days now. Any comments on why this is and when I can use the service reliably again? It is a very helpful companion for me.
@iam_chonchol Max. was 3 pages: As an AI [LM], I can provide you with information and examples, but I am not able to create a complete PowerPoint presentation or VBA script for you. My role is to assist and provide guidance, but ultimately it is up to you to create the presentation and script.