Local AI addiction starts innocently. I just want privacy.
Then you’re downloading 200GB checkpoints at 2am, comparing 4-bit quants, calculating memory bandwidth over breakfast, and calling a new GPU infrastructure.
Anyway, one last benchmark and I’m going to bed ☺️.
AI model names are getting out of hand.
Soon we’ll be downloading Qwen4-Ultra-Flash-Next-A3B-Instruct-Thinking-MTP-FP4.
The model will fit in memory. The filename won’t.
Qwen3.8-Flash is changing local AI.
Frontier scale models are no longer just something you rent from a cloud provider. More builders can now run them, test them and build with them.
AI moves so fast now that a model can be best in the world at breakfast and
last generation by dinner.
By the time you finish reading the benchmarks, something newer has already shipped.
Qwopus 3.8 27B NVFP4 with a BF16 LM head, vision preserved, and DFlash2 K=16 on one DGX Spark. 50+ tok/s where it matters. Not chasing a speed leaderboard, but the best quality-to-speed ratio for a model you can run every day.
https://t.co/eACb09NuLy
Just released a ModelOpt NVFP4 conversion of Qwopus3.8-27B-Flash for NVIDIA Blackwell local inference.
Fresh conversion from the pinned BF16 source, retaining native NextN/MTP self-speculation and BF16 vision components, with SGLang tool-calling guidance and a full conversion receipt.
Huge thanks to @KyleHessling1, for Qwopus3.8-27B-Flash,possible.
https://t.co/WedGO8oTF5
Just released a ModelOpt NVFP4 conversion of Qwopus3.8-27B-Flash for NVIDIA Blackwell local inference.
Fresh conversion from the pinned BF16 source, retaining native NextN/MTP self-speculation and BF16 vision components, with SGLang tool-calling guidance and a full conversion receipt.
Huge thanks to @KyleHessling1, for Qwopus3.8-27B-Flash,possible.
https://t.co/WedGO8oTF5
GPT-6 Astra is not a normal model update.
98.6% ARC-AGI-3.
97.6% FrontierMath T4.
95.9% BenchCAD.
74.1% DeepSWE.
64.6% Terminal-Bench Science.
The story is the range: reasoning, science, engineering and coding all moved at once.
A serious release from @OpenAI.
NVIDIA buying Hugging Face is a masterstroke.
It already owns AI compute. Hugging Face is where millions of builders discover models, share datasets, test demos and start deploying. The GPUs power AI. The platform shapes which AI gets used. That’s the real acquisition.
Working in AI right now:
Wake up: new model released.
Download it -Read the model card - Fix the runtime. - Run benchmarks - Write up the results.
Go to sleep.
Wake up: a better model released.
Keeping up with AI is now a full-time job performed after my full-time job.