NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately.
- Parallel box decoding
- 10× faster than Qwen3-VL
- Trained on 138M queries/785M boxes
- GUI, OCR, and document layout, dense detection
- Open source
Useful for computer-use agents and Physical AI,
Qwen3.8-27B needed 197 agent calls to finish this task. Its new post-train did it in 68.
@hcompany_ai recently released Holo4-27B, a new computer-use post-train of Qwen3.8-27B.
They built a self-playing Pac-Man game in Godot
Qwen3.8-27B
🐌 197 calls
🧠 11.4M tokens
...but Holo4-27B
⚡ 68 calls
🧠 2.4M tokens
That’s ~79% fewer tokens and 65% fewer agent calls on this task.
H post-trained Qwen3.8-27B specifically for computer-use/agentic work.
And there’s an official 16.9GB Q4_K_M GGUF.
Same Qwen3.8-27B base, but a different amount of work to finish the same job.
I have to try it to see if it loses accuracy.
🔗 Link in ALT
Just doubled the prompt processing speeds from 539 token/s to 1290 tokens/s.
Qwen3.8-Flash-Next on consumer GPU is now blazing fast.
Next on the list is further optimizations on kernels and swapping which should boost results by 3X-4X.
https://t.co/83diJSdDja
Still cooking on Saturday.
We compressed our most popular local cyber model down to 15.7 GB.
Meet OrcaSAQ-2 Cyber 27B Uncensored GGUF — built for defensive red teaming, vulnerability research, security coding, terminal workflows, and authorized security testing.
54.7 → 15.7 GB
262K context
94.4% Top-1 agreement
Cyber capability should not require sending your source code, logs, or vulnerabilities to someone else’s cloud. Run it locally. Keep the data local.
Our smallest cyber model yet and one of the most capable we’ve evaluated in this size class. Have fun!
https://t.co/QhItgfVVCy
You can run locally Qwen3.8-Flash-Next (125B MoE) on
a 12GB RTX 5070 in One-click install.
it a custom engine for this model and this kind of GSQ-RCO quants, and this CUDA + system-RAM setup.
- 128K context.
- 65 tokens/sec.
- Q2_0 to 65.1 tps
- IQ2_XS to 52.0 tps
- IQ3_XXS to 44.8 tps
- Prompt processing is 400–540 tps.
- OpenAI + Anthropic compatible API.
This guy was getting ~15 tps running Qwen3.8-Flash-Next on his 12GB RTX 5070.
Apparently that wasn't good enough. 😂
So he built his own inference engine. (as we all should)
Now he's reporting
🐌 llama.cpp → ~15 tps
🚀 Strata → up to 65.1 tps
And this isn't big workstation either
Specs
🎮 RTX 5070 12GB
🧠 64GB DDR5-5600
⚙️ Ryzen 5 7600
🪟 Windows
At 128K context his new Strata engine reports:
Q2_0 → 65.1 tps + 543 tps prompt
IQ2_XS → 52.0 tps + 472 tps prompt
IQ3_XXS → 44.8 tps + 414 tps prompt
For Qwen3.8-FLASH-NEXT. 👀
He built it specifically around this model and this kind of CUDA + system-RAM setup, paired with RCO-GSQ quants. 👉 And he open-sourced it.
I love these kind of Local AI projects.
🔗 Link in ALT
10 GITHUB REPOSITORIES THAT FEEL ALMOST ILLEGAL TO BE FREE
And yes, they're all open source.
1. Archify
Turn codebases into architecture, workflow, sequence and data-flow diagrams with AI.
https://t.co/3jLToYDzic
2. OpenMAIC
A multi-agent environment where AI agents can collaborate and work together.
https://t.co/0g98VBKflm
3. DeepSeek Harness
An open-source harness for building and running AI coding workflows.
https://t.co/c45mNs83y6
4. Ponytail
A simple approach to helping AI coding agents write less unnecessary code.
https://t.co/UdFeGXFvjL
5. Agent Skills
Reusable skills that give AI coding agents more specialized capabilities.
https://t.co/KkT71vIiR3
6. OmniVoice Studio
An open-source studio for experimenting with AI voice generation workflows.
https://t.co/zfiVHiHYWe
7. Scientific Agent Skills
Specialized skills for AI agents working on scientific research and analysis.
https://t.co/w8HCUPfmQ7
8. Orca
Run multiple AI coding agents in parallel inside one development environment.
https://t.co/rINx3HgPEZ
9. MiniMind
A compact project for learning how language models work by building one yourself.
https://t.co/sz3WQWSLQz
10. God's Eye View
A visual project exploring how AI can understand and represent complex systems.
https://t.co/H7tiOUFLLu
All open source.
Some are useful today.
Some are worth studying.
Some could become much bigger.
Bookmark this list for the weekend.
Introducing GAE — Geometry-Native Autoencoder.
The key choice is where generation happens. GAE lets video models generate directly in a geometry-native latent space.
We believe geometry is more fundamental than texture for future video and world models. Texture describes how a scene looks; geometry describes how it is structured and what becomes visible as the camera moves.
GAE learns a compact latent space from geometry features, then trains a video generator in that space. Camera trajectories guide generation, and the generated state decodes into RGB, depth, and 3D.
Generation in a geometry-native space is the core.
🎬 Watch the scenes unfold in the teaser below.
Paper: https://t.co/lMJ6aQua4S
Code: https://t.co/VEhxfTzBGa
Project: https://t.co/kCF8YeZuTH
⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.
Five models, one complete audio stack: understanding, generation, interaction & creation.
Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off.
Highlights: 🥳
- ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts.
- ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning.
- TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions.
- TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ads.
- Realtime: speak & listen at once with anytime interruption, just like a real call; it even slows down and responds empathetically when it senses a low mood.
Unlock the full potential of Qwen-Audio-3.1! 👇
- Blog: https://t.co/e0M8wvj8Kg
- Qwen-Audio-3.1-ASR:
https://t.co/88lPQTPxyz
- Qwen-Audio-3.1-Realtime:
https://t.co/Y63sdtK2AC
- More APIs: coming soon @qwen_cloud
"still filming?"
15s handheld Shibuya night walk with #MiniMaxH3.
Running locally on a 5070 12GB in #ComfyUI.
(Model & LoRA links ⤵️)
(Full prompt in replies⤵️)
#AIvideo
what the f*ck, new Uncensored Qwen-Image-2.1, actually abliterated this time
and you can run locally in your laptop (Q4_K_M is 4.68 GB only)
- the fix is in the text encoder, not the image weights
- refusals drop from 100 out of 100 to 5
- it does not get dumber on normal prompts
- edit with up to 10 reference images
- native transparency (stickers, cutouts, no extra mask step)
- ComfyUI + city96 GGUF loader
https://t.co/OQKQiYY7DK
DiffusionGemma-as-Jev (aka djev) running near-real-time vision detection from a mobile phone using its native vision tower.
Please don't fall down the stairs!
This is a banger node set for FastH3 or MiniMax.
runs chunks of video concurrently, giving you the option to generate anywhere from 30 to 120 seconds each generation. game changer!
LOW VRAM SAFE. rids weights from memory after each chunk completes.
https://t.co/1YFf1Oq7zq