after buying your first 3D printer, don’t go for a second one.
instead buy a 3-axis desktop CNC mill.
that shit is gonna chnage the paradigm for you.
thank me later.
100% hard agree.
Muslims need to be focusing on the *UNSOLVED* problems we face (and there's a TONNE trust me).
If we actually spread out our effort among these problems we'd be in a FAR better place today.
Qwen3.8-Flash-Next, a 125B MoE model, running locally on a 12GB RTX 5070.
And apparently, it’s a one-click install.
The trick isn’t somehow fitting 125B parameters entirely into 12GB of VRAM.
This setup uses a custom inference engine, GSQ-RCO quantizations, CUDA acceleration and system RAM together to make the model practical on consumer hardware.
The headline configuration supports up to 128K context.
And the reported speeds are surprisingly high:
Q2_0: 65.1 tok/s
IQ2_XS: 52.0 tok/s
IQ3_XXS: 44.8 tok/s
Prompt processing is also around 400–540 tok/s.
So you’re looking at a 125B-class MoE model with a huge context window, running from a 12GB gaming GPU, while using system memory to handle what the GPU can’t hold.
That’s the part that makes this interesting.
For years, local AI has mostly been a hardware game:
More VRAM → bigger model.
This approach flips that around.
The inference engine, quantization format, memory offloading and GPU/RAM split become just as important as the GPU itself.
And because the engine exposes OpenAI-compatible and Anthropic-compatible APIs, you don’t necessarily have to build a completely custom application around it.
You can plug the local model into tools that already understand those API formats.
That opens the door to using it with local coding agents, development tools and other workflows that normally expect a cloud model.
Obviously, these numbers are tied to this specific engine, quantization, hardware and workload. A 12GB GPU isn’t magically storing a 125B model in VRAM.
But that’s precisely why this is worth paying attention to.
125B MoE.
12GB RTX 5070.
128K context.
Up to 65.1 tok/s.
400–540 tok/s prompt processing.
OpenAI + Anthropic compatible API.
Local AI keeps finding ways to make hardware limitations less absolute.
Not Blender btw, it's my own software. Blender just never clicked for me lol
Might sell it on Gumroad as a one-time buy if you guys want it
~2 Big Macs' worth
Real talk, I'm broke and need cash for my indie anime, mostly VAs & sound
Drop a comment!
Não tem ninguém falando sobre, então vou ter que soltar o verbo:
Se vocês analisarem bem, o Kyle Chandler faria um trabalho infinitamente superior em TODOS OS PAPÉIS da carreira do Pedro Pascal!
Inclusive o Reed Richards…
@XiaomiMiMo Hey I won the 100T token Grant but it never got credited to my account. I had emailed but never got a response. Any possibility to rectify this please? Super excited for your new release.
Is anyone checking the accuracy of the models that inference providers are providing? 1000+ TPS is great but if it's a 10% hit on ability then is it really worth it?
@dhh@mickcodez@nvidia@OmarchyLinux Idea for a Benchmark:
Tasks like this being done across many different models and test them on how many runs to completion, no of tokens and cost per completion?
Rambler feature in Gboard on a Google Pixel 9
Most of the processing happens in the cloud, it can literally work on almost all pixel (and non pixel) devices, not just Pixel 11 series.
But offline mode will be broken because it definitely requires necessary hardware/models which other devices may not have.