@Anbeeld@MiaAI_lab seeing with Deepseek, 94Gb is the max you can take with 256k~ context without offload and single request, but i might be wrong, and looking at the DGX forum:
Qwen/Qwen3.8-Flash-Next-FP8 186 GB
Inferact/Qwen3.8-Flash-Next-NVFP4 183 GB
RadixArk/Qwen3.8-Flash-Next-NVFP4 135 GB
Today we’re launching Portable Computer on @NVIDIA DGX Spark.
Portable Computer is a fully local version of Perplexity Computer, where the entire runtime: orchestrator LLM, subagent LLM, agent harness all run on your local hardware. No cloud dependency.
@MiaAI_lab that's a suer job, tried for some time to actually make it faster on my end, trying mixed quants to keep the best KLD and ended tuning the most used experts etc, but nothing. This is a godbless, i was getting demoralized about not having a second spark. Thanks 🙏