Introducing Unsloth Desktop 🦥
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on https://t.co/xxBDOI9ApK and GitHub.
GitHub: https://t.co/2kXqhhvLsb
Blog and Guide: https://t.co/CYosNAQHva
You can now train and run 500+ models locally with our Unsloth Docker image! 🐳
Use our new GUI or notebooks workflow. No setup required. Works on NVIDIA and AMD.
Guide: https://t.co/ER5lOLGQa5
GitHub: https://t.co/aZWYAtakBP
Introducing Unsloth Desktop 🦥
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on https://t.co/xxBDOI9ApK and GitHub.
GitHub: https://t.co/2kXqhhvLsb
Blog and Guide: https://t.co/CYosNAQHva
Friendly reminder that you can fine tune 500+ open source models in a free Google Colab
You can even upload PDFs/CSVs and turn them into usable synthetic datasets.
1. Open the Google Colab below
2. Run the blocks to install Unsloth Studio
3. Choose a model
4. Upload a dataset
5. You're good to go!
And you can of course export your model afterwards.
@deepseek_ai Congrats DeepSeek on another epic release! Hopefully you guys will release smaller models for people to run locally. 🙏🐋
It's great that DeepSeek-V4.1 has 196B engram making it more accessible.
Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time!
The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥
GGUF: https://t.co/xIdNwm7CLQ
Guide: https://t.co/J2PwgMP6GZ
ICYMI, we're celebrating 1 BILLION+ downloads for @googlegemma 💎
How are developers actually using open models? @GoogleDeepMind’s @DynamicWebPaige caught up with devs and collaborators like @UnslothAI and @Qualcomm to hear how they’re building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
We made GLM-5.3-Flash run 3.3x faster locally!
Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction.
Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp.
Guide: https://t.co/axnDCrYeN7
GGUF: https://t.co/E73FKC7IKM
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
Hermes Desktop now sets up local models in one click.
It automatically reads your hardware, picks the best model for you, then downloads it and configures the runtime.
Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️
GGUFs can reach 170 tokens/s on a RTX PRO 6000.
MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change.
GGUFs: https://t.co/vXkjO3W0fj
Guide: https://t.co/LLMclyJTeL
Qwen3.8-Flash can now be run locally! 🔥
The 125B MoE model outperforms Claude-Opus-4.6 (Max).
Run on 75GB RAM via Unsloth GGUFs.
Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds.
Guide: https://t.co/LLMclyJTeL
GGUF: https://t.co/vXkjO3W0fj
⚡ Fine-tune and quantize AI models directly on NVIDIA Jetson.
Our new Jetson AI Lab tutorial shows how to use Unsloth and memory-efficient QLoRA to customize models, export quantized GGUF files, and run them locally with llama.cpp.
Follow hands-on examples for:
🔹 NVIDIA Nemotron 3.5 Lightning on Jetson AGX Thor
🔹 Qwen3.5-4B on Jetson Orin Nano
Start optimizing: https://t.co/hDUCiNnQcW
GLM-5.3 can now be run locally!
The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size).
Run on a 256GB Mac or RAM/VRAM setups.
GLM-5.3 is the strongest open model to date.
Guide: https://t.co/NLgb3CMB6A
GGUF: https://t.co/Sqwq3xjgo5
GLM-5.3 can now be run locally!
The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size).
Run on a 256GB Mac or RAM/VRAM setups.
GLM-5.3 is the strongest open model to date.
Guide: https://t.co/NLgb3CMB6A
GGUF: https://t.co/Sqwq3xjgo5
GLM-5.3 is now open-weight.
Our most capable model for agentic coding and cyber defense is now available to download, run, and customize.
Weights: https://t.co/v1IbWMXxg4
Tech blog: https://t.co/ekQkO83jCv
@Zai_org Congrats Z ai team, once again another amazing open-source release! 🥰 We're working on GLM-5.3 Unsloth GGUFs for those who can run it locally. https://t.co/bqGzanDcb3
GLM-5.3-Flash can now be run locally! ✨
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide: https://t.co/oItBKYNrl9
GGUF: https://t.co/E73FKC7IKM
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
The VRAM barrier is officially dead.
I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090.
21 tokens/sec decode. 364 t/s prefill.
no mtp. no dflash. no kv cache quantization!
We are running datacenter models on consumer hardware.
Tested on Ubuntu 22 | CUDA 13.0 | PCIe 4.0 x16 | 110 GB DDR4 System RAM with a continuous 28k prompt across all runs.
### The Benchmarks & Scaling
# 1. Hybrid Offload (-ncmoe 40 @ 80k Context)
Offloaded 40 expert layers to the GPU, pushing VRAM to the ceiling.
./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 80000 --port 8080 -v --fit off -b 4096 -ub 4096 -ncmoe 40
Prefill: 383.85 t/s | Decode: 22.52 t/s
Footprint: 23.85 GB VRAM | 97 GB RAM
# 2. Full CPU MoE Offload (-cmoe @ 80k Context)
Pinned all 512 expert layers to DDR4 RAM (-cmoe), keeping attention on the 4090.
llama.cpp flags: (Same as above, replace -ncmoe 40 with -cmoe)
Prefill: 355.72 t/s | Decode: 20.84 t/s
Footprint: 11.66 GB VRAM (12GB+ VRAM freed up!) | 110 GB RAM
# 3. The 180,000 Context Run
Prefill: 357.75 t/s | Decode: 20.98 t/s | VRAM: 15.6 GB | RAM: 110 GB
# 4. The 250,000 Context Absolute Ceiling
./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 250000 --port 8080 -v --fit off -b 4096 -ub 4096 -cmoe
Prefill: 364.29 t/s | Decode: 20.97 t/s
Footprint: 18.3 GB VRAM (Still ~5.7 GB of VRAM headroom!) | 110 GB RAM
### Key Insights:
-b 4096 -ub 4096: doubles the prompt ingestion from ~150 to 364+ t/s.
-cmoe Free Lunch: Shifting expert layers to DDR4 RAM slashes VRAM from 24GB to 11.6GB with virtually zero decode penalty (22.5 -> 20.9 t/s), enabling the 250k context ceiling.
Qwen 3.8 Flash-Next (UD-Q4_K_XL) is a massive 111.4 GB model split across 4 shards. To run this architecture, you must build from the experimental PR branch (#27742) by @danielhanchen:
git clone && cd llama.cpp
git fetch origin pull/27742/head:qwen-next && git checkout qwen-next
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DBUILD_SHARED_LIBS=OFF
cmake --build build --config Release -j $(nproc) --target llama-server
A single 4090 paired with 100 GB of cheap DDR4 RAM will comfortably serve production grade 125B inference.
While Qwen 3.8 27B (dense) still holds the crown for single 3090/4090 rigs, Flash Next proves 125B hybrid models are officially viable on consumer hardware.
Hugging Face GGUF link and complete performance telemetry graphs are dropped in the replies below.
GLM 5.3 Flash VS Qwen 3.8 Flash Next, which one takes the open weights crown this week?