GLM-5.3 can now be run locally!
The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size).
Run on a 256GB Mac or RAM/VRAM setups.
GLM-5.3 is the strongest open model to date.
Guide: https://t.co/NLgb3CMB6A
GGUF: https://t.co/Sqwq3xjgo5
We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy.
Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks.
We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM.
Blog: https://t.co/tHsBexyh2K
GGUF: https://t.co/xIdNwm7CLQ
Qwen3.8-27B can now be run locally! ✨
Run on 17GB RAM via Unsloth Dynamic GGUFs.
Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants.
GGUF: https://t.co/xIdNwm7CLQ
Guide: https://t.co/J2PwgMP6GZ
Introducing Unsloth Desktop 🦥
The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux
• Supports MLX, diffusion image/video, audio, GGUF
• Connect Claude Code and Codex to local LLMs
• 50% more accurate, self-healing tool calls + sandboxed code exec
• Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac
• Train models 2× faster with 70% less VRAM
• Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF)
• Use Unsloth’s OpenAI-compatible API and cloud models
• Securely deploy LLMs remotely and access anywhere
Unsloth Desktop is now available on https://t.co/xxBDOI9ApK and GitHub.
GitHub: https://t.co/2kXqhhvLsb
Blog and Guide: https://t.co/CYosNAQHva
Kimi K3 can now be run locally! ✨
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio + 128GB RAM device.
Kimi K3 is the strongest open model to date.
Guide: https://t.co/1mVwOMLpDW
GGUF: https://t.co/bt1c1ADdCZ
Introducing Unsloth for AMD 🚀
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on 3GB VRAM
GitHub: https://t.co/2kXqhhvLsb
Works on Radeon, Instinct, Ryzen and data center GPUs with up to 2× faster with 70% less VRAM and no accuracy loss via our custom Triton kernels and math algorithms. We also support optimized ROCm builds for GGUF & Safetensors inference.
Unsloth is an open-source local UI for faster LLM training and inference, with tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment. Connect local models to Claude Code, Codex agents and run the latest Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4 models.
🔗Blog + Guide: https://t.co/U9LqyRjFdj
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU.
Qwen3.6-27B NVFP4 runs on 24GB VRAM.
35B-A3B can hit 17,561 tok/s (B200).
We also improved accuracy, tool calling, agent use, and looping.
Guide: https://t.co/EEQIlFrR0c
Qwen3.6 NVFP4: https://t.co/RWflncpLPJ
We collaborated with @NVIDIA to make LLM training ~25% faster and wrote a guide on how we did it! 🤍
Huge thanks to the NVIDIA team for their open-source contributions to Unsloth, helping make training more accessible for everyone.
Introducing SubQ - a major breakthrough in LLM intelligence.
It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA),
And the first frontier model with a 12 million token context window which is:
- 52x faster than FlashAttention at 1MM tokens
- Less than 5% the cost of Opus
Transformer-based LLMs waste compute by processing every possible relationship between words (standard attention).
Only a small fraction actually matter.
@subquadratic finds and focuses only on the ones that do.
That's nearly 1,000x less compute and a new way for LLMs to scale.
I’m doing these for the people in the trenches like me, those without a datacenter at home
We don’t see enough benchmarks using quantized models, yet that’s the reality for 99% of users: a single GPU or even an older computer
Most benchmarks you see on X are run in full precision (which makes sense for a fair comparison), but 99% of actual users are going to run GGUF or MLX versions; because those are the ones that actually fit on their hardware
I’m currently running UD_IQ3-XXS GGUF (Qwen3.6 35B and 27B) on a RTX 5080. It’s an aggressive quant, but the only realistic option for 16 GB VRAM owners. Attached a recent analysis from @bnjmn_marie that shows even Q2_K_XL is "surprisingly recommendable" with the only real downside being that it will spend more tokens for the same tasks, as shown on the graph.
We're seeing much less precision loss than we used to see with previous generations at the same quant levels. Bartowski and @UnslothAI are on a crazy run and their GGUFs at lower quants are suitable for daily use today imo. If it's your only option, you don't have the choice anyway
Unsloth Dynamic v2 quantizations are pure wizardry and make a massive difference, it shines in Q3/Q2. Another huge win for the lads in the trenches with limited hardware 🏆
Getting these results on a single 16 GB VRAM card with a heavily quantized model is honestly insane to me
The “dumb AI” era for single-GPU users is finally coming to an end 🔥
Kimi K2.6 can now run on CPU, GPU and SSD setups! 🔥
We shrank the 1T model to 340GB via Dynamic GGUFs where important layers are upcasted.
Run at >40 tok/s on 350GB RAM/VRAM setups.
Run full precision on 610 GB.
Guide: https://t.co/nGGHiZhNPG
GGUF: https://t.co/nK6TkXRITf
Introducing Unsloth Studio ✨
A new open-source web UI to train and run LLMs.
• Run models locally on Mac, Windows, Linux
• Train 500+ models 2x faster with 70% less VRAM
• Supports GGUF, vision, audio, embedding models
• Auto-create datasets from PDF, CSV, DOCX
• Self-healing tool calling and code execution
• Compare models side by side + export to GGUF
GitHub: https://t.co/2kXqhhvLsb
Blog and Guide: https://t.co/ENuTWal5AA
Available now on Hugging Face, NVIDIA, Docker and Colab.
You can now fine-tune embedding models in our free notebook!
Improve retrieval and RAG with better semantic search & similarity.
Unsloth trains 2x faster, 20% less VRAM, >2x context & no accuracy loss
Blog: https://t.co/H9NfyU2MsJ
EmbeddingGemma (300M): https://t.co/MBu45lcpGr
You can now do reinforcement learning training with 7× longer context and no accuracy loss, via our new batching algorithms.
Long reasoning chains in RL are costly, but now we enable you to train gpt-oss with GRPO & reach 380K context on a 192GB GPU.
https://t.co/io8OUqGIbn
You can now train LLMs 3× faster with no accuracy loss, via our new RoPE and MLP kernels.
Our Triton kernels plus smart auto packing delivers ~3× faster training & 30% less VRAM vs optimized FA3 setups.
Train Qwen3-4B 3x faster on just 3.9GB VRAM.
Blog: https://t.co/kL6JM6skH1
You can now fine-tune DeepSeek-OCR with our free notebook!
We fine-tuned DeepSeek-OCR, improving its language understanding by 89%, and reduced Character Error Rate from 149% to 60%
Blog: https://t.co/wMyAualKbr
GitHub: https://t.co/2kXqhhvLsb
Colab: https://t.co/dtD3BSAFfa