1.45GB on disk and it cloned my voice on a CPU.
Qwen3-TTS is in mainline llama.cpp now. Two GGUFs load together, not one:
the talker and the audio tokenizer. At Q4_K_M that is 1.2GB + 255MB.
Apache 2.0. 24kHz mono out.
Every repost of this says "no quality loss at Q4." I gave it a 9 second
reference clip of my own voice and 40 sentences, English and Ukrainian.
model peak RSS 9s of audio real-time factor
1.7B Q8_0 6.7 GB 11.4 s 1.27x
1.7B Q4_K_M 4.9 GB 8.9 s 0.99x
0.6B Q4_K_M 3.1 GB 4.2 s 0.47x
Q4 on the 1.7B is real time on a normal desktop CPU. The 0.6B is twice
real time. No GPU touched at any point.
Where it broke: Q4_K_M held my timbre fine but flattened prosody on
questions. 6 of my 40 sentences came out with a falling intonation where
the reference clearly rises. Q8_0 got all 40. So: narration, take Q4.
Dialogue or anything conversational, the extra ~900MB is not optional.
Both files, or it will not load:
huggingface-cli download Serveurperso/Qwen3-TTS-GGUF \
qwen-talker-1.7b-customvoice-Q4_K_M.gguf \
qwen-tokenizer-12hz-Q4_K_M.gguf --local-dir models
customvoice is the mode that clones from a reference clip. base only does
the named speakers. That one word cost me twenty minutes.
https://t.co/xuFzUpFA3K
You don't need a GPU for fast studio grade voice cloning anymore.
Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution.
Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware.
Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU:
# Architecture & Model Setup
Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks.
# Real-World Memory Footprint
- Baseline RAM: 1.6 GB system idle.
- Peak Generation RAM: 8 GB RAM during active voice synthesis.
- Requirement: Any basic machine with at least 8 GB of system RAM can run this easily.
# Real World CPU Benchmarks
- Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds.
- Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!).
# Zero Shot Voice Cloning Quality
Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights.
To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!)
launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU.
Native C++ audio models are making edge based, offline AI voice agents a reality.
Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below!
Which models have you been running on your CPUs? What CPU hardware are you using for local inference?
@omarsar0 learning is great but framing it as a passion can backfire
when it becomes a grind to hit a daily quota it stops being curiosity and starts being homework?
@GaryMarcus@deanwball disagreement is fine but roasting one guy isnt a line-item veto
the white house deplatforming individuals is a much bigger leap than calling out bad posts