We have added Cluster inference support to Qvac Fabric. You can cluster together multiple machines and split the inference across them. Supporting both tensor and layer parallelism. Here 2x DGX Spark running Deepseek-v4-flash-0731 over RDMA.
https://t.co/FQkHF1e8Fg
Latest @qvac Fabric work added CoreML support to make Parakeet insanely fast. You can transcribe 25 mins audio in nearly 16 seconds on an Iphone 16.
check https://t.co/c3ZZzg8uci
QVAC Fabric now runs Parakeet's encoder through Core ML on Apple silicon, which moves it from the GPU to the Neural Engine.
On a MacBook Air M2, Parakeet TDT 0.6B v3 went from 78x real time on the Metal GPU to 164x through Core ML.
Parakeet Redux is compressed from NVIDIA's Parakeet TDT 0.6B v3, so we ran it against the original, which QVAC runs in Q8.
Both transcribed the same two hours of audio on the GPU of a MacBook Pro M5 Max.
QVAC took 22.3 seconds, 323x real time, and Redux took 44.2 seconds, 163x. Both times include loading the model.
QVAC SDK 0.20 is live. ๐
We added support for:
- ABot-World: generates a walkable world from a single photo
- MiniMax-H3: create video with synchronised soundtrack
- Video models can now run on GPUs with less VRAM
- TranslatePsy-AfriSLM: translate between English & 19 African languages on-device
- ACE-Step can now generate entire songs from a single sentence & more new features
We also merged the latest changes in llama.cpp into QVAC Fabric, to ensure continuous compatibility.
Learn more in the thread below.
Complete release doc here: https://t.co/N8hcETdJHF
The QVAC audio stack is continuously optimised. We compared its performance against several other open-source solutions, and QVAC comes out fastest on transcription, speech synthesis and music generation, across every GPU lane we measured.
If you want to build fast apps with local audio AI that fit on consumer devices, you should definitely check QVAC out.
https://t.co/3AcGpL3QRE
Deepseek v4 flash is now available on Fabric. We added Vulkan and Metal support, so you can run it on any gpu. Performance improvements come from optimized fused kernels. Go and try the new @deepseek_ai released version '0731' on https://t.co/OzCiML7whP
DeepSeek V4 Flash on Strix Halo just got a major Vulkan boost on Fabric.
Fusing different ops improved prefill way more than vanilla llama.cpp, closing the gap with @antirez ds4 ROCm implementation.
Coming soon to https://t.co/FQkHF1e8Fg.