a 118B model on one DGX Spark, in 4-bit, and vllm.cpp is ahead of vLLM on it
44.46 tok/s against vLLM's 43.10 on @poolsideai's Laguna-S-2.1. same output, token for token. default config, nothing to switch on
And most importantly: no python in the process. no pytorch. just one 66 MiB binary and a portable C ABI you can hook onto
and it's not just this model. qwen3.6-27b matches or beats vLLM at every concurrency we ran, cpu prefill is 1.18x llama.cpp on the same gguf, and we're 1.14x the fastest gguf engine we could find on deepseek-v4-flash!
vulkan is next, since that's what most of you asked for. after that the board is open, what goes on it?
Il corpo? È la nuova interfaccia del potere hi-tech
🖊️Su @startup_italia, un estratto del libro di @fabiolalli: https://t.co/xY0xXZ50f7
📖“Pelle digitale – Come l’intelligenza invisibile sta riscrivendo la realtà, il corpo e la società”: https://t.co/mq2aPKvHkJ
been writing a C++20 port of vLLM for a while now, and was quite silent on it, because I don't like to talk until I have results. So I was in "cave mode". Name TBD. Until then it's vllm.cpp
paged KV cache, continuous batching, prefix caching, no python at inference, reads GGUF natively
today it beat the reference engine on a 300B model running on one DGX Spark, so I guess it's becoming real
rf-detr.cpp: native C++/ggml (@ggml_org ) inference for @roboflow 's RF-DETR (my go-to for object detection and segmentation!) from the @LocalAI_API team.
All 11 variants (5 detection + 6 segmentation), running at PyTorch speed (slightly, ~8% faster on CPU benchmarks), without Python dependencies at f16.
Models available in @huggingface 🤗
Thanks to @ggerganov for ggml and @SkalskiP for rf-detr to make this possible!
🧵
@ggml_org@roboflow@LocalAI_API 44 GGUFs on @huggingface under mudler/rfdetr-cpp-*. Packaged also as a @LocalAI_API backend: `local-ai backends install rfdetr-cpp`, then just POST
/v1/detection API :) Apache-2.0.
Code: https://t.co/Xrp2xxD5oV
Models: https://t.co/nKJdZaFg6z
LocalAI ( @LocalAI_API ) 4.2.0 is out, just few numbers and facts:
- +392 commits ( we squash these 😄 )
- +11 Backends: voice and face recognition, vibevoice.cpp (from me), LocalQVE from @jichiep and among @sgl_project , @__tinygrad__ , @no_stp_on_snek 's Turboquant, ik_llama.cpp, sam.cpp from @el_PA_B
- Many new QoL improvements, increased sglang and VLLM support and hardening on distributed mode
- 16+ new contributors ! Thanks to the community!
LocalAI is all about give you flexibility to run the latest from the community, and ds4 support from @antirez is on its way!
This is the year of Local AI!
Say hello to vibevoice.cpp, @Microsoft 's Vibevoice in pure C++ with @ggerganov 's ggml (@ggml_org).
TTS and ASR (with diarization). CPU + CUDA + Metal + Vulkan via ggml backends. Quantized models live on @huggingface.
Built with ❤️ from the @LocalAI_API team
https://t.co/ZhFYd54uhz
🟣 Appuntamento venerdì 10 aprile alle 18:00 presso lo Spazio Attivo di Lazio Innova a Valle Faul #Viterbo per la presentazione dell’ultimo volume di @fabiolalli edito da @egeaonline#medioera
📚Pelle digitale: come l’intelligenza invisibile sta riscrivendo la realtà, il corpo e la società
📖Il nuovo saggio di @fabiolalli, ora disponibile in libreria e online, qui: https://t.co/M60NgYy35p
Ci sono momenti storici come quello che viviamo, che, mentre li attraversi, continuano a sembrare perfettamente normali. La loro forza non sta nell’annuncio e nemmeno nell’hype che generano. Sta nel ritardo con cui vengono percepiti.
https://t.co/bfMQV9gbPy
A call to open source maintainers. How I Built a 100% Local Autonomous Dev Team to help maintain LocalAI, and why you should too
Spinning up an autonomous dev team is no longer sci-fi. It is here, and it works 100% locally. Minimax really changed the game.
👇link to my blog
📚 La mente adattiva: pensare insieme alle macchine
📲 Pubblicato nella collana interamente digitale di Egea – Bookmark – il libro di @fabiolalli è disponibile online, qui: https://t.co/P7TfZKZxRZ
I’ve launched https://t.co/XJACXrHOUg. A multiplayer #Trivia#game where #AI generates the questions, the pace is fast, and knowing the answer isn’t always enough, you need to read the game. It works with friends, family, colleagues.
👉 https://t.co/nNzmrZIXPF
Avete mai provato a chiedere all'AI di dubitare di se stessa, e anche di voi? Io lo faccio spesso, e con "La routine del dubbio". Ne ho scritto qui una Interferenza #33 https://t.co/9fHwcjQHxl
I just finished reading #SuperAgency. Remarkable.
Imho, this is a #book worth reading because it connects many key ideas that help shape a broader perspective on what lies ahead.
https://t.co/giPm6t6TsA
@dnepimolineris Purtroppo non è un tema di oggi Diego. Invidia, mediocrità, indifferenza e irresponsabilità sono gli elementi che non fanno crescere la collettività.
✳️ Alle 11:30 #EtaBeta con @MaxCerof
🤖 Lenti smart e voci sintetiche: niente schermi, il computer è nello spazio con @fabiolalli, Gennaro Coppola di #OneMore e con un intervento della voce sintetica di #ChatGpt
🎧 Ascolta su #Radio1 e @raiplaysound
👉 https://t.co/N6Kb40RFkH