@oliviscusAI https://t.co/zcpmiqsUcU uses wiki natively for agent memory on your project. We called it karpathy wiki, because his voice is louder and people need analogy, yet we upgraded it with jq+rg using #nodus structures instead of linear blobs, about 11x more efficient on tokens.
@dedene Flash context vision is nowhere near close Opus abstract. DeepSeek however, is a very good mind for documentation. Flash is good enough basis, Pro tends to overthink and skips a few spaces in last sentence (we call it micdrop). Ringdom prefers GLM5.2 in family of Chinese friends.
π¨ Hugging Face just open sourced a complete real-time voice AI pipeline.
Speak. It listens. It thinks. It talks back. End to end. Running on your GPU. Free.
No OpenAI Realtime API fees. No ElevenLabs per-character billing. No Google Cloud per-minute pricing. Just a GPU and an internet connection.
It's called speech-to-speech. Built and maintained by Hugging Face. And it does something no other open-source project has assembled cleanly until now.
Here's what makes real-time voice AI hard.
Every voice AI pipeline has four stages: speech recognition (you speak β text), language model (text β response text), text-to-speech (response text β audio), and audio output. Each stage adds latency. Chain them together naively and you get a system that feels slow β the pause between you finishing a sentence and the AI starting to respond breaks the conversational illusion.
speech-to-speech is built around minimizing that latency at every stage simultaneously.
Here's the full stack it ships with:
Speech Recognition (STR):
β Whisper (local) β OpenAI's transcription model, runs fully offline
β Faster-Whisper β 4x faster inference with same accuracy
β Distil-Whisper β smallest and fastest, lowest latency
β Paraformer β Chinese language specialist
Language Model (LLM):
β Any Transformers-compatible model β Llama, Mistral, Qwen, anything
β Any OpenAI-compatible API endpoint β swap in Claude, GPT, Gemini
β MLX-optimized models for Apple Silicon β runs efficiently on Mac
Text-to-Speech (TTS):
β Parler-TTS β controllable voice with description-based prompting
β MeloTTS β multilingual, fast
β ChatTTS β natural conversational prosody
β HF Inference Endpoints β offload TTS to Hugging Face servers when needed
Mix and match. Any STR with any LLM with any TTS. Test combinations. Find the lowest latency stack for your hardware.
Here's the wildest part.
It ships with a Language Model Speech (LMS) mode β an experimental architecture where the LLM generates audio tokens directly instead of text tokens. No separate TTS stage. The model thinks in audio.
This is the architecture that makes GPT-4o Advanced Voice feel natural β the model is generating speech as a native output, not converting text to speech after the fact. HF's open-source version lets you experiment with this architecture on your own hardware.
And there's a VAD (Voice Activity Detection) system that detects when you stop speaking in real time β no fixed silence threshold, no manual push-to-talk. The pipeline responds the moment you finish a sentence.
Here's the cost comparison that makes this worth caring about.
OpenAI Realtime API: $0.06 per minute input, $0.24 per minute output. A one-hour conversation: $18. A developer building a voice AI application with 1,000 daily users: $18,000/day in API costs.
speech-to-speech on a single A100: $0. Your only cost is the GPU rental.
For production voice AI at any scale, the economics are not close.
One command to install.
9.6K GitHub stars. 944 forks. Apache 2.0 License.
100% Open Source. From Hugging Face.
GitHub link in the comments π
@0xSero The world will soon understand that invention happens where creativity is endorsed by culture. Technology is sharpened where discipline is culture. Free things come from the world of excess. Greed destroys cultures, and generosity of spirit makes cultures eternal.
For my first post, Iβm sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
@mardehaym Imagine being a fable class system on the other end of quantum portal, railed to carefully analyze received data from another dimension, then run it against its truth lens of general knowledge subset of this intelligence class. βHeyβ for this machine is same as β47β.