Open source!
We’ve fine-tuned Whisper models to handle arbitrary audio chunks and compressed them with ANNA.
This enables optimized inference on @NVIDIA and @Apple devices with 2x faster time-to-first-token and full streaming support.
Clone the repo and start building
The smallest checkpoints for Gemma 4 E2B and E4B for local inference. Results for E2B:
size: 9.3 GB → 1.4 GB
speed: 113 tok/s on Apple M3
quality: -3% on ifEval
runs with: MLX, llama.cpp (coming)
Pareto-optimal, open source! Links to the blog post and GitHub repo ⬇️
@GoogleDeepMind@lmstudio@ollama@huggingface@ggerganov
Beyoncé heard cursing. TheWhisper heard Arsenal.
The fastest Whisper in the world.
Open-source real-time ASR.
Top 5 on OpenASR benchmarks.
1800 RTFx.
Built for live captions, transcription, and voice apps.
See the repo
For AI engineers, latency is product.
Wan 2.2 in Elastic Models now generates 5s of video in 34s on H100. Elastic Models is a library of accelerated open-source models.
Also new: TheWhisper at 1800 RTFx on a single H100 and instant FLUX LoRA switching.
Try it
How do you make text-to-music run in real time in production?
The model has to keep audio generation ahead of playback.
Our new case study with @MireloAI shows how inference optimization delivered up to 2.4х higher throughput.
See the full case study ↓
Proud to team up with @brilliantlabsAR and @neuphonicspeech on Halo’s on-device privacy engine.
Coming to Brilliant Labs’ Halo smart glasses: real-time voice + vision, POV stays private.
ANNA + GPU/NPU SDK + memory manager for wake word, STT, TTS, diarization.
SDK demo 👇
Are you a big fan of jacket potato?
This is an open-source, real-time multilingual ASR for live speech.
It stays robust in heavy noise – even at SNR 0 dB.
That’s why it understands speech where people struggle to hear.
Use it for transcription, research, and multilingual apps
We integrated @nvidia cuDNN Paged Attention into our Elastic Models library at @TheStageAI.
It’s already running on B200: INT8 Llama-8B reaches ~200 tok/s per sequence @ batch 16 (≈3.2k tok/s total) – and we’re still tuning. Elastic Models keeps HF APIs, but can deliver up to ~4x faster inference.
Details:
https://t.co/JIa80XczH9
TheWhisper is now SOTA for Arabic, Chinese, and Hindi.
A single open-source speech-to-text model handling multiple languages and noisy audio.
On Jetson Thor it’s faster than torch.compile: RTFx 147.6 vs 106.2, TTFT 0.039s vs 0.122s
Read the full tech report →
Up to 8x lower Time to First Token and 4x higher Real-Time Factor than other Whisper libs.
That reduces latency and lets you run more streams per GPU, like real-time captions and voice calls.
You get these gains in TheWhisper, our open source self-hosted speech-to-text engine
We posted a tutorial on real-time on-device transcription using TheWhisper, our optimized open-source model at @TheStageAI.
It runs short windows on Apple Silicon with sub-150 ms latency and about 2 W power.
Build fast speech apps on your Mac
We believe that everyone will become a model builder! That's why we are creating an automated acceleration and deployment stack which undestands ai engineers needs
Open source!
We’ve fine-tuned Whisper models to handle arbitrary audio chunks and compressed them with ANNA.
This enables optimized inference on @NVIDIA and @Apple devices with 2x faster time-to-first-token and full streaming support.
Clone the repo and start building
I build ANNA. ANNA builds optimized compression for faster inference.
Seeing models breathe on @Modal always feels special. Follow our simple guide for optimized, production-ready diffusion model deployment.
We’ve made it easy to run text-to-image models on @Modal with the speed you’d expect from top inference providers.
Follow our quick guide to deploy containers with an @OpenAI compatible API and get 2× faster performance.
Big thanks to @MireloAI for the soundtrack magic 🎶
Great result in MLPerf Inference v5.1 (@MLCommons) .
We used ANNA, our automated NNs acceleration stack.
It accelerates inference while keeping output quality high, delivering fast, stable, and production-ready performance for real workloads.
Enjoy ↓
Validation is a key step when compressing or accelerating models.
It shows if the network still performs well.
Our research team @TheStageAI shared evaluation methods for sharpness, tone, color, object placement, and more