The world’s smallest Transformer-based TTS model?
We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint.
Flux 3 Prompt: archival footage of the Wright brothers’ first flight in 1903. No sound
Flux 3 understood the simplicity of the moment. A fragile aircraft, an open field, a small group of witnesses, heavy clothing, rough wind, and a camera positioned far enough away to observe rather than dramatize.
The result did not feel like a cinematic reenactment. It felt like an improbable machine briefly lifting off the ground while someone happened to be recording.
Flux 3 is especially impressive when it turns a familiar historical photograph into a living moment.
Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro.
>It can clone a voice from 5 seconds of audio
>2x faster than Cartesia & 1/6th the cost of Eleven Labs
>most expressive model with word level control over emotion, intonation, pacing etc
We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production.
If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free.
Book a demo: https://t.co/vHkyZf9JoG
To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.