Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro.
>It can clone a voice from 5 seconds of audio
>2x faster than Cartesia & 1/6th the cost of Eleven Labs
>most expressive model with word level control over emotion, intonation, pacing etc
We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production.
If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free.
Book a demo: https://t.co/vHkyZf9JoG
To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.
@jxnlco maybe it auto routes to be more efficient but we can still say in the settings we prefer saving tokens or better quality in general, to nudge it in a general direction for our budget
If you have a few minutes today, watch this short film. I highly recommend it. You won’t be disappointed.
It’s one of the winners of AIF and a film that everyone I meet keeps mentioning. It’s mesmerizing, with a brilliant concept and exceptional execution.
Costa Verde. Directed by Léo Cannone & New Forest Film.
A fundamental problem with extending Codex/Cowork/Code to all knowledge work is that they remain very "software-brained" where the end result (the software) is what is important & that code serves as a source of truth.
For a lot of other knowledge work, the process is at least as important as the outcome. This includes researching what is known, an exploration of alternatives, failed efforts, prototype branches, experiments, etc. All of those things are valuable, so you cannot use the PowerPoint at the end the way you can use a codebase, nor is progress on a to-do list sufficient context post compaction. You work in learning loops, refining your perspectives as you go.
In some ways, this makes long-running models like Fable hard to use for deep knowledge work, since they are designed to deliver product to you in the end. You can prompt your way around this problem, but everything about the Codex and Code harnesses want you to be a software developer and you have to fight them. There is a real disconnect between how a manager or analyst thinks about problems and how the agentic software tools approach solving them. Addressing this is critical to breaking out of the coding niche for these tools.
@emollick@TheRealAdamG In my experience normal text 4o is even better than whatever AVM gives you. I miss the option for a slow and unnatural but smart voice mode. Specially now in the winter.
@nickaturley Not sure what the UX could look like for this, but my ideal “temporary” solution would be an easy switch, like in the model switcher. AVM pronunciation and tone improved with the last update but many times i’d rather have a hands-free gpt-5 with lag than a dumb but cool voice