Barely finished testing Gemini 3.5 Transcribe, and here we are with a new model. This time I am testing @Meta Muse Voice Transcribe, with @ai_coustics Quail VF 2.2 and Tyto.
Try it out here: https://t.co/ZnbLGKDkRw
Wow
“Make Red Alert 2 (YR) compile and run natively on iOS & macOS.”
Codex (GPT-5.6 Sol) worked for 26 days, analyzed the EXE and rebuilt the entire game in 624K lines of C++
This is by far the hardest task I’ve ever given Codex as the game's code was never released and is believed to be lost.
>
@thorwebdev It's pretty solid! Kudos team! It can even clean up the interfering speech by itself.
Tried to give it some hard time, but it did pretty well. 💪
https://t.co/wkGNId0CeL
@Google's Gemini 3.5 Transcribe is superb but audio input is the foundation, and plenty can go wrong before the model ever sees it.
I put it to test with @ai_coustics Quail Voice Focus 2.2 for the audio enhancement, and Tyto for input insight.
The same recording goes through Gemini both ways.
Robots don't have to be scary!
Microduck discovers Reachy Mini. Give me your best scenario for these two and I'll see what I can do.
Don't push me too hard though, or I might end up making a whole movie.
Robots don't have to be scary!
Microduck discovers Reachy Mini. Give me your best scenario for these two and I'll see what I can do.
Don't push me too hard though, or I might end up making a whole movie.
SF voice AI + robotics crowd: save the date. 👨💻
Audio Layer 3.0 lands in San Francisco on Sept 15, co-hosted with @tavus.
What to expect? Networking, food and drinks, and a panel on where voice AI and robotics meet, and what that means across the stacks. On the panel (joining our co-founder Fabian): @quinnfavret, @chenosaurus, @ConstanceGriso and @aliattar.
Hope to see you there!
RSVP: https://t.co/AdKbujP5eZ
@Google's Gemini 3.5 Transcribe is superb but audio input is the foundation, and plenty can go wrong before the model ever sees it.
I put it to test with @ai_coustics Quail Voice Focus 2.2 for the audio enhancement, and Tyto for input insight.
The same recording goes through Gemini both ways.
A robot doesn't live in the same acoustic world as voice agents. Navigating a three-dimensional acoustic sound field is far more complex than operating on a monophonic audio stream.
As human beings, we spend a lifetime learning to navigate, recognize and differentiate all of these auditory cues. We know where a sound comes from, which voice is meant for us, and what to ignore. Current robotics applications, however, are mostly guided by sensory input based on vision and movement, while voice, and audio in general, is still the overlooked layer.
At @ai_coustics we believe that training those auditory cues into an audio intelligence layer for machines and applications will open the door to more natural conversations and real spatial awareness, as well as entirely new physical AI applications, all based on the same audio understanding we rely on every day.
We are co-hosting Audio Layer 3.0: Voice AI x Robotics with Tavus at their SF office, Tuesday, September 15th.
I'll be joined on the panel by @quinnfavret (Co-founder at @tavus), @chenosaurus (GM Robotics at @livekit), @ConstanceGriso (Chief Growth Officer at @GradiumAI) and @aliattar (Founder at @lightberry).
Join us for some demos and a conversation across the voice AI and robotics stack, plus a chance to catch up with familiar faces and meet new ones over food and drinks.
If you're in SF and working on Voice or Robotics, we would be happy to see you there. 🙂
🗓️ RSVP here: https://t.co/Q9Q5IjtU5c
We couldn't even hear each other, so how could the voice agent hear me then? Tested it with @ai_coustics Tyto, and here is how it worked.
demo: https://t.co/EGnMlR0Sbd
docs: https://t.co/k98Q3qn3tY
Audio 👏 input 👏 breaks 👏 your 👏 agents 👏
→ before the signal even reaches the core stack!
We heard it from many companies - they have no way to measure just how much interfering voices and noise are breaking their stacks. And that's even without mentioning other audio artifacts.
So we built Tyto - a lightweight audio insight model to help you understand your agent input before it causes you lost revenue.
If we got your attention, here is a simple way to batch diagnose hundreds of calls you already have and find potentially fatal patterns. @GrungeCoder walks you through it below.
▶️ https://t.co/3nq2e8ETE2
--
🔗Post-call batch analysis with Tyto: https://t.co/FFdguT5oO6
🔗Link to analysis dashboard: https://t.co/UTlBH29GFB
Phonon, our 100M-parameter on-device TTS model, and our in-house on-device ASR handle the speech, with @ai_coustics cleaning the mic input before the ASR. Gemma 4 from @googlegemma does the reasoning.
All of it runs at once on one machine.