@VaibhavSisinty One of the most dangerous models is video models cause they can easily fool older flocks and these people have money
unlike the youth that don't have money but won't get fooled
There is a subtle architecture shift happening in voice AI.
The voice stack is becoming part of the agent's execution loop.
@cartesia is combining the listening and speaking paths around that loop.
Sonic-3.6 turns text into speech (90ms latency) and Ink-2 turns speech into text (100ms transcript latency), faster than anything else streaming.
now holds the #1 spot for both speaking and listening models.
Becasue, voice agents are one of those systems where 100ms in the wrong place is very noticeable.
Cartesia is attacking both sides of that loop at once.
super interesting work by @krandiash and team.
Jensen Huaan: The data-center industry failed to engage communities early enough.
"water consumption, that's a myth now. these data centers use less water evaporative. They don't use evaporation anymore. It uses less water than a swimming pool evaporating all year long."
newer data centers recycle their cooling water, use much less evaporation. He says operators can also add power generation, make the grid stronger, lower electricity costs, and create the economic demand needed for solar, hydro, fission and fusion projects.
----
From "CBS Sunday Morning" YouTube channel, (full video link in comment)
GOOGLE CEO SUNDAR PICHAI: "IF YOU DON'T LEARN HOW TO ORCHESTRATE AGENTS NOW, YOU'LL SPEND 2027 CATCHING UP TO PEOPLE WHO STARTED TODAY."
30 minutes on why the best engineers stopped writing code line by line and started orchestrating agents instead.
Most people think building an agent requires an engineering degree.
It doesn't.
It requires one guide and one afternoon.
Watch the interview. Then read the article below.
One guide. One afternoon. That's all it takes.
Nobody Reported OpenAI’s Biggest Announcement 🚨
It reads like a boring company update. It is actually a countdown.
Their agents already do three days of work for every one human day, at six hundred dollars a day per researcher.
More than half of those tasks still need a human to step in.
For now. March 2028 is the month they expect that to stop being true.
One of the stranger assumptions in AI is that the same model architecture should work equally well for thinking and interacting.
A voice model has a weird job, it cannot just produce the right answer. It has to keep up with a person while the sequence keeps getting longer.
In voice AI "speech" may not the right abstraction for the hard part, rather continuous state could be it. Here, Cartesia's founder talking how they came into the space from sequence modeling rather than speech research, which is probably why they focused so heavily on state space models.
---
(Full video on “The Neon Show” YT channel, link in comment)