parakeet-redux is a really interesting beast. You can extract the VAD head and then, you discover that has better recall than Silero. Despite the footprint compared to it, offline might have other uses.
@vikhyatk cooked hard here, congrats for the release, and thanks for working in the open!
By the way, all Parakeet-redux models are now in parakeet.cpp (links below). And you can use parakeet now as standalone VAD with both.
We mapped out voice AI in healthcare: 52 companies across 4 workflows.
The same technology is doing very different jobs depending on who is speaking: a patient booking a visit, a care team following up after discharge, a clinician documenting an appointment, or a billing team calling an insurer.
We grouped the companies by workflow:
1. Front office - patient access
@assort_health, @hyro_ai, @elise_ai, @HelloPatient, @Relatient, @phreesia, @Zocdoc, @notablehealth, @actiumhealth, @Infinxinc, @confidohealth, @keonahealth, @parakeethealth, @eClinicalWorks
These products answer patient calls, manage appointments and referrals, and route requests that need staff attention.
2. Middle office - care outreach and follow-up
@hippocraticai, @EllipsisHealth, @cipherhealth, @lumahealthhq, @Artera_io, @concurrencehq, Sagecare, @clarionhealthai, Kouper Health, @innovaccer, @Lumeris, Cadence, @Qventus, Flagler Health
Here, voice helps teams reach patients about screenings, missed visits, care plans, and needs that surface after discharge.
3. At the point of care - clinical documentation
@AbridgeHQ, @SukiHQ, @nabla_ai, @AmbienceAI, @DeepScribeAI, @Microsoft Dragon Copilot, @OracleHealth, @CommureOS, @HeyEpic, @athenahealth, Freed, @doximity, @knowtexai, @sullyai
These tools use speech to turn clinical conversations into draft notes or structured documentation for clinicians to review.
4. Back office - payer and revenue cycle calls
@InfinitusAI, @R1RCM, Raintree Systems, SuperDial, @TennrOfficial, VoiceCare AI, @AdonisRcm, @CedarNY, Collectly, Prosper AI
Their voice workflows tackle benefits verification, authorization, and claim follow-up. These calls often involve phone trees, hold times, and payer-specific rules.
The companies may all work with voice, but the mistakes they need to catch are different.
• A scheduling agent might book the wrong type of visit.
• A follow-up agent might miss something a clinician needs to hear.
• A payer agent might come back with the wrong coverage details.
Those are the kinds of calls we test at Cekura. We check whether the agent gets the job done, including when the conversation takes an unexpected turn.
Bookmark this and share it with someone who would find this list useful.
Also, let us know if we missed any companies that should be on here.
Since everyone is open-jevving. I set ML Intern (on hugging chat) to work. After a few experiments and a little guidance and we did ok.
Qwen3.5-4B fine-tuned as a decision scorer: given some context, a question, and possible answers, it scores each answer. Temperature scaling then calibrates the probabilities, so its confidence better reflects how often it’s right.
The inference setup is half of the thing, so needs a demo.
https://t.co/z7M4zlfofe
As the first publicly disclosed agent cyberattack victim, we've had a front-row seat to this new risk. I formalized my thinking about it below.
I'll be in DC tomorrow to share more with policymakers and at decoded summit by @politico!
“no one is forcing anyone to code with ai”
if you want to be competitive on job market you need to embrace agentic coding. otherwise you are 10x less productive. just a reality.
YuE2 is here! A 3B parameter song general model that rivals Suno 5.5
YuE2-3B works by first writing a symbolic score, then rendering it.
the [Instrumental] really is instrumental - and then the lyrics are composed on top of it
▶️ https://t.co/Dy3k4u83Ly
WanGP v13.00 — It’s Your Lucky Day! 🚀
��� A faster, refreshed UI with 5 themes and voice dictation for prompts.
🌍 Start a generation on your PC, follow progress and manage jobs from your phone or another PC.
📊 See what’s happening at every stage, with cancellation during preparation too.
📁 Organize media into project workspaces that stay saved across restarts.
📱 Deepy now has a mobile web app: chat, upload photos or recordings, and follow your results from your phone. Add it to your home screen for quick access.
🎵 YuE2 turns your lyrics and musical style into full songs with vocals.
🎙️ AuK generates speech, clones voices, edits spoken words and cleans up recordings.
https://t.co/5MQjylADRk
YuE2-3B turns lyrics into complete songs through an editable musical score.
🤖 https://t.co/i9imLmzNYt
📄 https://t.co/Ud2W6G7IZx
🏆 YuE2 best-of-8 reaches 6.9632 SongBench Avg on WildSongBench, the highest score across the evaluated open and proprietary systems. Suno v5 scores 6.8721.
🎼 It plans melody and chords before generating vocals and accompaniment. Edit the score or supply your own ABC notation.
🎙️ The same checkpoint supports direct song generation, zero-shot covers, and agent-guided edits to the score, style, and lyrics.
⚡ Generate 48 kHz stereo audio locally on a 24GB GPU without quantization.
📜 Weights: CC BY-NC 4.0. Code: Apache 2.0.
"First, we need to put aside the tools we’ve used for years and start from scratch.
Every project should start by trying really hard to solve the problem with existing tools."
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia?
These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach.
The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date.
The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant.
At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on “old” problems that have a long history. This history is based on assumptions about how the “vision problem” will be “solved”. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems.
So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models.
Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field.
If we want there to be a “field” of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section.
Concretely, I think papers should include a new section analogous to “Related Work” where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models.
I'm interested in your thoughts.
Vaani timestamps real background noise to the millisecond: 122+ hours, 106,892 events, 58 Indian languages, and verified/unverified tiers. It turns robust speech into a testable claim. https://t.co/C0z3Ddqv77
Voice AI benchmarks.
All in one place.
I built this: https://t.co/cBOlQ9jhul
One place where you can find the best benchmarks for your TTS, LLM, STT, EoT, etc.
(with a cool UI just because it's fun)
It's all open source so feel free to open a PR to add a benchmark or improve an existing entry.
hope you like it!
Astra can do segmentation
this is pure VLM result. no expert models (like SAM) were used
No other VLM even comes close to this quality
- high effort
- avg input tokens / image: 2,052
- avg output tokens / image: 4,685
- avg cost / image: $0.255
- median time / image: 78.3 s
↓ GPT-6 Astra segmentation deep dive