When several people talk at once, a transcript can get messy fast.
Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
@GoogleDeepMind I strongly hope to see Google Meet's transcription accuracy improved, and a lite transcribe model cheaper than the current one would make integrating it into Meet economically viable.
Your job may be surviving on inertia. Your resentment won’t extend the runway. What do you actually want to do with the intelligence now at your disposal? https://t.co/vikawjCmZD
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command.
Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec.
It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle.
Get the code and instructions here: https://t.co/E3agzrW2hZ
we used ChatGPT 6 Astra to do Blender previs on stream today. funny enough, more detail in our Blender previs made Seedance worse.
the posed people gave it stiff motion to copy. we replaced them with blocks and used image refs for the look.
comparison from today’s ImagineArt LIVE. anyone else getting this?
I’m significantly older than you. I started coding in the late 60s. My current strategy is to not read any of the code written by my agents. That’s the only way I can take advantage of their productivity. What I do instead is to surround the agents with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, test coverage, and a plethora of others. In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.