@NVIDIAAI@huggingface But how many speakers ?
For audio (specially Steaming audio) --- how many speakers can be individually identified is also a very important point. If this is being used for a meeting intelligence system - having more than 8 speakers will cause a issue with the transcript ownership
@abdullah_twt23@DevanshuXi I don't. I mostly do AI and dev on daily driver, and occasionally explore dsa concepts - mostly for learning though.
Don't know if I would call it managing these 3.
@byteHumi For codex particularly - it loads the entire chat and lot of unused stuff - that the agent isn't really using -- all onto ram. So u can actually just ask claude to take a big chunk, compress and save it.
This is if your session has been going for too long.
@SarvamAI@c_engines what do you think ?
I understand the different factors that create limitations -- reaching a state of optimum accuracy with limited number of speakers. But still, nothing that can be done ?
Can anything give a decent accuracy in Speaker Diarization with Streaming(Online) Audio - where like 50-100 people are there and speaking concurrently - suppose in a scenario of a Meet/Zoom - only via audio - so it's platform/usage agnostic ?
What about u @SarvamAI ?
@HowToPrompt__ I can think of a very intersting usecase for this. Infact quite a few interesting usecase - but the scale would be tremendous - if we wanna embed loads of content.
Minimax H3 Max has generates video faster than you can watch it so I hooked it to a twitch livestream! Now you can watch infinite interdimensional cable - link to the stream below