Hey Bangkok, walking around the city is my favorite activity and I'm sure many of you like it too.
I want to preserve some routes and share it (many places around are changing so fast) as well as I want to learn other people favorite routes or places. 1/
We're hiring a Research Engineer at @Arsenal β½π΄βͺ to work directly with our Men's First Team!
We're building state-of-the-art AI models for the football domain. This role will focus on building the application layer for our research to advance coaching and analysis workflows.
have little idea what's happening, but it's very entertaining... "teaching a big language model to hear and eventually speak." keeps me pressing Y to everything it asks.
Thank you, now back into the terminal :)
I was digging around, found https://t.co/4DFXQALDwZ you might like to explore, same core thesis as your's Gemma approach (a single autoregressive LLM decides listen/speak/stop, no external VAD or turn-taking), but different design point.
This is what people call a high-value reply! Thank you, it's concrete enough to validate several things I've been wrestling today.
One tiny engineering clarification: per 80 ms step, what conditions the depth transformer - Gemma's hidden state at the 2nd 40 ms audio token, a pool of the two 40 ms tokens, or a dedicated frame token?
Sleepless nights are back... teaching GPUs to generate speech sounds correctly.
Think of it like teaching someone to make the right mouth shapes and sounds - getting the pronunciation perfect.
What a time to be alive.
@matthen2 That idea of audio in / out without the encoder occupied me this weekend. Good thing training isnβt that long/expensive + plus some labs opened their datasets as a bonus to ease tedious part.
@LorenzoIotti2 came here with the same question! Was listening to founder(s) of Sesame and year ago they were saying that it's going that way and @sesame has a nice demo recently https://t.co/K2QclRkycP