We’re excited to release @basiscompany’s first open source dataset: 1,500 hrs of track separated, multi-speaker conversations across 22 languages. The largest ever of its kind and first for several languages.
Basis has built up a network of over 1 million contributors submitting audio, video, annotations, ratings, transcriptions and will keep open-sourcing interaction data to cover all the weird/ambiguous/natural modes of human interaction (more soon!).
This release includes 70+ of hours with 3+ live speakers and 2645 native speakers meeting, sharing, arguing, laughing, trolling in English, Spanish, Japanese, Korean, Georgian, Mingrelian, German, French, Italian, Hebrew, Russian, Hindi, Arabic, Ukrainian, Dutch, Xhosa, Zulu, Portuguese, Chinese, Turkish, Polish, and Kazakh.
For some of these languages, this is the first ever open release of duplex conversational data.
100 hrs are densely labeled with 113k human judgements on subtleties that models struggle with (was this backchannel sympathetic or frustrated? Was this silence awkward or turn-holding? Was this utterance directed or general? …)
Reach out if you’re interested in more!
ჯგირი მორაგადეს ჯგირი მარჩქილე ოკონია
@rnafng Wait the 감자 in this case actually doesn’t mean potato 😭 it’s referring to a pork bone with the same name. There’s a double meaning except potato is way more common.
excerpts from average cs student day:
"all these git commits but she still has commitment issues"
"spend all day making relations in databases, but still can't make any in real life"