Introducing Midcentury.
We’re building the data and simulation infra for physical AI.
Today, we’re coming out of stealth with a $15M Series Seed to scale robotics beyond polished demos.
We’re already supporting frontier labs with:
→ The world’s largest egocentric dataset: 2M+ hours, 50+ environments, 20,000+ tasks
→ Matrix: a frontier simulation platform scaled with our real-world data to evaluate and post-train policies at scale
@abidlabs@matias Thanks @abidlabs! We looked into gated datasets - the main blockers were per-recipient canaries and wanted a bit more control over the card design. Might revisit and publish the full dataset on HF!
Open Yap 1K captures 1,000 hours of natural two-speaker English speech.
🤖 https://t.co/XuGT4bNBDh
🎙️ Friends, families, partners, and colleagues talk freely. Interruptions, overlaps, backchannels, pauses, and laughter stay intact.
🔊 Each speaker has a separate 48 kHz track on the same timeline, with transcripts and word-level timing.
🌍 The full corpus contains 1,602 conversations from 239 speakers across 34 countries. Built for speech-to-speech, expressive TTS, ASR, and audio understanding.
📜 Sample: CC BY 4.0. Full corpus: free for approved commercial and research use under the Open Yap 1K Data Use Agreement.
To teach AI full-duplex conversation, let it hear the cues.
Open Yap 1K brings 1,000 hours of English two-person conversation for training and evaluating full-duplex voice models.
Speakers talk freely with people they already know, without assigned topics—with interruptions, overlapping laughter, and the little “mhm” that lets a story continue.
Recording — 239 speakers, recorded at 48 kHz, with each speaker captured on an independent track sharing the same timeline.
Overlap — the team reports median overlap of 8.3% of voiced time, reaching 20.9% at the 95th percentile.
Transcripts — Deepgram Nova-3, with word-level timestamps; not human-verified.
Full-duplex AI needs to learn when to join in, when a brief acknowledgment is enough, and when to keep listening.
Access: An 8.9-hour sample is directly downloadable under CC BY 4.0. The full 1,000 hours are free for commercial and research use, subject to application, approval, and a Data Use Agreement.
Credit to @CVestergaard_ and @matiasdrejer at @agenticdataco for the release. Also, love the new company logo—an approving “mhm” from us.
Listen and explore the sample:
https://t.co/dDpDYfUUpS
Project and full-dataset access:
https://t.co/cY4gmyaQeP
#BackchannelSignals
This is a serious alternative to speech corpora such as Fisher, Switchboard, and CANDOR. Not only on objective quality measures (48kHz, dual-channel etc), and on the human elements of speech. The interruptions, overlaps, laughter, backchannels and variance in prosody.
We designed it specifically for full-duplex and interaction models in collab with researchers to capture how people speak together naturally in real environments. Not strangers performing prompted conversations.