Great to see the MedScribe benchmark we designed with @ValsAI and Harvard used to judge current SOTA models.
Today, Protege holds over 2,000,000 visits worth of medical audio, constituting over 200,000 total audio hours, along with 1,000,000,000's of medical notes that correspond to hundreds of millions of SOAP notes.
For this benchmark and from this corpus, we curated medical records and SOAP notes and test the model's ability to generate accurate SOAP notes from a medical transcript.
Most importantly.... none of the benchmark cases have ever appeared in any foundation model training!
Getting this right is huge for any medical AI application. A SOAP (Subjective, Objective, Assessment, Plan) note is a key endpoint in care delivery frameworks, where a doctor writes the care plan and assessment for an office visit.
We'll continue to see models improve on real-world tasks like these so long as they have access to the real-world data that accurately reflects real-world conditions.
we’re hosting an AI researcher event in London next week, only a few spots left.
60 people from top labs, startups, data companies talking about the state of post training.
we have a few spots left - register here and lmk so I can make sure you’re added
https://t.co/hqR0tJLiKh
Today’s most capable models are increasingly constrained by access to data. Public datasets are largely exhausted. The internet has been scraped. AI models now need high-quality, real-world data to continue advancing.
That’s why we’re excited to announce a $30M Series A investment in Protege, the platform building the real-world data infrastructure for AI.
@withprotegeai connects the world’s leading AI builders with massive, multimodal datasets across healthcare, video, audio, motion capture, and more, unlocking data that has historically been fragmented, inaccessible, or prohibitively slow to use at scale.
Protege has already found an impressive product-market fit, and is a core data partner to the majority of MAG7 public companies, as well as many of the largest private players in AI.
We couldn’t think of a better team to build this company than Bobby Samuels (cofounder & CEO) and Travis May (cofounder and Chairman), who spent a decade building Datavant and LiveRamp into two of the biggest exits in the data space, alongside fellow cofounders Engy Ziedan (Chief Scientific Officer) and Richard Ho (CTO), who bring deep data and technical expertise.
By @daisydwolf and Eva Steinman
@BobbySamuels@engyziedan
@WillManidis great post, have felt the same way about Boston. It needs to turn around — universities/hospitals have such much gravitational pull into the city but there’s little hook outside of that for people to stick around
I’ll be at NeurIPS in sd from wed->friday of this week and would love to meet up with anyone in the space!
I spend most of my time thinking about how external data can be used in AI (pretraining, RL, evals, etc).
I’m also excited about the space of open source models (and infrastructure that supports them) or generally in bio/healthcare.
I think the Reflection AI announcement may be one of the most underrated developments in the AI space in a long time. raised $2B to build pure open weights models and a commitment that we haven't seen from any American company today besides AI2/NVIDIA.
I think this could be huge for start-ups and other companies building on these models.
inference providers will greatly benefit from investment in open-weights models as well.
What am I missing? Why isn’t this being talked about more?