@lefthanddraft@ianchanning I think Balrog is the best evidence for those kind of tasks for how much gain can you get by switching from image to text
https://t.co/LkLP4XGhfT
@DimitrisPapail They are not rank 2 though. You are taking a tiny subset of highly correlated evals. There are thousands of them, and we don't even have to go far. Just taking 5 examples from life bench:
@joejanizek You probably saw this one. In your opinion, how much does this benchmark tell us about "replacing" radiologists? How "Last" is this exam?
https://t.co/RsaCtMp1H2
๐ฅ Gemini 3.0 vs Radiologists: RadLE Benchmark Results Are OUT!
โ ๏ธ Is it game over for Radiology? Let us find out! โฌ๏ธ
๐ซจ Since yesterday, Gemini 3.0 has been everywhere for crushing benchmarks. My inbox exploded asking: โBut how did it do on the hardest visual reasoning benchmark in healthcare?โ
So we ran it!
And here you go. ๐
โก๏ธ Gemini 3.0 Pro on RadLE v1:
โ 51% accuracy; first time a general-purpose model has beaten radiology residents
โ Radiology residents: 45%
โ Board-certified radiologists: ~83%
โ Shows clean step-by-step reasoning in some tough cases (appendix localization, mimics ruled out, etc.)
๐ This is the first time ever that a generalist model has crossed the trainee bar on RadLE v1!
Congratulations to @GoogleDeepMind and @Google team including @vivnat, @alan_karthi and all others for cooking this time!
Full breakdown here:
๐ Link in comments / bio
๐ฅ Huge shoutout to Lakshmi, Divya, Upasana, Hakikat, Kautik & the entire #CRASHLab team at @KCDH_A for turning around in under a day.
๐ If you are a medical AI lab and want to improve your performances and want our expert insights, reach out!