DIRL: Domain-Invariant Representation Learning for Sim-to-Real Transfer
*Best Systems Paper Award Finalist at #CoRL2020*
paper: https://t.co/Fz76bLvpK5
website: https://t.co/Hvo8iOlbnB
Also, a shout out to the organizers and participants for the amazing work @corl_conf
🚨 Big news 🚨
Together with a set of amazing folks we decided to start a company that tackles one of the hardest and most impactful problems - Physical Intelligence
In fact, we even named our company after that: https://t.co/t0eFmHoro9 or Pi (π) for short
🧵
I’m very excited to share our work on Gemini today! Gemini is a family of multimodal models that demonstrate really strong capabilities across the image, audio, video, and text domains. Our most-capable model, Gemini Ultra, advances the state of the art in 30 of 32 benchmarks, including 10 of 12 popular text and reasoning benchmarks, 9 of 9 image understanding benchmarks, 6 of 6 video understanding benchmarks, and 5 of 5 speech recognition and speech translation benchmarks. Gemini Ultra is the first model to achieve human-expert performance on MMLU across 57 subjects with a score above 90%. It also achieves a new state-of-the-art score of 62.4% on the new MMMU multimodal reasoning benchmark, outperforming the previous best model by more than 5 percentage points.
Gemini was built by an awesome team of people from @GoogleDeepMind, @GoogleResearch, and elsewhere at @Google, and is one of the largest science and engineering efforts we’ve ever undertaken. As one of the two overall technical leads of the Gemini effort, along with my colleague @OriolVinyalsML, I am incredibly proud of the whole team, and we’re so excited to be sharing our work with you today!
There’s quite a lot of different material about Gemini available, starting with:
Main blog post: https://t.co/NzSycJl7aE
60-page technical report authored by th Gemini Team: https://t.co/CEdMRyYSLo
In this thread, I’ll walk you through some of the highlights.
We remain committed to our partnership with OpenAI and have confidence in our product roadmap, our ability to continue to innovate with everything we announced at Microsoft Ignite, and in continuing to support our customers and partners. We look forward to getting to know Emmett Shear and OAI's new leadership team and working with them. And we’re extremely excited to share the news that Sam Altman and Greg Brockman, together with colleagues, will be joining Microsoft to lead a new advanced AI research team. We look forward to moving quickly to provide them with the resources needed for their success.
Need help in the real world? RoboVQA can guide robots and humans through long-horizon tasks on a phone via Google Meet.
We release a dataset of 800k (video, question/answer) with robots & humans doing various long-horizon tasks.
Data: https://t.co/TbRRrDwQ6O
@GoogleDeepMind
Delighted to join NUST US Alumni on the 5th Anniversary of NUSTIAN 🇺🇸@nustianusa in Silicon Valley. Impressed by their commitment to build community, create networking opportunities for the 1200+ NUST alumni in the US and do impactful philanthropic work in 🇵🇰.
GitHub Copilot is already helping developers code faster in their IDEs. But what’s next?
Our answer is GitHub Copilot X. It’s our vision for the future of AI-powered software development. Check it out ⬇️ https://t.co/3Xrn7dAPgi
What happens when we train the largest vision-language model and add in robot experiences?
The result is PaLM-E 🌴🤖, a 562-billion parameter, general-purpose, embodied visual-language generalist - across robotics, vision, and language.
Website: https://t.co/ouMkeQiGr5
Some nice work by many @GoogleAI and @DeepMind researchers to assess the quality of answers to a variety of clinical QA datasets.
In general, larger models+careful fine tuning on medical text performs better (many things to figure out before real use, but improving fast).
Excited to share Med-PaLM, a large language model aligned to the medical domain to generate safe and helpful answers.
Our work advances SOTA in 7 medical question-answering tasks, including achieving 67% on MedQA USMLE improving prior work by >17%.
https://t.co/FSSpzATotz
Delighted to share our new @GoogleHealth@GoogleAI @Deepmind paper at the intersection of LLMs + health.
Our LLMs building on Flan-PaLM reach SOTA on multiple medical question answering datasets including 67.6% on MedQA USMLE (+17% over prior work).
https://t.co/jZZuFDrxGw
💡New paper - Large Language Models Encode Clinical Knowledge💡 Our work @GoogleHealth@GoogleAI @DeepMind advances state-of-art in 7 medical question-answering tasks - including achieving 67% on MedQA (USMLE qs) improving prior work by >17%
https://t.co/ugplzJvLaY
1/n
Almost ♾ unlabeled data is the “secret sauce” for today's ML, but how do we use uncurated datasets in robot learning?
Conditional Behavior Transformer makes sense of "play" style robot demos w/ no labels and no RL to extract conditional policies!
https://t.co/2uCJtol5Wt 🧵
RepsNet: Combining Vision with Language for Automated Medical Reports
https://t.co/sBfbwW4G4B
by Ajay Kumar Tanwani et al. including @the_real_dfree
#ComputerVision#PatternRecognition
Super excited to introduce SayCan (https://t.co/NWyvPubhmE): 1st publication of a large effort we've been working on for 1+ years
Robots ground large language models in reality by acting as their eyes and hands while LLMs help robots execute long, abstract language instructions
@CVPR Because CMT has been down for the last hour before the CVPR 2022 deadline, we decided to extend the paper submission deadline to 11:59am PDT (noon) on 11/18/2021. We need to wait for the CMT team to bring the submission site back first thing in the morning.