Happy to share that the @GoogleDeepMind Gemini team is starting a new research team in Singapore!
This new team will be focused on advanced reasoning, LLM/RL and improving bleeding edge SOTA models such as Gemini, Gemini Deep Think and beyond. ๐ฅ
This team will be led by yours truly and reports up to Quoc Le (@quocleix)'s broader team in Mountain View which was recently in the center of both IMO gold medal and ICPC gold medal breakthroughs with Gemini Deep Think, amongst many other significant Gemini advancements. ๐
Weโre starting out with a very small but intensely capable force because talent density is key over anything else in the LLM era. Over the past few months, we have gone around and gathered the best of the best talent (in the region and beyond) and Iโm confident weโll have a super cracked team very soon.
If you are interested in joining and have made truly exceptional contributions in any domain or area, (engineering and/or research etc) please contact me.
This is quite an exciting time, with the Gemini / GenAI team at Google Deepmind leading the charge at the frontier. This is also the best opportunity to be on the critical path to AGI from the sunny island of Singapore. ๐๏ธ
Many thanks to leadership support from @quocleix@JeffDean@benoitschilling, @EugenieRives and @demishassabis for the support of this team.
Wonderful and fun image generated by Nano Banana ๐
Had a really wonderful time hosting @JeffDean, @quocleix, @benoitschilling and @denny_zhou in Singapore for the @GoogleDeepMind Gemini Singapore ๐ธ๐ฌ event last week! ๐ฅ
The event went super well imo, the vibes were on-point and an overwhelming number of people told me directly they were moved and very inspired! ๐
Moreover, the consensus amongst GDM speakers was also that the level of questions from the community were solid and reassuring that something good is going on here. Basically, signal to noise ratio was generally pretty good.
As I also said in my opening remarks, I think Singapore has a lot of raw pure talent, but today it is physically and spiritually far away from the true AI frontier for a myriad of reasons. Hence, increasing "frontier awareness" and "aura diffusion" (aka ecosystem building ๐) was something particularly high ROI (for us too!)
The way I see it is that SG is like a reasonably strong base pretrained model that hasn't been properly post-trained yet. It feels like the gears have started turn though, hopefully we'll make it happen. ๐ค
We also had a lot of fun roaming SG. Visiting gardens by the bay on day-0, having 2-michelin star "tarvar" as the finale dinner, pickleball with Jeff on top of a random building, spotting anteaters at the night safari and last but not least, chatting with senior minister Lee Hsien Loong and visiting the Istana.
Was an epic week! ๐
More details of the 3/9/2025 event in the thread below ๐
Excited to share that I'll be hosting some of the world's best AI researchers and engineers for our @GoogleDeepMind Gemini event next week in Singapore ๐ธ๐ฌ!
Join @JeffDean, @quocleix, @benoitschilling, @melvinjohnsonp and @denny_zhou for a day of technical conversations, panels and talks about AI, reasoning and our mission to build a world class AI frontier lab in Singapore.
If you're in town and would like to attend, please check the RSVP link below๐. Note, subject to capacity constraints and you'll need to be approved to join.
Evals are notoriously difficult to get right but necessary to move the field forward. ๐
As part of our commitment to science, weโre releasing a subset of our internal evals. ๐
Vibe-Eval is an open and hard benchmark comprising 269 image-text prompts for measuring the progress of multimodal chat models.
Vibe-Eval has two goals:
๐นvibe-checking existing models on day to day tasks ๐นdeeply probing into the capabilities of frontier models
These very difficult prompts were constructed by AI experts (us ๐คญ) with the intention to deliberately find examples where our model (Reka Core) is unable to solve.
On 50% of the prompts in the hard set, none of the existing models (including frontiers) can fully solve. ๐ฒ.
We benchmark 13 representative multimodal models using both human evaluation and automatic evaluation.
We also share our findings and insights about the challenges of curating and evaluating hard prompts.
Check out our paper and resources below โฌ๏ธ
Our @RekaAILabs Tech Report / Paper is out! ๐ฅ
Tech reports with completely no information are kinda boring so weโre revealing some interesting information on how we train our series of Reka models including tokens, architecture, data & human evaluation workflows. ๐
We tried our best to give a behind-the-scenes experience ๐. In particular, if you enjoyed my previous blog post about training LLMs in the wilderness, thereโs a dedicated section on that in this report! ๐ด
We canโt disclose literally everything but we tried our best to make it interesting, I promise. ๐
Hereโs a rundown summary of some of the highlights.
๐นEdge and Flash are outrageously strong 7B and 21B models. They are trained on 4.5-5T tokens in total. Also, they have been improved significantly since their first public appearance! They outperform many popular faces. Some data mixture information is in the report.
๐นWe discuss our internal human evaluation workflow, prompt distribution, and how we use Core for model development and automatic evaluation.
๐นWe describe our infrastructure setup for training large models, quantifying node failures, and report loss curves for training our models.
๐นAside from the hardware lottery, we also show how this affects node stability across time. Once we were told our cluster became less stable because there were "big guys" moving things around the data center. ๐
On performance which you might have already seen on other threads.
๐นCore approaches frontier-class models like Claude3 Opus and GPT4-V. It outperforms Claude3 Opus on third-party blind human evaluation for multimodal chat, outperforms Gemini Ultra on video QA, and is quite competitive to other frontier models on core text metrics. It also matches GPT4-V on MMMU!
๐นCore ranks #2 on our internal multimodal chat leaderboard, right after GPT4-V. On text, it ranks #3 just behind Claude Opus and GPT4 Turbo. Core outperforms GPT-4 (0613) on this ranking.
This has been a focused and concentrated effort of a small team of ~20 people in the past 4 months (yes, we got access to 90%+ of our compute only late December last year! ๐).
This tech report tells our story. Enjoy! Happy to answer any questions in replies or DM!
PS: it was nice writing in latex after one whole year!
PPS: I had quite some fun writing this ๐. There's some puns and easter eggs and interesting tidbits in there. Trust me. ๐
Link: https://t.co/Zve6RtRRGR
Didn't get much chance to share this yesterday with everything else going on with the Reka core launch but here's the most non-cherry picked showcase of Reka Core vs GPT-4 vs Claude Opus on multimodal chat tasks. ๐
We put together this showcase with examples our team created. How's it not cherry picked?
Our team just manually crowdsourced some prompts and examples, ran all 3 models, and load them on this simple website UI. I didn't even look at most of these.
You can get the sense of how all 3 models perform on multimodal chat from this site.
link: https://t.co/Q6KrpysToH
Meet Reka Core, our best and most capable multimodal language model yet. ๐ฎ
Itโs been a busy few months training this model and we are glad to finally ship it! ๐ช
Core has a lot of capabilities, and one of them is understanding video --- letโs see what Core thinks of the 3 body trailer.๐
*Fully funded PhD position opening @NTUsg *๐ฅ I have two openings for Ph.D. students to work on Deep Learning for Drug Discovery and Multi-Modal Large Language Models in Medicine.
Apply here ๐๐ฝ https://t.co/h1EnQLw1FH
Thanks for sharing!
We are excited to share Reka Flash โจ, a new state-of-the-art 21B multimodal model that rivals Gemini Pro and GPT 3.5 on key language & vision benchmarks ๐.
We've trained this model from scratch and ground zero with a small (but amazingly capable team ๐งโโ๏ธ) and relatively finite resources. We're amazed at how strong it is ๐ฆพ. I'm proud of our financially optimal LLM team.
Abandoning one's comfort zone is surely difficult and having to redo things from scratch is often scary & daunting. Many things in the wilderness don't work from the get go and it was often a huge pain in the neck ๐ข.
I should write a separate post someday of how much we have we've had to rebuilt (and suffered ๐คฃ). Everything from robust training infra, proper (human) evaluation pipelines and proper RLHF setups. I am thankful of the crazy talented team we have here โบ๏ธ.
Meanwhile, our largest most capable model Reka-Core is finishing soon and we're already very excited by early results ๐. More to come very soon!
9 months in. Excited to be back at the frontier ๐ฅ.
Check out our blogpost here: https://t.co/sJGIgp5CkW
We are excited to announce the 1st version of our multimodal assistant, Yasa-1, a language assistant with visual and auditory sensors that can take actions via code execution ๐ช.
Yasa-1 can understand text, images, videos, sounds & more! ๐
Check out more details below๐
Excited to share our latest work at @GoogleAI on "Transformer Memory as a Differentiable Search Index"!
TL;DR? We parameterize a search system with only a single Transformer model ๐. Everything in the corpus is encoded in the model! ๐
Paper: https://t.co/yA6jSEBcdZ
We explore how AI can generate sequences across natural language and proteins with attributes that go beyond the training data. Thanks to @thisismadani, @benwkrause and @nikhil_ai for the wonderful internship experience in @SFResearch!
Find out more here: https://t.co/MZO4x5EI7q
Can generative AI learn to extrapolate? We explore how to generate sequences that enhance desired attributes-- beyond what was seen in training. Works pretty well in #NLP and #proteins!
Blog: https://t.co/dUAgZuX0W5
Paper: https://t.co/kY4TruHUTu
Code: https://t.co/jFNccTxW0M
Inspired by the dizzying number of efficient Transformers ("x-formers") models that are coming out lately, we wrote a survey paper to organize all this information. Check it out at https://t.co/nAaTLG8wOp.
Joint work with @m__dehghani@dara_bahri and @metzlerd. @GoogleAI ๐๐
It is our pleasure to announce the Call for Papers for ICLR 2021. For the first time there will be an abstract submission (due September 28) before the full paper deadline (October 2).
For details see: https://t.co/k4iUdjt9dr
Looking forward to many exciting submissions!
Thrilled to share that our #CVPR2020 oral paper
โWhat It Thinks Is Important Is Important: Robustness Transfers Through Input Gradientsโ
will be live today at 10am/pm (PST time)!
Video: https://t.co/R12MbjDQ6X
Paper: https://t.co/WSUmQW0dlG
Code: https://t.co/dfnFEWuulm