The Generator-Verifier Gap. When verifying a solution is easier than generating it (think Chess moves, coding, math), scaling up inference time compute gets better results — @OpenAI's Noam Brown speaking at the Simons Institute's workshop on Transformers as a Computational Model.
This Grothendieck quote is the best-kept secret about mathematics:
Always start with your naive intuition, even if it’s plainly dumb, and then refine it by asking a barrage of ���stupid” questions. Waiting in silence until you “get it right” will only lead to paralysis.
Bard is now called Gemini; and Gemini Advanced with Ultra 1.0 has launched! It was the most preferred chatbot in blind evaluations with third-party raters. And it’s now available on mobile on Android and iOS. Have fun trying it out!
@dwnews Now some poor souls will have to clean all this up, being paid from taxes collected from the same farmers. You really showed those politicians their place!
I’m very excited to share our work on Gemini today! Gemini is a family of multimodal models that demonstrate really strong capabilities across the image, audio, video, and text domains. Our most-capable model, Gemini Ultra, advances the state of the art in 30 of 32 benchmarks, including 10 of 12 popular text and reasoning benchmarks, 9 of 9 image understanding benchmarks, 6 of 6 video understanding benchmarks, and 5 of 5 speech recognition and speech translation benchmarks. Gemini Ultra is the first model to achieve human-expert performance on MMLU across 57 subjects with a score above 90%. It also achieves a new state-of-the-art score of 62.4% on the new MMMU multimodal reasoning benchmark, outperforming the previous best model by more than 5 percentage points.
Gemini was built by an awesome team of people from @GoogleDeepMind, @GoogleResearch, and elsewhere at @Google, and is one of the largest science and engineering efforts we’ve ever undertaken. As one of the two overall technical leads of the Gemini effort, along with my colleague @OriolVinyalsML, I am incredibly proud of the whole team, and we’re so excited to be sharing our work with you today!
There’s quite a lot of different material about Gemini available, starting with:
Main blog post: https://t.co/NzSycJl7aE
60-page technical report authored by th Gemini Team: https://t.co/CEdMRyYSLo
In this thread, I’ll walk you through some of the highlights.
@dmvaldman Good question, chatgpt might suggest that it would be the generalist, but IMO that's an illusion, LLMs cannot really generalize and synthesize new knowledge, and I suspect we're not even close here. Also, specialist is already kinda becoming obsolete (e.g. AlphaGo or AlphaFold).
@unusual_whales Interest payments become too high to manage -> print money to avoid default -> inflation raises -> increase interest rates to lower inflation -> interest payments become even higher -> ... 🤡
@BartoszMilewski Since the molecules could technically fall into a "looped" chain of state configurations, you would have to wait infinitely long to check whether my molecule will ever visit the 2nd container.
@BartoszMilewski 1/n In fact, you can consider a simpler case. Take a 10x10m vacuum room with a 1x1m container of some gas in one corner of the room, and an empty container in a different corner or the room. Now, I pick one particular gas molecule from the first box, and then release them all
@BartoszMilewski Actually, giving a 1 hour bound for this problem doesn't work, since you could just wait. So let's say instead the question is whether my molecule will visit the second container at all.