I’m excited to share the news of Gemini Deep Think’s gold-medal level performance 🥇 at the International Math Olympiad! It has been an absolute blast building Deep Think this year and then scaling it to the IMO.
We’re presenting the first AI to solve International Mathematical Olympiad problems at a silver medalist level.🥈
It combines AlphaProof, a new breakthrough model for formal reasoning, and AlphaGeometry 2, an improved version of our previous system. 🧵 https://t.co/SYaLPSbIyj
@gneubig Thank you for this great effort! I am achieving a 78% score on the `gsm8k-python` benchmark using Gemini Pro and this colab notebook: https://t.co/ZFJ024EvN1.
`gsm8k-python` was featured in the PaLM paper (https://t.co/bRRegQkmve) and PAL paper (https://t.co/Hb8YprAWJY).
Extremely excited to announce Promptbreeder 🌱, a self-referential self-improving system that can automatically evolve effective domain-specific prompts in a given domain. Prompts evolved with Promptbreeder outperform Chain-of-Thought and Plan-and-Solve Prompting on a range of arithmetic and commonsense reasoning benchmarks.
Creating self-referential self-improving systems is a holy grail of AI research, but prior self-referential approaches rely on costly model parameter updates—it's unclear how to scale them to the vast number of parameters in modern LLMs, let alone how to make them work with LLMs whose parameters are hidden behind APIs. Instead, in Promptbreeder the substrate for self-improvement is natural language, which allows us to modify mutation operators using an LLM in a self-referential way. Thus, Promptbreeder not only improves prompts, but over multiple evolutionary generations it also improves the way it is improving prompts.
I keep on coming back to Figure 4 in the original Chain-of-Thought Prompting paper: as the LLM gets bigger and better, the gains from prompt strategies increase significantly. I believe the future is going to be wild as we will see increasingly open-ended self-referential self-improvement systems that continue to scale with ever larger and more capable LLMs. Promptbreeder is an important stepping stone in this direction.
Really enjoyed this collaboration with amazing co-workers @chrisantha_f@2ne1@hmichalewski@sindero
arXiv: https://t.co/o3ZjsgpY6a
🔬Found a brilliant #PhysicsCalculator on Google Search! Solves queries like energy calculation & estimates gravitational force, while explaining the science behind. Ideal for students, teachers, and curious minds. A gem for physics exploration! 💡🌌🔭 #GoogleSearch#ScienceTool
Very excited to present Minerva🦉: a language model capable of solving mathematical questions using step-by-step natural language reasoning.
Combining scale, data and others dramatically improves performance on the STEM benchmarks MATH and MMLU-STEM. https://t.co/bQJOyMSCD4
Super excited to share Minerva!! – a language model that is capable of solving MATH with 50% success rate, which was predicted to happen in 2025 by Steinhardt et. al. (https://t.co/DFV2nARhD0)!
#Minerva
1/
Thrilled to announce🦉Minerva: a large language model capable of solving mathematical problems using step-by-step reasoning in natural language.
See blog here: https://t.co/eDtHy9oXci and samples here: https://t.co/GGECkO5Noo (1/n)
How can we effectively train generalist multi-environment agents? We trained a single Decision Transformer model to play many Atari games simultaneously and compared it to alternative approaches: https://t.co/IRGp6kMYAK
I am looking for students who are interested in doing a Ph.D. in ML/NLP/DL/RL at @UCL@UCLCS starting in Fall 2019. See instructions below and get in touch via PM if you have questions. Application deadline is 5th of January 2019. https://t.co/WVViYMibFo