Our new paper unravels the mysteries of LLM overthinking through the lens of sub-thoughts!
https://t.co/nXIZ62e3wk
great work by our intern @FrederickXZhang on leading this project!
How do LLMs really navigate the thinking space? Straight off to a final answer OR follow a wiggly path? Definitely commit OR get stuck to “infinite” self-doubting?
In our latest study, we unravel (over-)thinking through the lens of sub-thoughts: https://t.co/Wb5AIcbI6a
more in 🧵
Introducing FRAMES: a challenging benchmark for testing retrieval + factuality + reasoning.
Developed by our amazing intern @SatyaScribbles at @GoogleAI
🚀 Excited to share the research I worked on during my summer internship at @GoogleAI! We developed FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), a challenging high-quality benchmark for evaluating retrieval-augmented large language models. FRAMES tests LLMs on retrieving relevant info, reasoning across documents, and providing factual responses to complex questions. 🧵👇 #AI #RAG #Eval
Dataset link: https://t.co/krEijwOist
Paper link: https://t.co/WVyV2izxZk
7). FRAMES - a unified framework to evaluate an LLM’s ability to provide factual responses, assess retrieval capabilities, and the reasoning required to generate final responses...
https://t.co/UwOvSCgu3Q
Exciting News from Chatbot Arena!
@GoogleDeepMind's new Gemini 1.5 Pro (Experimental 0801) has been tested in Arena for the past week, gathering over 12K community votes.
For the first time, Google Gemini has claimed the #1 spot, surpassing GPT-4o/Claude-3.5 with an impressive score of 1300 (!), and also achieving #1 on our Vision Leaderboard.
Gemini 1.5 Pro (0801) excels in multi-lingual tasks and delivers robust performance in technical areas like Math, Hard Prompts, and Coding.
Huge congrats to @GoogleDeepMind on this remarkable milestone!
Gemini (0801) Category Rankings:
- Overall: #1
- Math: #1-3
- Instruction-Following: #1-2
- Coding: #3-5
- Hard Prompts (English): #2-5
Come try the model and let us know your feedback!
More analysis below👇
Gemini 1.5 Model Family: Technical Report updates now published
In the report we present the latest models of the Gemini family – Gemini 1.5 Pro and Gemini 1.5 Flash, two highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio.
Our latest report details notable improvements in Gemini 1.5 Pro within the last four months.
Our May release demonstrates significant improvement in math, coding, and multimodal benchmarks compared to our initial release in February.
Furthermore, the 1.5 Pro Model is now stronger than 1.0 Ultra.
The latest Gemini 1.5 Pro is now our most capable model for text and vision understanding tasks, surpassing 1.0 Ultra on 16 of 19 text benchmarks and 18 of 21 of the vision understanding benchmarks. The table below highlights the improvement in average benchmark performance for different categories in 1.5 Pro since Feb, and also shows the strength of the model relative to the 1.0 Pro and 1.0 Ultra models. The 1.5 Flash model also compares very well against the 1.0 Pro and 1.0 Ultra models.
One clear example of this can be seen on MMLU
On MMLU we find that 1.5 Pro surpasses 1.0 Ultra in the regular 5-shot setting scoring 85.9% versus 83.7%. However with additional inference compute, via majority voting on top of multiple language model samples, we can get a performance of 91.7% versus Ultra’s 90.0%, which extends the known performance ceiling of this task.
@OriolVinyalsML and I are very proud of the whole Gemini team, and it’s fantastic to see this progress and to share these highlights from our Gemini Model Family.
Read the updated report here: https://t.co/CTzTHND4nQ
FYI I know the rental market in NYC is absurd rn, but as a principle please don’t overbid. It sets the price of the unit/area in stone and trending up. That’s how people get priced out of the city.
Having kids today must be so stressful. Like not only do you need to worry about all the normal stuff but you also need to make sure they don’t become a LinkedIn influencer.
@obsessivelyocd I'm not as mentally strong as I should be and my OCD often does get the better of me. It can get very very lonely. Sending love, you are way stronger than me. My DMs are open!
As I sit at a coffee shop in the Mission (SF) I see 30 year o tech dudes talking at 5x speed about how to improve their Medium. No thanks I’m glad I left the bay
URGENT for journalist followers - @uscis is now expected to DENY most Afghan humanitarian parole cases - more than 30K pending - because they are applying standards most applications won't meet. Here's why - 1x