Big news – Gemini 1.5 Flash, Pro and Advanced results are out!🔥
- Gemini 1.5 Pro/Advanced at #2, closing in on GPT-4o
- Gemini 1.5 Flash at #9, outperforming Llama-3-70b and nearly reaching GPT-4-0125 (!)
Pro is significantly stronger than its April version. Flash’s cost, capabilities, and unmatched context length make it a market game-changer!
Huge congrats to @GoogleDeepMind on the incredible
Gemini launches! Can't wait to see what new applications Gemini unlocks!
More breakdown analysis below👇
Congrats, @JeffDean@GoogleDeepMind! Gemini 1.5 Pro has shown substantial improvements from Feb to May, scoring 63.9% on our #MathVista (https://t.co/kf2dU6ATDn), outperforming humans and GPT-4o, which was out 4 days ago!🚀
AI Progress has never been this rapid and impressive!🌟
Gemini 1.5 Model Family: Technical Report updates now published
In the report we present the latest models of the Gemini family – Gemini 1.5 Pro and Gemini 1.5 Flash, two highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio.
Our latest report details notable improvements in Gemini 1.5 Pro within the last four months.
Our May release demonstrates significant improvement in math, coding, and multimodal benchmarks compared to our initial release in February.
Furthermore, the 1.5 Pro Model is now stronger than 1.0 Ultra.
The latest Gemini 1.5 Pro is now our most capable model for text and vision understanding tasks, surpassing 1.0 Ultra on 16 of 19 text benchmarks and 18 of 21 of the vision understanding benchmarks. The table below highlights the improvement in average benchmark performance for different categories in 1.5 Pro since Feb, and also shows the strength of the model relative to the 1.0 Pro and 1.0 Ultra models. The 1.5 Flash model also compares very well against the 1.0 Pro and 1.0 Ultra models.
One clear example of this can be seen on MMLU
On MMLU we find that 1.5 Pro surpasses 1.0 Ultra in the regular 5-shot setting scoring 85.9% versus 83.7%. However with additional inference compute, via majority voting on top of multiple language model samples, we can get a performance of 91.7% versus Ultra’s 90.0%, which extends the known performance ceiling of this task.
@OriolVinyalsML and I are very proud of the whole Gemini team, and it’s fantastic to see this progress and to share these highlights from our Gemini Model Family.
Read the updated report here: https://t.co/CTzTHND4nQ
Gemini-1.5-flash as a Copilot in VSCode is amazing!
You can now use this model by connecting CodeGPT with Google AI Studio.
@codegptAI + @googleaistudio
In this video, I show how CodeGPT manages to get the entire context of the "Quick Fix" section and Gemini provides a complete solution to the potential installation error of the @langchain library in my notebook 🙌
Excellent work @OfficialLoganK and the whole @GoogleAI ! 👏
📢Berkeley Function Calling Leaderboard Update: Discover the enhanced performance and cost-efficiency of @GoogleDeepMind's Gemini-1.5-pro and Gemini-1.5-flash, alongside @OpenAI's new gpt-4o models ⚡️Gemini sets a new benchmark in function-calling 🏆and improves its capabilities in relevance detection - ability to say NO! Parallel-multiple function calls continue to be elusive 👀
Play around: https://t.co/hcwjnTajFp
📢 HELM now supports VLM evaluation to evaluate VLMs in a standardized and transparent way. We started with 6 VLMs on 3 scenarios: MMMU, VQAv2 and VizWiz. Stay tuned for more - this is v1!
✍️ Blog post: https://t.co/kkYae5dvFs
💯 Raw predictions/results: https://t.co/eHRJtAXo3r
Gemini 1.5 Pro surges on the @lmsysorg leaderboard, surpassing GPT-4-0125! It's a new top-tier contender, especially for long prompts.🔥 Also, Llama3 score is still increasing.
@GoogleDeepMind#GeminiAI#LLM#LMSYS#llama3
More exciting news today -- Gemini 1.5 Pro result is out!
Gemini 1.5 Pro API-0409-preview now achieves #2 on the leaderboard, surpassing #3 GPT4-0125-preview to almost top-1!
Gemini shows even stronger performance on longer prompts, in which it ranks joint #1 with the latest GPT-4-Turbo🔥
Big congrats to @GoogleDeepMind on shipping this powerful Gemini API to developers & community. Very excited to see what app can be built on top!
This past weekend, students at @MHacks pushed the boundaries of Gemini 1.5, unlocking inspiring use cases across safety, privacy, accessibility, and creativity. Thanks to all who participated and to @UMichPrezOno for swinging by ✨ Read up on the winners below:
had a raspberry pi laying around and built an ai wearable called insight at @Google x @mhacks hackathon this weekend.
insight uses gemini 1.5 pro to answer questions based on what you see and hear, and it remembers those memories for you.
repo in comments
🎙️📹Audio & Video Structured Extraction with Gemini♊️
Google's Gemini 1.5 Pro officially came out of preview yesterday, with support for audio and video inputs, and they work with function calling!
We just released a new short YouTube video outlining how to perform structured extraction on YouTube videos, and audio clips using LangChain JS/TS 🦜🔗!
Watch the video here: https://t.co/TixM82aAFT
Read the docs 👇
🎉 It’s a big day for @Google Gemini.
Gemini 1.5 Pro now understands audio, uses unlimited files, acts on your commands, and lets devs build incredible things with JSON mode! It’s all 🆓. Here’s why it’s a big deal 👇
🔈 Gemini can hear
Gemini understands audio (up to 9.5 hours of audio) but not just the words you’re saying but the tone and emotion behind the audio. In some cases it even understands some sounds such as dogs barking and rain falling.
👩🏫 TEACHERS: Upload a recording of your lecture to create a 10 question quiz on the most important content.
👨💼 CONSULTANTS: Upload a recording of an entire day’s offsite and create a new team strategy document.
🚀 FOUNDERS: Record your startup pitch and get specific feedback on how you could improve your next round.
🗂️ Gemini can use unlimited files.
When working with Gemini, you can upload nearly unlimited files (images, video frames and audio) to ask Gemini questions against and it’s free.
👨🎨 CREATIVE PROFESSIONALS: Upload your mood board to get a theme and color for your next product.
👩🎓 STUDENTS / RESEARCH: Upload documents and photos from papers or notes to summarize your thesis.
❤️ FAMILY STUFF: Upload hundreds of family photos and Gemini can figure out the best holiday card photos.
⚒️ Better function calling and system instructions.
2023 was the year of text based chat with ChatGPT. 2024 will be the the year of AI agents taking action on behalf of people. Gemini can understand thousands of actions and figure out what to do next for people.
👩💻 TECH COMPANIES: Create a general digital assistant like Siri or Alexa that actually can do thousands of tasks.
👩🔧 CUSTOMER SERVICE: Create a call center bot that doesn’t infuriate people.
⛩️ Finally, JSON mode.
JSON mode has been the biggest ask for developers since the team started giving access to the power of Gemini. In short it lets developers pull information out of text, speech or videos in a structured way. Today, it’s live and I’m so proud of the team getting this in.
👯♀️ It’s public. It’s free. No waitlist.
The largest roadblock to using Gemini was the waitlist. Google is showing its incredible capacity to host these models and today is the day we let the floodgates open. No more waitlist and it still includes all the magical features that exist in Gemini:
📚 1 million token context window. Ask Gemini about ten books of text, 9 hours of audio or a movie worth of video frames.
⚖️ Advanced reasoning, understanding and performance.
🚦 Set safety limits or remove them all together.
💸 If you want to build a new app or experience on Gemini, there’s a new paid tier for even higher rate limits.
Let me know what you do with #Gemini. I’ll be posting a few things I’m doing with it soon.
Sergey’s inspiring q&a @ AGI house - pt 1 (@elonmusk watch this, instead of the boob guy - who flew in from Mexico to attend our event and has nothing to do w Google)
For me, Google is the #1 AI company in the world and kickstarted the whole revolution with "Attention is all you need", BERT, T5 and many more.
I'm ecstatic to be able to collaborate with @sundarpichai@ThomasOrTK@JeffDean@fchollet and the whole Google org to democratize AI thanks to open-source AI and open science! Let's go!
📢 #NotebookLM heads: our team just dropped some new features powered by #GeminiPro ✨ Your notebook can now generate study guides, map outlines, offer instructive feedback, and more. We'd love to see how you use these new features (tag us!) https://t.co/Z55vDrKEdR