Llama 4 Intelligence Index Update: We have now replicated Meta’s claimed values for MMLU Pro and GPQA Diamond, pushing our Intelligence Index scores for both Scout and Maverick higher
Key update details:
➤ We noted in our first post 48 hours ago that we noticed discrepancies between our measured results and Meta’s claimed scores for our multi-choice eval datasets (MMLU Pro and GPQA Diamond)
➤ After further experiments and and close review, we have decided that in accordance with our published principle against unfairly penalizing models where they get the content of questions correct but format answers differently, we will allow Llama 4’s answer style of ‘The best answer is A’ as legitimate answer for our multi-choice evals
➤ This leads to a jump in score for both Scout and Maverick (largest for Scout) in 2/7 of the evals that make up Artificial Analysis Intelligence Index, and therefore a jump in their Intelligence Index scores
➤ Scout’s Intelligence Index has moved from 36 to 43, and Maverick’s Intelligence Index has moved from 49 to 50.
Overall, we continue to conclude that both Scout and Maverick are very impressive models and a significant contribution to the open weights AI ecosystem.
While DeepSeek V3 0324 maintains a small lead over Maverick, we continue to note that Maverick has ~half the active parameters (17B vs 37B), and ~60% of the total parameters (402B vs 671B), while also supporting image inputs.
All our tests have been performed on the Hugging Face release version of the Llama 4 weights for both Scout and Maverick, including testing via a range of third party cloud providers. None of our eval results are based on the experimental chat-tuned model provided to LMArena (Llama-4-Maverick-03-26-Experimental).
We can also share that we have observed third party cloud APIs generally stabilizing over the last 48 hours. We will soon release endpoint-level comparison data to allow developers to understand whether any cloud providers are still serving versions of Llama 4 with accuracy issues.
BREAKING: Meta's Llama 4 Maverick just hit #2 overall - becoming the 4th org to break 1400+ on Arena!🔥
Highlights:
- #1 open model, surpassing DeepSeek
- Tied #1 in Hard Prompts, Coding, Math, Creative Writing
- Huge leap over Llama 3 405B: 1268 → 1417
- #5 under style control
Huge congrats to @AIatMeta — and another big win for open-source! 👏 More analysis below⬇️
Struggling to harness the power of AI for your business? 🚀
Introducing the First Llama Incubator in Singapore—an initiative by Meta, industry leaders, and the Singapore government. This program empowers startups and SMEs to scale faster and operate smarter using open-source AI.
🏆 Congratulations to first place from the #Llama3Hackathon this weekend: Open Glass AI!
More details and open source code ➡️ https://t.co/Z55m87Ny6M
It’s inspiring to see such innovative work from the open source community, built with Meta Llama 3!
Introducing Meta Llama 3: the most capable openly available LLM to date.
Today we’re releasing 8B & 70B models that deliver on new capabilities such as improved reasoning and set a new state-of-the-art for models of their sizes.
Today's release includes the first two Llama 3 models — in the coming months we expect to introduce new capabilities, longer context windows, additional model sizes and enhanced performance + the Llama 3 research paper for the community to learn from our work.
More details ➡️ https://t.co/nFll4exicO
Download Llama 3 ➡️ https://t.co/Ps0OAHt0RR
We're collaborating with some of the best in the industry to set the standard for identifying AI generated content.
In the near future, we will label images on Facebook, Instagram, and Threads informing people when they're engaging with GenAI content.
https://t.co/r7Z0LmV9xd
Today, Meta is releasing its 2nd annual human rights report. It shares insights and actions from 2022, and reports the progress we’ve made in implementing our human rights policy across the company. It includes details of the APAC Human Rights Defenders fund, which to date has helped over 500 people with urgent needs.
https://t.co/qIESX347L8
Sad to hear of the passing of sister Shuhada Sadaqat, also known as Sinéad O'Connor. She was a tender soul, may God, Most Merciful, grant her everlasting peace. Inna lillahi wa inna ilayhi rajioon - Verily we belong to God, and verily to Him do we return. 2:156
We support an open innovation approach to AI. Responsible and open innovation gives us all a stake in the AI development process, bringing visibility, scrutiny and trust to these technologies. Opening today’s Llama models will let everyone benefit from this technology.
Today, we are announcing the availability of Llama 2 for commercial use. Llama 2 is a foundational large language model (LLM). Users can access Llama 2 by signing up through Meta or through other cloud computing platforms, including Microsoft’s Azure.
https://t.co/gIU7vqPLR6
In my latest article, I write about the misconception that data is “the new oil”. This idea increasingly underpins digital nationalism around the world, which threatens the free flow of data across borders that is fundamental to how the open internet works.
Meta is proud to have joined in the launch of the Alliance for Afghan Women’s Economic Resilience. We're committed to providing digital skills to women and vulnerable communities in Afghanistan and the diaspora to improve Afghan women’s participation in the global digital economy
We’re delighted to announce that Meta has committed ongoing financial support for the Board, including a new $150 million contribution to the Oversight Board Trust to support our operations. 🧵