Today is the start of a new era of natively multimodal AI innovation.
Today, we’re introducing the first Llama 4 models: Llama 4 Scout and Llama 4 Maverick — our most advanced models yet and the best in their class for multimodality.
Llama 4 Scout
• 17B-active-parameter model with 16 experts.
• Industry-leading context window of 10M tokens.
• Outperforms Gemma 3, Gemini 2.0 Flash-Lite and Mistral 3.1 across a broad range of widely accepted benchmarks.
Llama 4 Maverick
• 17B-active-parameter model with 128 experts.
• Best-in-class image grounding with the ability to align user prompts with relevant visual concepts and anchor model responses to regions in the image.
• Outperforms GPT-4o and Gemini 2.0 Flash across a broad range of widely accepted benchmarks.
• Achieves comparable results to DeepSeek v3 on reasoning and coding — at half the active parameters.
• Unparalleled performance-to-cost ratio with a chat version scoring ELO of 1417 on LMArena.
These models are our best yet thanks to distillation from Llama 4 Behemoth, our most powerful model yet. Llama 4 Behemoth is still in training and is currently seeing results that outperform GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on STEM-focused benchmarks. We’re excited to share more details about it even while it’s still in flight.
Read more about the first Llama 4 models, including training and benchmarks ➡️ https://t.co/9G3QgVdCkB
Download Llama 4 ➡️ https://t.co/eVomRvEr0w
@WesKambale@ngamita@UCC_ED@UCC_Official These are absolutely valid points and concerns. I skimmed through the comments, and discussing whether people understand LLMs, whether they are on Twitter, etc., made the important point you wanted to communicate vanish.
How to write the Introduction?
As a junior student, writing the introduction of a research paper is arguably the most daunting part of paper writing. 😱
Here is a simple template I find useful:
3 Figures 🖼️ + 5 Questions 🤔
Best of AI Twitter (Sept. 18-25):
- OpenAI releases Whisper (human-level speech-to-text)
- Adversarial, interactive deepfakes
- Twitter "gossip" on how OpenAI gathers GPT-4's trillions of training tokens
- The guts of Google's multibillion-parameter ad model
... and more:
1/15
🎉 StatsForecast Exponential Smoothing (ETS) is 32% more accurate and over 100 times faster than NeuralProphet🔥
We hope this exercise by @nixtlainc helps the forecast community avoid adopting forecasting methods that are not thoroughly tested💡
#forecasting#python#timeseries
.
@DeepMind released this GOLD a couple days back. If you ever wanted to study Transformers from scratch I think this would be that one resource you wouldn't want to miss:
https://t.co/ORtQAinlUj
@UmemeLtd On a serious note, Pole replacement and conductor upgrade in this area if a serious contractor was awarded shouldn't take more than 2 days. Now its two months, Do you assess the work man hours wasted and its effects on the businesses. Sort your contracting!.
We are eager to participate in @phpulpocon'22 in September in #Vigo! Our contribution is twofold: @anafrio will deliver a talk, and we are supporting the conference through a gold sponsorship. Moitas ganas!
https://t.co/lSC0Y4WdU3
In the spirit of the PhD admissions season ending, I'm making my state of purpose public.
I learned a lot from reading @nelsonfliu and @ssgrn's SoPs, and so I'd like to pay it forward.
https://t.co/awYteBgHow
https://t.co/MhJQc5nBnP
Prior SoPs below:
Introducing the Common Voice-based Speech-to-Speech translation corpus, CVSS, which includes 2657 hours of speech-to-speech translation sentence pairs from 21 languages into English. Read about its development and how we used it to train baseline models ↓ https://t.co/dxtif8v9bM