On #ICML2025 16 Jul, 11 AM
We present Meta Locate 3D: a model for accurate object localization in 3D environments.
Meta Locate 3D can help robots accurately understand their surroundings and interact more naturally with humans.
Demo, model, paper: https://t.co/8ZhV21TDxq
🐾 In the wild
Deployed on a Boston Dynamics Spot 🤖: 8️⃣/🔟 successful “find & pick the plush-toy” trials in a multi-room apt —no manual resets. Check our demo! 🎥
🏆 Results
New SOTA on SR3D + NR3D + ScanRefer: 61.7 → 49.4 %@25/50 IoU (prev. best 58.5/52.5). 🚀 Beats GPT-4o & other VLM agents by >20 pts while using only raw sensor clouds
💡 Core idea
We pre-train a 3D-JEPA encoder that masks & predicts in latent space, turning lifted CLIP+DINO point-cloud features into contextualized scene reps—no meshes, no proposals, just sensor RGB-D.
Introducing Meta Locate 3D: a model for accurate object localization in 3D environments.
Learn how Meta Locate 3D can help robots accurately understand their surroundings and interact more naturally with humans.
You can download the model and dataset, read our research paper, and even try a demo! https://t.co/GgJEPXTH8W
New work from the Robotics team at @AIatMeta . Want to be able to tell your robot bring you the keys from the table in the living room? Try out Locate 3D!
interactive demo: https://t.co/aS9WPPmhcF
model & code & dataset: https://t.co/oMWc32VrH9
llama-4-scout-17b-16e-instruct
prompt: write a p5.js script that shows a ball bouncing inside a spinning hexagon. The ball should be affected by gravity and friction, and it must bounce off the rotating walls realistically
👀 Accelerate performance of @AIatMeta Llama 4 Maverick and Llama 4 Scout using our optimizations in #opensource TensorRT-LLM.⚡
✅ NVIDIA Blackwell B200 delivers over 42,000 tokens per second on Llama 4 Scout, over 32,000 tokens per seconds on Llama 4 Maverick.
✅ 3.4X more performance and 2.6X lower cost compared to Hopper H200.
BREAKING: Meta's Llama 4 Maverick just hit #2 overall - becoming the 4th org to break 1400+ on Arena!🔥
Highlights:
- #1 open model, surpassing DeepSeek
- Tied #1 in Hard Prompts, Coding, Math, Creative Writing
- Huge leap over Llama 3 405B: 1268 → 1417
- #5 under style control
Huge congrats to @AIatMeta — and another big win for open-source! 👏 More analysis below⬇️
Introducing our first set of Llama 4 models!
We’ve been hard at work doing a complete re-design of the Llama series. I’m so excited to share it with the world today and mark another major milestone for the Llama herd as we release the *first* open source models in the Llama 4 collection 🦙. Here are some highlights:
📌 The Llama series have been re-designed to use state of the art mixture-of-experts (MoE) architecture and natively trained with multimodality. We’re dropping Llama 4 Scout & Llama 4 Maverick, and previewing Llama 4 Behemoth.
📌 Llama 4 Scout is highest performing small model with 17B activated parameters with 16 experts. It’s crazy fast, natively multimodal, and very smart. It achieves an industry leading 10M+ token context window and can also run on a single GPU!
📌 Llama 4 Maverick is the best multimodal model in its class, beating GPT-4o and Gemini 2.0 Flash across a broad range of widely reported benchmarks, while achieving comparable results to the new DeepSeek v3 on reasoning and coding – at less than half the active parameters. It offers a best-in-class performance to cost ratio with an experimental chat version scoring ELO of 1417 on LMArena. It can also run on a single host!
📌 Previewing Llama 4 Behemoth, our most powerful model yet and among the world’s smartest LLMs. Llama 4 Behemoth outperforms GPT4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Llama 4 Behemoth is still training, and we’re excited to share more details about it even while it’s still in flight.
A big thanks to all of our launch partners (full list in blog) for helping us bring Llama 4 to developers everywhere including @huggingface, @togethercompute, @SnowflakeDB, @ollama, @databricks and many others👏 This is just the start, we have more models coming and the team is really cooking – look out for Llama 4 Reasoning 😉
A few weeks ago, we celebrated Llama being downloaded over 1 billion times. Llama 4 demonstrates our long-term commitment to open source AI, the entire open source AI community, and our unwavering belief that open systems will produce the best small, mid-size and soon frontier models. Llama would be nothing without the global open source AI community & we are so ready to begin this next chapter with you. 🦙
Read more about the release here: https://t.co/7mbK3uggjO, and try it in our products today.
Today is the start of a new era of natively multimodal AI innovation.
Today, we’re introducing the first Llama 4 models: Llama 4 Scout and Llama 4 Maverick — our most advanced models yet and the best in their class for multimodality.
Llama 4 Scout
• 17B-active-parameter model with 16 experts.
• Industry-leading context window of 10M tokens.
• Outperforms Gemma 3, Gemini 2.0 Flash-Lite and Mistral 3.1 across a broad range of widely accepted benchmarks.
Llama 4 Maverick
• 17B-active-parameter model with 128 experts.
• Best-in-class image grounding with the ability to align user prompts with relevant visual concepts and anchor model responses to regions in the image.
• Outperforms GPT-4o and Gemini 2.0 Flash across a broad range of widely accepted benchmarks.
• Achieves comparable results to DeepSeek v3 on reasoning and coding — at half the active parameters.
• Unparalleled performance-to-cost ratio with a chat version scoring ELO of 1417 on LMArena.
These models are our best yet thanks to distillation from Llama 4 Behemoth, our most powerful model yet. Llama 4 Behemoth is still in training and is currently seeing results that outperform GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on STEM-focused benchmarks. We’re excited to share more details about it even while it’s still in flight.
Read more about the first Llama 4 models, including training and benchmarks ➡️ https://t.co/9G3QgVdCkB
Download Llama 4 ➡️ https://t.co/eVomRvEr0w
Llama 4 is a milestone — fast, smart, and open-source.
It’s been incredible working on the vision side of this launch.
Try it now at https://t.co/ETRvMc2OSR
Let’s build the future of AI — together.
Six months ago, I joined the Llama Multimodal team to work on the vision side of the model.
Today, team is launching Llama 4 — redesigned from scratch and natively multimodal.
This is a huge step forward for open-source AI.
We’re releasing:
•Llama 4 Scout: 17B params, MoE, native vision, 10M+ context, runs on a single GPU
•Llama 4 Maverick: best multimodal model in its class — beats GPT-4o & Gemini Flash
•Llama 4 Behemoth (preview): already outperforming GPT-4.5 & Claude on STEM