📣 Bambu Farm Manager is here!
A streamlined, cloud-free tool to manage large printer fleets directly over a local network.
Now available for X1C, P1, and A1 – for free!
Learn more 👇
https://t.co/xT55Fs0YeZ
#BambuLab#PrintFarm#3DPrinting
Meta 🤝 Apple
Llama 4 + Apple Silicon is a match made in heaven.
Here's why: Like DeepSeek V3/R1, all of the new Llama 4 variants are massive sparse MoE models. They have a massive amount of parameters, but only a small number of those are active each time a token is generated. We don't know in advance which parameters are going to be active so all parameters need to be ready in high-speed GPU memory.
GPUs have fast memory but GPU memory is expensive. Apple Silicon, however, uses Unified Memory and UltraFusion to fuse dies - a tradeoff that favors a large amount of medium-fast memory at a cheaper cost.
The M3 Ultra Mac Studios released 1 month ago push this all the way to 512GB of Unified Memory. However, pushing the memory this far means memory bandwidth lags behind. For the 512GB model, the memory refresh rate (the number of times per second the GPU can cycle through all its memory, or simply the ratio of memory bandwidth to memory) is only 1.56/s. Comparison with other hardware:
NVIDIA H100 (80GB): 37.5/s
AMD MI300X (192GB): 27.6/s
Apple M2 Ultra (192GB): 4.16/s (9x less than H100)
Apple M3 Ultra (512GB): 1.56/s (24x less than H100)
Ideally the properties of the workload match the properties of the hardware. Otherwise, either the hardware is over-provisioned for the workload (wasteful) or under-provisioned for the workload (hardware-bottlenecked). The equivalent property to look at for the workload (in this case batch_size=1 inference) is the model sparsity. Sparsity is defined here as 1 - (Active parameters / Total parameters). A dense model has sparsity 0% (since Active Parameters = Total Parameters). Sparsities for different models:
Llama 3.3 405B: Total=405B, Active=405B, Sparsity=0%
DeepSeek V3/R1: Total=671B, Active=37B, Sparsity=94.4%
Llama 4 Scout: Total=109B, Active=17B, Sparsity=84.4%
Llama 4 Maverick: Total=400B, Active=17B, Sparsity=95.75% (!!)
Llama 4 Behemoth: Total=2T, Active=288B, Sparsity=85.6%
In general, higher sparsity is a better fit for Apple Silicon because of its lower memory refresh rate. Clearly then, Llama 4 Maverick is the best fit for Apple Silicon.
Llama 4 Scout and Behemoth have a lower sparsity. Also, Behemoth is so large (2T/2,000B parameters) that it requires running on multiple macs. It would require >8 M3 Ultra 512GB Mac Studios to fit it all in memory at fp16. Normal pipeline parallel is bottlenecked by the speed of one machine, in this case 800GB/s which would run the full model at fp16 at only 1.39tok/sec (basically unusable!!).
Now, that doesn't count Apple out. Keep in mind that Apple Silicon would be the most cost effective solution for running the model since unified memory if much cheaper per GB than GPU memory:
NVIDIA H100: 80GB @ 3TB/s, $25,000, $312.50 per GB
AMD MI300X: 192GB @ 5.3TB/s, $20,000, $104.17 per GB
Apple M3 Ultra: 512GB @ 800GB/s, $9,500, $18.55 per GB
Consider that it would require 50 H100's just to fit the entire (fp16) Behemoth model in GPU memory. That would cost $1.25M. With MI300X, it would cost $420,000. With M3 Ultras, it would cost just $76,000.
There is a solution here: better distributed parallelization strategies for running MoE models at high batch_size=1 TPS on macs connected with Thunderbolt 5.
Currently the only feasible distributed parallelization strategy for Apple Silicon is pipeline parallel because it requires just one network hop for each connected mac (e.g. 2 hops for 2 macs). However, other parallelization strategies like expert parallelism or tensor parallelism require many network hops per layer of the model. Llama 4 Maverick has 48 layers, meaning >48 network hops per token generated. Thunderbolt 5 empirically has a latency of ~0.5ms between macs meaning at least 24ms of additional latency per token generated. Behemoth will likely have many more layers, probably more than 100 meaning at least 50ms of additional latency per token generated.
The bottleneck is latency and @exolabs have been experimenting with solutions here (details soonTM). A setup with 10 M3 Ultra 512GB Mac Studios could run the Llama 4 Behemoth (fp16) at a theoretical maximum of 27 tok/sec. The actual achievable number is probably somewhere between 10 and 20, depending on the constraints of inter-mac latency.
Here's a full breakdown of theoretical performance:
Llama 4 Scout: 1 x M3 Ultra 512GB Mac Studio, $9,500, 23 tok/sec (pipeline parallel)
Llama 4 Maverick: 2 x M3 Ultra 512GB Mac Studio, $19,000, 23 tok/sec (pipeline parallel), 46 tok/sec (experimental advanced parallelization with @exolabs - theoretical maximum)
Llama 4 Behemoth: 10 x M3 Ultra 512GB Mac Studio, $95,000, 1.39 tok/sec (pipeline parallel), 27 tok/sec (experimental advanced parallelization with @exolabs - theoretical maximum)
Introducing our first set of Llama 4 models!
We’ve been hard at work doing a complete re-design of the Llama series. I’m so excited to share it with the world today and mark another major milestone for the Llama herd as we release the *first* open source models in the Llama 4 collection 🦙. Here are some highlights:
📌 The Llama series have been re-designed to use state of the art mixture-of-experts (MoE) architecture and natively trained with multimodality. We’re dropping Llama 4 Scout & Llama 4 Maverick, and previewing Llama 4 Behemoth.
📌 Llama 4 Scout is highest performing small model with 17B activated parameters with 16 experts. It’s crazy fast, natively multimodal, and very smart. It achieves an industry leading 10M+ token context window and can also run on a single GPU!
📌 Llama 4 Maverick is the best multimodal model in its class, beating GPT-4o and Gemini 2.0 Flash across a broad range of widely reported benchmarks, while achieving comparable results to the new DeepSeek v3 on reasoning and coding – at less than half the active parameters. It offers a best-in-class performance to cost ratio with an experimental chat version scoring ELO of 1417 on LMArena. It can also run on a single host!
📌 Previewing Llama 4 Behemoth, our most powerful model yet and among the world’s smartest LLMs. Llama 4 Behemoth outperforms GPT4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Llama 4 Behemoth is still training, and we’re excited to share more details about it even while it’s still in flight.
A big thanks to all of our launch partners (full list in blog) for helping us bring Llama 4 to developers everywhere including @huggingface, @togethercompute, @SnowflakeDB, @ollama, @databricks and many others👏 This is just the start, we have more models coming and the team is really cooking – look out for Llama 4 Reasoning 😉
A few weeks ago, we celebrated Llama being downloaded over 1 billion times. Llama 4 demonstrates our long-term commitment to open source AI, the entire open source AI community, and our unwavering belief that open systems will produce the best small, mid-size and soon frontier models. Llama would be nothing without the global open source AI community & we are so ready to begin this next chapter with you. 🦙
Read more about the release here: https://t.co/7mbK3uggjO, and try it in our products today.
NEWS: Humanoid robot company Figure released a new video of its robot working on BMW's production line.
"This isn't a test environment—it's real production operations."