Should you preorder the new Mac for local AI?
Be careful about hype vs reality, because in one day the Mac Studio M3 Ultra 512GB dropped from being worth $20,000 to being $3,000 overnight.
I researched every tier of the new Mac mini and Mac Studio so you do not have to. Every RAM SKU, what actually loads, predicted speeds with MTP, and what the same money buys used this week.
I also forward-predicted what you will be able to run once Qwen 3.8 Flash Next and GLM-5.3-Flash land on local stacks, both released today with no Mac benchmarks yet.
I run the daily 27B on a Mac mini M4 Pro 64GB at 27 tok/s (MTP), so the small boxes are measured against a real desk experience.
Rule of thumb: buy the memory that holds the model. The chip name only tells you how fast it streams once it fits.
REMEMBER it's called "unified memory" : macOS eats 8-12GB before your model loads, so subtract that from every RAM number. Every userspace, every chrome instance that runs on the computer will EAT your memory.
WHAT APPLE IS ACTUALLY SELLING
Four different upgrades under one word, AI.
1. Neural Accelerators in every GPU core. Matrix units inside the GPU that speed up prefill (time to first token). Apple's 4.8x (M6 vs M4), 3.9x (M5 Max vs M4 Max), 4.0-4.3x (M5 Ultra vs M3 Ultra) are all prefill, measured in LM Studio prompt processing. My M4 Pro takes ~80 seconds of dead silence before the first token on an 8K system prompt. Neural Accelerators are the fix. If you run agents with real system prompts, this is the upgrade that matters more than the bandwidth number.
2. Memory bandwidth. Generation speed. M6 Mini 170 GB/s (+42% vs M4). M5 Pro 307 (+12% vs M4 Pro). M5 Max 40-core 614 (+12% vs M4 Max). M5 Ultra 1.2 TB/s (+47% vs M3 Ultra). Not 4x. The Ultra is the only bandwidth jump that changes class.
3. Unified memory capacity. Apple did not raise these ceilings. M6: 32GB. M5 Pro: 64GB. M5 Max: 128GB. M5 Ultra: 512GB. They restocked high-RAM options after this year's memory cuts.
4. Everything else. Faster SSD (model load, not tok/s). Thunderbolt 5 clustering (prefill gains, decode can go backwards). M5 Ultra is Apple's first quad-die at 4.4 TB/s inter-die. Wi-Fi 7, Bluetooth 6.
Neural Accelerators shorten the wait before the first token. Bandwidth shortens the wait after it. RAM decides if the model is even there.
USED MARKET
Street prices Aug 25-26.
Used 4090 ~$2,500 (24GB, 1008 GB/s). Wins raw 27B speed against every Mac here, cannot touch 70B.
Used 3090 ~$1,275, two at ~$2,550 hold 70B and decode faster than any 64GB Mac. Loud.
DGX Spark $4,699 list (128GB, 273 GB/s). Two Sparks with DSpark speculative decoding hit 61 tok/s on V4-Flash, competitive with the Ultra 256GB. A year-old $9,400 pair matches a $10,800 Ultra on decode. The Ultra wins on bandwidth headroom.
THE CUDA TAX
Most frontier open-weight models are built and optimized for CUDA first. When they land on Apple silicon via MLX or llama.cpp, the kernels are ports, not native. Quantization formats that fly on Blackwell tensor cores (NVFP4, MXFP4) do not exist on Metal. MoE expert routing that NVIDIA spent years tuning hits generic fallback paths on the Mac.
A model that runs at 60 tok/s on a 4090 might run at 35 on a Mac with the same bandwidth, not because the hardware is worse but because the software was not written for it. MLX is getting better fast, DwarfStar writes native Metal kernels, and the gap is closing.
SKIP THESE
Overpaying for a chip and starving it of memory.
- 16GB M6 · $899 — 8B only. 27B will not load.
- 24GB M6 · $1,099 — 27B knife edge, long context will not fit. Pay $200 more.
- 32GB M6 · $1,299 — ~20GB usable after macOS. A used 3090 has 24GB dedicated VRAM at 936 GB/s, this has 20GB shared at 170 GB/s. Slower and smaller than a $1,275 GPU card. Not a local-model machine.
- 24GB M5 Pro · $1,699 / $1,899 — same problem, faster pipe that cannot help.
- 36GB Max Studio · $2,499 — no 70B. Same price as a used 4090, less pipe.
- 48GB Max Studio · $3,099 — still no 70B. 40-core GPU starved.
- 96GB Ultra · $5,499 / $6,799 — V4-Flash Q4 does not fit. Same money buys the Max 128GB.
THE REAL AI COMPUTERS
Mac mini M5 Pro 48GB · $2,299 / $2,499
27B comfortable (~36GB usable). Predicted ~55-65 tok/s with MTP. Buy it if 27B is the forever default and you want a small always-on box. If 70B is on your horizon, spend the next $400.
I'd go with the RTX 3090/4090 second hand here really. The Ampere chips have excellent compute, and CUDA's ecosystem, offload to system ram for video models. You might as well just pay $400 more here to get to 64GB.
Mac mini M5 Pro 64GB · $2,699 / $2,899
70B Q4 tight (~52GB usable, context eats headroom). Predicted ~30 tok/s on 27B base, ~55-65 with MTP. 70B Q4 ~9-17.
The honest upgrade from my M4 Pro 64GB: decode will not feel new (+12% bandwidth), but prefill will. My 80-second wall before the first token on 27B dense becomes ~20 seconds with Neural Accelerators.
A clean pick for a fleet orchestrator, run 1-2 small models at decent speed. Like the Qwen 3.8 27b Dense or the Orinth 1.5 35B Moe
Mac Studio M5 Max 64GB · $3,499
Fastest way to hold 70B without going to 128GB.
614 GB/s pipe, way better prefill when you paste. Predicted ~58-70 tok/s on 27B with MTP (Youssofal timed 58.7 sustained on the same chip). 70B Q4 ~9-17. Buy it if 70B is the goal and you will not spend up to 128GB.
Mac Studio M5 Max 128GB · $5,099
If you want frontier level, this machine will just miss it.
70B with real context. V4-Flash Q2 (80-91GB) fits with room, DwarfStar already measured 34 short / 26 long. Q4 at ~153GB does not fit. Orinth 1.5 35B MOE will run at easy 80 tok/s and massive concurrency.
I think this one lands on the weird space a single DGX Spark lands on. You have to go to 2-bit for the models you really want. (DSV4F, Qwen 3.8 Flash, GLM 5.3 Flash), and
Honestly 2-bit has too much loss vs the 4-bit versions.
Mac Studio M5 Ultra 256GB · $9,499 / $10,799
The V4-Flash Q4 machine, and MoE is where the quad-die earns its money.
V4-Flash Q4 predicted 40-50 tok/s base, 60-80 with MTP. Base scaled from Viticci's 35 and DwarfStar's 35.5/26.6 on the M3 Ultra, times 1200/819 (1.47x). Not 120. Viticci's 120+ is a prompt-processing multiplier misapplied to decode.
Two Sparks with DSpark already hit 61 on the same model. If MTP on the Ultra closes that gap, this box pulls ahead to 80 tok/s.
GLM-5.3-Flash (320B/18B active, released today, MIT) should fit here at Q4 with ~244GB usable, making this the first box that holds two frontier MoEs at real quants. No Mac benchmarks yet. GLM-5.2 at 4-bit (~466GB) does not fit. 27B and 70B will fly here but that is not why you spend this.
Mac Studio M5 Ultra 512GB · late October · predicted $15,000-$19,000 at 36/80 (1TB)
GLM-5.2 4-bit machine (~466GB, fits with ~500GB usable). K3 Q1_0 at 466GB also fits. MacStories confirmed K3 on a single 512GB Studio at up to 3 tok/s. Every K3-on-Mac speed is a load demo, not a daily model.
GLM-5.2 4-bit predicted 18-28 tok/s, scaled from M3 Ultra numbers (12-19 tok/s) times 1.47x. Buy this for the mid 20s. I think when everything else is combined it probably hits 40 tok/s for 4 bit on GLM 5.2
Qwen 3.8-Flash-Next also fits here easily and should be extremely fast with all that headroom.
HOW I PREDICTED THE ULTRA
Decode scales with bandwidth. Prefill scales with the compute multiplier. I predicted them separately.
Decode. M3 Ultra ran V4-Flash at 35 tok/s on 819 GB/s. M5 Ultra runs 1.2 TB/s. 1200/819 = 1.47x, so V4-Flash Q4 lands at ~51 short, ~39 long. Dense models hit 70-90% of theoretical bandwidth (Contra Collective measured 70B Q4 on M5 Ultra at 21.1 single-die, 27.3 with TP=2). MoE is different: V4-Flash routes to scattered experts across 284B of weights, so the memory controller cannot pipeline reads efficiently. M3 Ultra hits ~28% of theoretical on V4-Flash.
The M5 Ultra's quad-die has more memory controllers across four dies, which is exactly the hardware that helps scattered MoE reads. The 40-50 base prediction uses 28%. If the quad-die improves MoE efficiency, the real number beats 50.
Prefill. Apple's 4x vs M3 Ultra and Viticci's 4.4x on the M5 iPad are the anchors. M3 Ultra V4-Flash prefill was ~449 tok/s. 4x that is ~1,300-1,800 prefill tok/s. That is the Ultra's real new product.
Software caveat. These predictions assume MLX on the M5 Ultra achieves M3 Ultra-level bandwidth utilization. Contra Collective's 70B receipt shows the stack works on the quad-die. But MoE kernels on day one could be less mature. If the launch number is below 40, check back after one MLX update.
Think hard before buying into the hype. You do not want to end up like the people who were told "buy the 512GB M3 Ultra Mac Studio and run Kimi K2.5" and then it ran at 10 tok/s with the CUDA tax on top.
Buy the RAM that holds the model you will actually use, at the speed you can actually live on.
Do not buy local AI hardware with the expectation that the value will hold, buy it because you thought long and hard about what models you will want to run.
Sources in reply 👇
Should you preorder the new Mac for local AI?
Be careful about hype vs reality, because in one day the Mac Studio M3 Ultra 512GB dropped from being worth $20,000 to being $3,000 overnight.
I researched every tier of the new Mac mini and Mac Studio so you do not have to. Every RAM SKU, what actually loads, predicted speeds with MTP, and what the same money buys used this week.
I also forward-predicted what you will be able to run once Qwen 3.8 Flash Next and GLM-5.3-Flash land on local stacks, both released today with no Mac benchmarks yet.
I run the daily 27B on a Mac mini M4 Pro 64GB at 27 tok/s (MTP), so the small boxes are measured against a real desk experience.
Rule of thumb: buy the memory that holds the model. The chip name only tells you how fast it streams once it fits.
REMEMBER it's called "unified memory" : macOS eats 8-12GB before your model loads, so subtract that from every RAM number. Every userspace, every chrome instance that runs on the computer will EAT your memory.
WHAT APPLE IS ACTUALLY SELLING
Four different upgrades under one word, AI.
1. Neural Accelerators in every GPU core. Matrix units inside the GPU that speed up prefill (time to first token). Apple's 4.8x (M6 vs M4), 3.9x (M5 Max vs M4 Max), 4.0-4.3x (M5 Ultra vs M3 Ultra) are all prefill, measured in LM Studio prompt processing. My M4 Pro takes ~80 seconds of dead silence before the first token on an 8K system prompt. Neural Accelerators are the fix. If you run agents with real system prompts, this is the upgrade that matters more than the bandwidth number.
2. Memory bandwidth. Generation speed. M6 Mini 170 GB/s (+42% vs M4). M5 Pro 307 (+12% vs M4 Pro). M5 Max 40-core 614 (+12% vs M4 Max). M5 Ultra 1.2 TB/s (+47% vs M3 Ultra). Not 4x. The Ultra is the only bandwidth jump that changes class.
3. Unified memory capacity. Apple did not raise these ceilings. M6: 32GB. M5 Pro: 64GB. M5 Max: 128GB. M5 Ultra: 512GB. They restocked high-RAM options after this year's memory cuts.
4. Everything else. Faster SSD (model load, not tok/s). Thunderbolt 5 clustering (prefill gains, decode can go backwards). M5 Ultra is Apple's first quad-die at 4.4 TB/s inter-die. Wi-Fi 7, Bluetooth 6.
Neural Accelerators shorten the wait before the first token. Bandwidth shortens the wait after it. RAM decides if the model is even there.
USED MARKET
Street prices Aug 25-26.
Used 4090 ~$2,500 (24GB, 1008 GB/s). Wins raw 27B speed against every Mac here, cannot touch 70B.
Used 3090 ~$1,275, two at ~$2,550 hold 70B and decode faster than any 64GB Mac. Loud.
DGX Spark $4,699 list (128GB, 273 GB/s). Two Sparks with DSpark speculative decoding hit 61 tok/s on V4-Flash, competitive with the Ultra 256GB. A year-old $9,400 pair matches a $10,800 Ultra on decode. The Ultra wins on bandwidth headroom.
THE CUDA TAX
Most frontier open-weight models are built and optimized for CUDA first. When they land on Apple silicon via MLX or llama.cpp, the kernels are ports, not native. Quantization formats that fly on Blackwell tensor cores (NVFP4, MXFP4) do not exist on Metal. MoE expert routing that NVIDIA spent years tuning hits generic fallback paths on the Mac.
A model that runs at 60 tok/s on a 4090 might run at 35 on a Mac with the same bandwidth, not because the hardware is worse but because the software was not written for it. MLX is getting better fast, DwarfStar writes native Metal kernels, and the gap is closing.
SKIP THESE
Overpaying for a chip and starving it of memory.
- 16GB M6 · $899 — 8B only. 27B will not load.
- 24GB M6 · $1,099 — 27B knife edge, long context will not fit. Pay $200 more.
- 32GB M6 · $1,299 — ~20GB usable after macOS. A used 3090 has 24GB dedicated VRAM at 936 GB/s, this has 20GB shared at 170 GB/s. Slower and smaller than a $1,275 GPU card. Not a local-model machine.
- 24GB M5 Pro · $1,699 / $1,899 — same problem, faster pipe that cannot help.
- 36GB Max Studio · $2,499 — no 70B. Same price as a used 4090, less pipe.
- 48GB Max Studio · $3,099 — still no 70B. 40-core GPU starved.
- 96GB Ultra · $5,499 / $6,799 — V4-Flash Q4 does not fit. Same money buys the Max 128GB.
THE REAL AI COMPUTERS
Mac mini M5 Pro 48GB · $2,299 / $2,499
27B comfortable (~36GB usable). Predicted ~55-65 tok/s with MTP. Buy it if 27B is the forever default and you want a small always-on box. If 70B is on your horizon, spend the next $400.
I'd go with the RTX 3090/4090 second hand here really. The Ampere chips have excellent compute, and CUDA's ecosystem, offload to system ram for video models. You might as well just pay $400 more here to get to 64GB.
Mac mini M5 Pro 64GB · $2,699 / $2,899
70B Q4 tight (~52GB usable, context eats headroom). Predicted ~30 tok/s on 27B base, ~55-65 with MTP. 70B Q4 ~9-17.
The honest upgrade from my M4 Pro 64GB: decode will not feel new (+12% bandwidth), but prefill will. My 80-second wall before the first token on 27B dense becomes ~20 seconds with Neural Accelerators.
A clean pick for a fleet orchestrator, run 1-2 small models at decent speed. Like the Qwen 3.8 27b Dense or the Orinth 1.5 35B Moe
Mac Studio M5 Max 64GB · $3,499
Fastest way to hold 70B without going to 128GB.
614 GB/s pipe, way better prefill when you paste. Predicted ~58-70 tok/s on 27B with MTP (Youssofal timed 58.7 sustained on the same chip). 70B Q4 ~9-17. Buy it if 70B is the goal and you will not spend up to 128GB.
Mac Studio M5 Max 128GB · $5,099
If you want frontier level, this machine will just miss it.
70B with real context. V4-Flash Q2 (80-91GB) fits with room, DwarfStar already measured 34 short / 26 long. Q4 at ~153GB does not fit. Orinth 1.5 35B MOE will run at easy 80 tok/s and massive concurrency.
I think this one lands on the weird space a single DGX Spark lands on. You have to go to 2-bit for the models you really want. (DSV4F, Qwen 3.8 Flash, GLM 5.3 Flash), and
Honestly 2-bit has too much loss vs the 4-bit versions.
Mac Studio M5 Ultra 256GB · $9,499 / $10,799
The V4-Flash Q4 machine, and MoE is where the quad-die earns its money.
V4-Flash Q4 predicted 40-50 tok/s base, 60-80 with MTP. Base scaled from Viticci's 35 and DwarfStar's 35.5/26.6 on the M3 Ultra, times 1200/819 (1.47x). Not 120. Viticci's 120+ is a prompt-processing multiplier misapplied to decode.
Two Sparks with DSpark already hit 61 on the same model. If MTP on the Ultra closes that gap, this box pulls ahead to 80 tok/s.
GLM-5.3-Flash (320B/18B active, released today, MIT) should fit here at Q4 with ~244GB usable, making this the first box that holds two frontier MoEs at real quants. No Mac benchmarks yet. GLM-5.2 at 4-bit (~466GB) does not fit. 27B and 70B will fly here but that is not why you spend this.
Mac Studio M5 Ultra 512GB · late October · predicted $15,000-$19,000 at 36/80 (1TB)
GLM-5.2 4-bit machine (~466GB, fits with ~500GB usable). K3 Q1_0 at 466GB also fits. MacStories confirmed K3 on a single 512GB Studio at up to 3 tok/s. Every K3-on-Mac speed is a load demo, not a daily model.
GLM-5.2 4-bit predicted 18-28 tok/s, scaled from M3 Ultra numbers (12-19 tok/s) times 1.47x. Buy this for the mid 20s. I think when everything else is combined it probably hits 40 tok/s for 4 bit on GLM 5.2
Qwen 3.8-Flash-Next also fits here easily and should be extremely fast with all that headroom.
HOW I PREDICTED THE ULTRA
Decode scales with bandwidth. Prefill scales with the compute multiplier. I predicted them separately.
Decode. M3 Ultra ran V4-Flash at 35 tok/s on 819 GB/s. M5 Ultra runs 1.2 TB/s. 1200/819 = 1.47x, so V4-Flash Q4 lands at ~51 short, ~39 long. Dense models hit 70-90% of theoretical bandwidth (Contra Collective measured 70B Q4 on M5 Ultra at 21.1 single-die, 27.3 with TP=2). MoE is different: V4-Flash routes to scattered experts across 284B of weights, so the memory controller cannot pipeline reads efficiently. M3 Ultra hits ~28% of theoretical on V4-Flash.
The M5 Ultra's quad-die has more memory controllers across four dies, which is exactly the hardware that helps scattered MoE reads. The 40-50 base prediction uses 28%. If the quad-die improves MoE efficiency, the real number beats 50.
Prefill. Apple's 4x vs M3 Ultra and Viticci's 4.4x on the M5 iPad are the anchors. M3 Ultra V4-Flash prefill was ~449 tok/s. 4x that is ~1,300-1,800 prefill tok/s. That is the Ultra's real new product.
Software caveat. These predictions assume MLX on the M5 Ultra achieves M3 Ultra-level bandwidth utilization. Contra Collective's 70B receipt shows the stack works on the quad-die. But MoE kernels on day one could be less mature. If the launch number is below 40, check back after one MLX update.
Think hard before buying into the hype. You do not want to end up like the people who were told "buy the 512GB M3 Ultra Mac Studio and run Kimi K2.5" and then it ran at 10 tok/s with the CUDA tax on top.
Buy the RAM that holds the model you will actually use, at the speed you can actually live on.
Do not buy local AI hardware with the expectation that the value will hold, buy it because you thought long and hard about what models you will want to run.
Sources in reply 👇
Codex limits are NERFED.
With my $200 ChatGPT Pro plan I used to never worry about my limits using Codex.
I've been using GPT 5.5 the past 2 days and am already 45% through my weekly limit.
This is a completely different Codex then what we had last week.
Who else is noticing this?