Twelve machines. 2.4 terabytes of unified memory. One coordinated inference problem.
Seven DGX Sparks: 896 GB, CUDA, brutal at compute.
Five Mac Studios: 1.5 TB, Metal, brutal at bandwidth and unbeatable per watt.
Plus three Mac minis, a 5080, some V100s.
At 4-bit, that Studio memory holds a 1 to 2 trillion parameter model with room left over for a 900,000-token context. The weights fit. That was never the problem.
The problem is prefill. Before a model with that context says one word, it has to read everything you gave it. On Apple Silicon that is roughly 400 tokens a second. A 900K-token load is over half an hour of silence. Decode is fine, 25 to 30 tokens a second all day, quiet, 300 watts. It’s the first word that costs you.
The Sparks prefill four to five times faster and can’t hold the model. The Studios hold the model and can’t prefill. Everyone with mixed silicon owns both halves of the answer and no way to join them.
So we joined them. NVIDIA prefills, Apple decodes, one request. Two engines that share no cache format, no framework, no vendor. Instead of transferring a cache neither can read, the prefill box computes the decoder’s finished cache using the decoder’s own weights and writes it into the decoder’s prefix store. About 10 KB per token crosses the wire, over plain 10 gigabit Ethernet through the two switches in the first picture. No RDMA, no Thunderbolt.
Measured today on DeepSeek-V4-Flash, 284B, 241,000-token cold load:
Mac Studio alone, 12 minutes to the first word.
Two Sparks feeding it, 3 minutes.
Same prompt again, 19 seconds.
Decode identical. Answers identical.
That ratio is what makes the goal real: prefill 900K on the Sparks for a trillion-parameter model living on five Studios. Tonight we took the prefill window from 262K to 524K. Not finished, and every number gets posted either way.
Why this is a paradigm shift for us: our agent is persistent and has 54MB of memory files..and then we drop a transcript or a codebase on top of it mid-conversation. The wait was the product’s real cost. It isn’t anymore.
All credit to everyone who contributed to these concepts before us. We distill knowledge from all the greats and give credit to all. Standing on the shoulders of GitHub wizards unapologetically without fear of failure or judgement. Local ai must win!
https://t.co/X94wAZsND8
#localai #heterogeneousinference #dgxspark #applesilicon