Again - people miss the point of the Spark. Either deliberately because they want clicks, or through ignorance. Sparks are not top line inference machines. They were designed as - and are brilliant as - prototyping and experimental benches to build on the nVidia enterprise stack. You build on Spark and it translates directly to the nVidia datacenter stack. They are relatively cheap to buy and run. It is a happy accident they are being used for local inference by an enthusiastic and creative community. No one has ever argued a Spark is faster than an xx90 or RTX pro. But 128GB is no joke and allows you run multiple models easily (robotics sim, embedding, rerank, creative workflow experiments, etc) which - again - is the purpose the Spark was designed for. The notion of local inference being done on an xx90 at more than a hobby level is also ridiculous. I have a 5090 (and 4 Sparks) and that thing, while extremely fast, consumes an enormous amount of power during any long run. 32GB is also a serious bottleneck. So yeah - if running a 27B or similar class MoE is your bag and you don’t game or do imaging/video work that compete for those 32GB? Go nuts! It is really fast and a great product. The RTX Pro 6000 = 5090 on steroids. But it is pushing the same price as 3 or more Sparks for 96GB. The reality is the memory capacity and flexibility are worth more than speed in most cases. If you simply want to run one model - and it fits cleanly with KV and desired context window inside 96GB you would be a fortunate person indeed to have a Pro 6000. That does not, in any way, discount the value of the DGX Spark. Most people aren’t following any conversation with an agent at 200 tokens/second and multi agent work is truly fantastic on a Spark. I am hoping for a Pro 6000 next month and, with a bit of luck this Fall - a shiny new Mac Studio from Apple with a wallet-shattering amount of memory and a 5 Ultra. I will still be using my Sparks happily every day even with those systems. For exactly what they were designed for… experimentation and prototyping of datacenter-ready applications and loads. And if I have my agents plugged into a local version of Ox Alpha, @MiaAI_lab ‘s DS4 or some other model for a spell? Well I guess I’ll be a happy idiot in the eyes of some. An enthusiast, hobbyist, and someone who loves his little golden boxes by others. And that’s just fine by me. That second group of people are a lot of fun and I’d rather hang out with them anyway…
It’s fat. Yes - i am going there… Skill clipping and management should be automatic and I find I’m constantly resetting configurations after updates. And too much memory discipline needs to be built and constantly monitored rather than simply ‘being there’. The functionality is there but enforcing it is sometimes inconsistent.
Breaking News 2027! - White House says ‘Jensen - no more open weight Chinese models on platforms owned by USA companies.’ I think this is really bad long term. And even if I trust nVidia, the USA gov has already shown a willingness to interfere in this market, and has been consistently in conversations about limiting it over the last six months. If nVidia is sincere they’d buy it and set it up offshore as a trust - out of reach.
Again - people miss the point of the Spark. Either deliberately because they want clicks, or through ignorance. Sparks are not top line inference machines. They were designed as - and are brilliant as - prototyping and experimental benches to build on the nVidia enterprise stack. You build on Spark and it translates directly to the nVidia datacenter stack. They are relatively cheap to buy and run. It is a happy accident they are being used for local inference by an enthusiastic and creative community. No one has ever argued a Spark is faster than an xx90 or RTX pro. But 128GB is no joke and allows you run multiple models easily (robotics sim, embedding, rerank, creative workflow experiments, etc) which - again - is the purpose the Spark was designed for. The notion of local inference being done on an xx90 at more than a hobby level is also ridiculous. I have a 5090 (and 4 Sparks) and that thing, while extremely fast, consumes an enormous amount of power during any long run. 32GB is also a serious bottleneck. So yeah - if running a 27B or similar class MoE is your bag and you don’t game or do imaging/video work that compete for those 32GB? Go nuts! It is really fast and a great product. The RTX Pro 6000 = 5090 on steroids. But it is pushing the same price as 3 or more Sparks for 96GB. The reality is the memory capacity and flexibility are worth more than speed in most cases. If you simply want to run one model - and it fits cleanly with KV and desired context window inside 96GB you would be a fortunate person indeed to have a Pro 6000. That does not, in any way, discount the value of the DGX Spark. Most people aren’t following any conversation with an agent at 200 tokens/second and multi agent work is truly fantastic on a Spark. I am hoping for a Pro 6000 next month and, with a bit of luck this Fall - a shiny new Mac Studio from Apple with a wallet-shattering amount of memory and a 5 Ultra. I will still be using my Sparks happily every day even with those systems. For exactly what they were designed for… experimentation and prototyping of datacenter-ready applications and loads. And if I have my agents plugged into a local version of Ox Alpha, @MiaAI_lab ‘s DS4 or some other model for a spell? Well I guess I’ll be a happy idiot in the eyes of some. An enthusiast, hobbyist, and someone who loves his little golden boxes by others. And that’s just fine by me. That second group of people are a lot of fun and I’d rather hang out with them anyway…
@MiaAI_lab@plotarmordev Great - thank you Mia! I’ll have 4 by the end of next week. So looking forward to whatever Ox Alpha is going to be - as long as it can run on sparks 😁