Hm, good question... I think both things can be true.
HBM supply is already massively strained so everyone is basically get the scraps. Anything that has memory continues to get more expensive (which is pretty everything lol).
Unless NVIDIA has unlocked extra supply, they'll have to ration the scraps between the DGX Spark and the RTX Spark (i.e. reduced memory in the Spark) or choose to only manufacture one of the two.
Either way, long-term, because of memory supply chain constraints AND the reasons you, yourself stated, the DGX Spark will most likely appreciate in value 🙂
One of the most important graphs when picking a model for coding 👀
Omniscience score measures accuracy vs hallucination (higher is better).
You don't need the latest and greatest, for some languages!
@alexocheema@0xSero any chance you guys can let me in? 🙏🏼 I had a few friends try to register but they got stuck on the waitlist so I haven't quite made it in yet
got a spark recently and itching to see what sorts of results people have measured for various models 👀
Have you ever thought about how easy it would be to hack you?
To get into your accounts, mess with your finances, disrupt your business, leak customer data, reveal personal information, etc?
Well however much you thought about it before, you should think a lot more about it now.
Unrestricted open source models are going to be as good as Sol / Mythos / Astra within 3-12 months, and that's basically going to be a Thanos Guantlet situation.
"Ruin this guy."
*snap*
If someone utters your name and says, "Ruin their life."
*snap*
And all the power of Mythos++ is unleashed on every attack surface you have in life, how would you hold up?
You should start thinking about it.
We all should.
Low effort approach (with risks): stick the machine on it's own zone/VLAN + invite folks via teleport. Edit Zone Based Firewall rules so that it can't talk to anything else + tighten up mDNS/UPnP/IGMP Proxy network settings (might cause some issues). Assuming you're granting them root access to the baremetal device if it has wireless radios, it's technically possible to execute bluetooth/wifi attacks regardless of your firewall rules.
High effort (close to bulletproof):
1. Still put it on it's own zone/VLAN
2. Install Proxmox VE on it
3. Setup a VM and pass all the connect GPUs through it
4. Disable wifi/bluetooth at the VM level (they can't re-enable with sudo even if they wanted to)
5. Configure VM + Datacenter level firewall to make it impossible for anything else but SSH/Web traffic
6. Harden sshd
7. Provision access via tailscale with ACL restrictions
Bonus:
- you can setup hourly/daily/weekly snapshotting in case someone messes up the box and you want to restore it to a certain point
- VM acts an extra layer of security by having its own separate kernel
If these are trusted friends you've known for years and are cybersecurity savvy, I'd go with the first approach. Otherwise, I'd do #2, especially given recent AI-driven cybersecurity incidents. 🙂
I'm curious what @NetworkChuck and @DanielMiessler would do...
When you combine this with the fact that the base models that are being quantized are getting smarter from day one, useful local intelligence really feels accessible.
Looking forward to contributing to the space as I learn more about inference engineering 🙂
It’s been really inspiring watching the local AI community over the past few months.
One person quantizes
One person fine tunes
Another person applies a new technique
Then another person incorporates their method
You end up upwards of 20% performance gains left and right!
It’s been really inspiring watching the local AI community over the past few months.
One person quantizes
One person fine tunes
Another person applies a new technique
Then another person incorporates their method
You end up upwards of 20% performance gains left and right!
@theo I doubt there will be a compute crisis for long, everything is getting more efficient and there’s new players in the game.
There’s been a boom in demand, people were unprepared, as usual people take on the mantle and get to producing what we need.
I’m very optimistic about it
💯It is a tool. You always have a choice between cognitive surrender and cognitive agency. The challenge is that society at large has poor cogsec as a result of B2C/social media economics have shaken out.
Energy
- way more intelligence per watt
- don’t have to build anything
- sips power compared to the 300W PER 3090
- my office won’t sound like a data center/need AC 🤣
Recently came to the decision to buy a single DGX spark after spending some time considering whether or not to turn my current single RTX 3090 setup into a cluster by adding ~4 more for the same price.
Reasons why I went for the Spark instead:
I’m seeing more and more 3090’s pop up for sale
Everyone is coming to terms that they are going to need more than one spark
3 sparks or this? What would you buy?
Functionality
- can hypothetically disaggregate refill and decode between the two
- gives a dedicated box for MoE models in addition to my current RTX for dense models
- adding a 2nd spark down the road makes DSV4F 0731 at full precision easy
- latest Blackwell architecture with support for NVFP4
@tmophoto Personally - sparks. RTX 3090s have an edge in terms of memory bandwidth. But I can’t imagine running a rig like that 24/7, at least where I live.
At 6 cards (300W) + that CPU, you’re pulling 2100W per hours. Also I imagine it’s loud/hot in whatever room it’s in.
Hbu?