There are many DeepSeek-on-Spark projects like it, but this one is mine.
3ร DGX Spark. Switchless CX7 ring. Native DS4.1F weights. Upstream-aligned vLLM. Some unique patches.
92 tok/s aggregate at 8ร concurrency.
Now it can be yours:
https://t.co/EOx0AuCX10
Thanks for posting! I still haven't put my 3x switchless spark ds41f repo up.
Every time I come across a repo like yours I take a quick scan for what I can do to improve mine.
Have you tried RoCEnante? I was able to get a couple percent improvement over NCCL for shorter cx lengths.
@Hikari_07_jp I went with the Ugreen 480t and nvme drives. It's super quiet.
Be careful with the OS that these providers ship. For example uGreen made extensions to the operating system that meant mounting it elsewhere wasn't possible without recovery. Just install Ubuntu.
@burkov .5 days? What a timeline we live in :).
Back in the day we would spend weeks to setup something much less ;).
It took me about 1.5 days to get it working, btw.
@8my41@iam_zachi which abliteration are you using? did you find that it was necessary compared to the normal weights? did it have any unintended consequences on intelligence?
@trawasthi_ai Why the gatekeeping? It's great that more people are learning and trying.
The messages, when written by their AIs, do often sound a little over-confident though.
Itโs just for me :). It is literally next to my tv in the living room.
Total throughput divided by concurrent threads is lower than the performance of a single session for sure, though. I think this is โthe nature of the beastโ.
I am using the full weights from DS. No extra quant.