Qwen3.8-Flash-Next might not run well on one dgx spark at all π¨
it has a huge 51B Engram table that even when compressed to 4 bit would use around 25gb of RAM, but because of DGX spark unified memory the total the entire model would use 105β125GB of unified memory total making it very close
@maria_rcks fair lol it is extremely token inefficient making it extremely expensive with intelligence literally worse than Gemini somehow on AA which is why I think it belongs down there with Google
βquietly deleted their postsβ meanwhile them apologizing and publicly saying they deleted it lmao. This guy sketches me out, constantly hating on people and glazing Elon like what?