@UnslothAI@Alibaba_Qwen I would download this just to try it in RAM only vs GPU or unified RAM (I don't have any setup like that) ***BUT*** all your HuggingFace downloads are impossible to succeed nowadays. I don't know why. Throttling, pass-through hosting, no idea. Has anyone looked into this yet?
@mr_r0b0t@loktar00 Might stretch longer since the 6000 uses HBM. But not much longer. I think we have 2 more months and then we're in a sticker shock reality state for at least another year. The bizarre part is AI service companies are not delivering much added value since the RAM crisis started.
@OnceGreatAm@BabyD1111229 Often I screw something up like scan an item more than once or forget to scan my card before, then I have to wait for sn employee to come over and fix the setup. I'd rather have robots scanning and checking me out as a cashier than this self-serve BS.
@atharvabuilds@M1Astra I's poossibly not the text changing. It's that all output gets hashed for future lookup like a library where someone feeding text to Anthropic will allow it to spit back if it was generated by it. If they are somehow embedding signals in text it's the dumbest thing I ever heard.
For all the Spark vs GPU vs cloud arguments: I use all 3. I have 3 systems with multiple large GPUs, 2 Sparks, and use cloud models. Where do I fit into the arguments?
@thdxr You're calculating transient non-concurrent usage to 24/7 concurrent usage. You can't compare a single user doing inference for coding 6-8 hours/day sporadically to someone running 6 or more concurrent agentic sessions 24/7.
@bpatters@TheAhmadOsman Also, knowing the quality of local models will never be as good as frontier, just have extra systems running other models to cross-check work. Doesn't have to be the same. Even 85% is still significantly better than where frontier was 18 months ago.
@bpatters@TheAhmadOsman I use 27B for the code and logic. Nemotron Super or Qwen 122B for orchestration. I've not yet done a 2 Spark DeepSeek setup but if it's good at both then my only reason to not use it is based on context, concurrency and long horizon workflow.