@jun_song First: I agree. Second: It resonates a bit with how I understood wikipedia decades ago - and look what that has turned into. So Third: How do we make sure Open Source gets as good as it can and we all hope for?
@0xSero money would not be a problem it honesty would not normally get diluted by it. that rarely happens irl. so if in doubt, aim for trust instead. me and many others honor that.
@kyzoroX Yes, thats why I thought about splitting tasks into en/decode heavy on 30k feet, then only transmit the results of tasks, not the full kv cache. maybe a little shared knowledge via qdrant or so. a bit like committee of autonomous agents vs true distributed inference.
@h100envy Much appreciated! Comparing to real life, how do you weigh the person that always contradicts 'because', the person that is distracted, the person that tries to please everyone? can we overlay, say, DISC model on these personas?
@ksuniri@jun_song workload dependent. encode/prefill -> spark, decode/gen -> ultra. also better power efficiency and fuller OS on Mac, better compatibility to DC workloads on nvidia. And what do you mean by 'or'? π
@jun_song I don't agree. Had they stated they would not work with your data, fine. But they didn't. It was clear in the open for anyone who cared. Yes, a different default would have been better, but loosing credibility because users don't check basic settings? I don't know.
@SpaceXAI Yeah, if you 'OK' every step of the installation, do not RTFM / EULA / default settings, that's what you kind of get. scandal where? Sounds like a hitpiece fabricated to fling excrements now that grok 4.5 gains traction.
@selfhosted_ai@digitalix@ivanfioravanti depends on the workload. prefill heavy, spark. decode heavy, ultra. i decided to split things on the task level and delegate accordingly.