@andrewho03 Ran into this problem myself with an "additional comments" section on a spreadsheet (That I hand reviewed every line!!) of instruments. No, Luna, we do not need to note that the serial number had a slight scratch.
@ollama Fix the quantization you guys are using on GLM 5.3 Flash. I've used IQ4_XXS locally and even then it never loops like this. What in the world are you guys using?
@ArtemisI0@owl_posting https://t.co/og9xpeUMFC
Yes and no. Plenty plenty PLENTY of stuff that will remain for lab techs.
But the data for biology AI models is SUPER limited and bottlenecked by the cost. High throughput methods like this will change the game for training data in the future.
@CultOfRoyal@SIGKITTEN Maybe you could improve it, but from my perspective, either create something actually worthwhile or use the AI to tutor yourself. By offloading entirely and using tokens entirely as a crutch, that demo prototype is as good as it gets.
@CultOfRoyal@SIGKITTEN Yeah, I guess if this were a teenager that made it, I'd make some suggestions along the line of "You gotta take a creative direction, because this looks like beige." But, since this is a large language model we are talking about. Good luck prompting artistic direction.
@CultOfRoyal@SIGKITTEN Nah. IDK how to put my finger on it, but nearly every aspect of this is just off. I'd rm rf the repo, and use Astra to made a vs. multiplayer mod on top of FTL. Then you'd actually have a fun game at the end of the day, not irreparable slop.
@SIGKITTEN https://t.co/Sokk1DeRIo
I don't know how to explain it to you but its the same ship and same humanoid shape with like 3 polygons and a coat of paint. Compare to the real handcrafted deal:
@Presidentlin It also really reminds me of the GPT-5 release that one also apparently blew his mind as well, and then he had to really walk it back in the wake of a horrible launch.
@1kartikkabadi1@OmedVibeCodes It makes sense until you start comparing $20 on Ollama Tokens to that same $20 on Openrouter. Going price for Deepseek v4 Flash 0731 is $0.05 / $0.16per 1M, whereas Ollama CLoud prices the same at $0.44 / $1.32. You get a 3x discount on a 8x more expensive product.
@1kartikkabadi1@OmedVibeCodes I agree! But my weekly limit contained 4k GLM 5.3 Flash requests and 5k Deepseek v4 Flash requests, now I hit the limits with only 3k. Perhaps its a symptom of longer context workloads, but I'm deeply suspicious of the subscription's value now.
@buildwithhassan The part you're missing is they are above market prices for many comparable models on openrouter. Grateful I got grandfathered in with the old pricing structure, but the future is clear. What I need mocing forward is to run these models locally with a DGX Spark or Max Studio.
@ollama My line of thinking is that for many tasks I don't gain economic value by having them done faster. I would prefer a node thats slower but amortizes the node across more people, even if it crawls at 10 TPS or enforces 256k max context.
@ollama Ouch! The Deepseek V4 Flash 0731 pricing is nearly 5x what slower competitors charge!
If its not too much to ask, would you guys look into a slower endpoint in with less interactivity but better economics for long-horizon tasks? Same thing with GLM 5.3 Flash.
@adibvafa You can do /swarm on in Kimi Code. Pretty useful now that we have cheap models like Deepseek V4 Flash that can do genuinely decent work. But I have a hunch its not a time-saver, since you have to wait for ALL subagents to be done rigidly, it feels more like hurry-up-and-wait.
@ollama You guys seem to be pretty responsive tonight, but I'm curious, why is Gemma 31B served, but not Qwen 3.8 27b as well? Is it a lack of capacity?