@thereisnobeth bro have you run any of your thinking through an LLM? There are actually ways to calculate how much heat you can dissipate with what size radiator in space (as EM radiation) it does work and they’ve thought this through
@shiraeis Nah there’s a latency throughput tradeoff to batch sizes and the way you parallelize across GPUs, speculative decoding barely adds any cost and I’m sure they’re doing that for all models
@xw33bttv Wdym it didn’t pay off? Codex 5.2 is a beast (WeirdML seems pretty representative of the more technical coding it can do) I’m sure they’re making quite the profit off of it. Let them cook.
@esrtweet Consumer AI isn’t where the money is though, money is in, say, fusion and drug design and eldritch sci fi tech. Doesn’t matter if it runs on the average person’s phone, average person doesn’t have money.