@0xjohnho@0xSero Only decode is bandwidth bound; prompt processing is still compute bound. And 3090 has a TDP of 350W, also a 3090 doesn't have Blackwell's tensor cores, so you can't run NVFP4 quants correctly like you can with that 4000.
@0xjohnho@0xSero If you throttle it to 150W you're still 80W above with TOPS and TFLOPS reduced to about 950 and 30, respectively, along with 8GB less non-ECC VRAM
@Johannsmitcn4s@Deived3339@p8stie Everyone knows, that's just how LLMs work, they can't keep track of things like that. OpenAI would have to record what kind of person you are to ChatGPT and give everyone's ChatGPT session access to that data, which would land them in a whole heap of trouble.
@AiArtFactory@ronaldmannak 3.8 27B at Q4 is great, I've done a lot of work with it, I've also heard the GSQ-RCO quants are even smaller and still very usable too but I haven't bothered with them