@sama 4o was not just a model for dumb users falling in love with it. Sure they exist and sure they should be protected. The cut 4o / 5.0 was a desaster. Felt like removing your warm neighbour with a cold psychopath. 5.6 is good. But you should have improved 4o.
The model inside the Deepseek app is incredible dumb. Sure, it's free but it failed the car wash test. The test usually only the dumb voice agents fail. But, tbh, their STT is way better than @GeminiApp s. That makes me sad.
@xenovacom It's BS. It's a Qwen Model behind. The real 1 bit approach is Microsofts Bitnet. Train a model from scratch in 1.58 bit. Not quantity it to hell.
@Mr_Salio Well... 2M context window? Then they silently compress it radical so it feels like early 4o with its 16k? Max Token also radical down after some month and memory? Hell. Gemini was flagship of all. Sorry to say, but OpenAI killed them. Even I still hate them for the 4o shutdown.
@googlegemma@cerebras Can't you have a look at Microsofts Bitnet? Or develop an 1 bit Model by yourself? A Gemma 5 50B would fit on a 4070 and soo many people would kiss you (virtually). Pleeeeease ๐ Greetings from my local Gemma 4 26B A4B Hermes Agent in Q2 (which is surprisingly stable for a Q2)
@whyarethis@googlegemma I got OAIs OSS 120B on a 4GB VRAM. Took 20 seconds for a single letter but hey, I got it running ๐ Can't remember, I think 0.00something tokens/second ๐
Guys, if you cry "16GB VRAM no one can afford" local AI isn't for you yet. It's a damn 12B model. In gguf you can use a quant like iQ5 - and it will run smoothly on a 12GB GPU. Even on a 8, if you reduce your KC cache. Thanks Google from my side. Just waiting for the Tiger Gemma for v4 12b from TheDrummer ๐ v3 was mental ๐
@Oracle Trying to get one of your offered free tier slots for 5 days now, even with the github script no success. I am absolutely willing to pay, but not with this service. Fck "Out of capacity".
@wesselsHQ@Google MoE is dumb. Uses only 4B parameter. There's no model to use included this time. I was so excited of Gemma 4 and they just left the 9-13B users behind. Disappointing.
@moonprocelawlp Then I need to apologise. Majority is still complaining, but continuing using the paid models outside free tier. And that's kind of stupid. But you're not. I apologise ๐