I have tested ThinkingCap when it came out, I even had to disable MTP layers as Lucebox uses Dflash. It looked promising initially, but its far more inconsistent than plain Qwen 3.6. Yes there is token efficiency, but running 20-40M token goals I value reliability over efficiency. Btw on old 3090s at 27k context I can run 27B Q6 TG at 60 t/s, Q5 at 86 t/s, Q4 not a quant I would use for real work. Much faster than your M5 Max with 256k context, that is why for M5 gen I only went for M5 Pro as I noticed I have used less of less of M4 Max and my workload has shifted to CUDAs https://t.co/mJJUijUZ3X
Ok tried the updated NVFP4 again of Laguna, much better but so far initial results points to similar results, older and smaller Qwen 3.6 35B is still better. I will keep Laguna for a few more million tokens. Qwen 3.6 27B are still my workhorses running @luceboxai especially with my recent PR that makes 3090s NVLink close to my 5090 at 75% speed (60 t/s vs 80 t/s) but double the context of 256k... 48GB > 32GB at Q6_K.
@mfranz_on Try 27B, it’s better than 35B the V100 has enough grunt to run it. My patch on ik_llama.cpp accidentally made it 20-30% faster for V100 users for 27B. Read my comments on that PR.
@jaita@loktar00 Unfortunately qwen 122b and 397b were only released as 3.5, 3.6 27B is almost as good on my tests and public benchmarks. I do miss both models as back in Feb 2026 I use 397B as my primary smartest local model.
@1337hero What do you think of my GB10? Gigabyte is much more minimalist. I do love Thinkpads, my T520 is still on active duty. 15 years of hard work and now part of a Proxmox cluster. https://t.co/fjHDQwcNEL
I have tested ThinkingCap when it came out, I even had to disable MTP layers as Lucebox uses Dflash. It looked promising initially, but its far more inconsistent than plain Qwen 3.6. Yes there is token efficiency, but running 20-40M token goals I value reliability over efficiency. Btw on old 3090s at 27k context I can run 27B Q6 TG at 60 t/s, Q5 at 86 t/s, Q4 not a quant I would use for real work. Much faster than your M5 Max with 256k context, that is why for M5 gen I only went for M5 Pro as I noticed I have used less of less of M4 Max and my workload has shifted to CUDAs https://t.co/mJJUijUZ3X
Ok tried the updated NVFP4 again of Laguna, much better but so far initial results points to similar results, older and smaller Qwen 3.6 35B is still better. I will keep Laguna for a few more million tokens. Qwen 3.6 27B are still my workhorses running @luceboxai especially with my recent PR that makes 3090s NVLink close to my 5090 at 75% speed (60 t/s vs 80 t/s) but double the context of 256k... 48GB > 32GB at Q6_K.
@rafaelcaricio@MiaAI_lab I better try again. Initially Laguna looked promising but after some real world coding and agentic tasks it was very inconsistent and eventually I couldn’t trust it. Switch back to qwen 35b.
@MiaAI_lab Unfortunately, Laguna S 2.1 118B struggled on an IRL coding task. codex report: "it lost scope and exhausted context. Qwen 3.6 35B finished quickly with coherent edits, tests and a commit—though it missed one regression." both in NVFP4
@MiaAI_lab@NVIDIAAI I find Qwen Agent World to be slightly better. But I dropped to nvfp4 variant as I need the vram for something else. For orchestration and basic terminal use Agent World is better for me.
@ivanfioravanti@pupposandro I have 3x 3090 and DGX Spark. I didn’t understand at first. I use them differently. 3090 with 27b dense model or image gen models. Spark is for 35b or higher sparse/MoE models and always on image generation. I also run different agents. 3090 is worker agent, Spark is scout agent.
@MiaAI_lab Now running Orinth 397B Q6, faster than Minimax M3 and GLM-5.2 using ik_llama.cpp for hybrid GPU+CPU PP: 113 tok/s, TG: 9 tok/s on 8 year old hardware
@ornith_ Is this a fine tune of Qwen 3.5 397B? I hope someone quants it to Q6. I can run this at home, I have also fixed ik_llama.cpp not to crash on 397B for old V100 gpus.
Big changes in AI have happened in just the last few months. About 30 years ago, some of us fought for the open source internet we have today. Like it or not, AI will affect you even more. If you don’t care, at least be aware of it. #openintelligence