@cozybearlog I tried Q4_K_XL GGUF, updated the halogen, and changed the harness from OMP to opencode2. So far, everything is working flawlessly, but I think more testing is needed. Also, while Q4 is slower, I'm satisfied with its quality, which is better than the original halogen.
qwen3.8 flash next q4-k_xlでunslothさんとこのmtpをprのバッチ当ててllama.cppで試してみた。ついでにrocm10で。長いエージェントループだとバラツキは大きいが平均のppが数十から200。tgは9-17。mtp無しのx1.3-1.7というのはリアルな感じ。タスク完了まで数時間かかかるので3割減るのは大きい