@spiritbuun Thank you for the great work! I happened to bump onto this repo randomly from a tweet and I'm enjoying it so far.
JFYI: I tried your new soft cache mode suggestion from the new PR but auto cache mode is still the faster decode path for me on my 2x RTX 3080 system
GLM-5.3-Flash is surprisingly resilient and performant in my tests so far even at Q2_K (custom recipe) sized 111GB
For the below result, it reasoned for about 120k tokens and analysing those traces shows no sign of hallucination or repetition. Quite solid!
This is not being talked about enough! Getting useful results with free web search for agents was always a major handicap on my local workflows.
Enter @Tiny_Fish!
COMPLETELY FREE Search + Fetch for agents
30 Search requests/min = up to 43,200 searches a day!
We just killed Exa, Tavily, SerpAPI, and Brave.
Your agent can now search & fetch any webpage for 100% FREE.
Them: $7 per 1,000 searches.
Us: $0. No subscriptions, no quotas.
Humans search Google for free. Agents shouldn't have to pay either.
Made possible by @Tiny_Fish and @MonidHQ.
I'm seriously starting to question whether I need a cloud subscription anymore...
Was already impressed with 3.8 27B and DSV4-Flash but now it feels like we're stepping into a realm where these things can potentially replace cloud for most.
(As per the benchmarks at least!)
I absolutely love how GLM has finally released a smaller model with such fantastic capabilities!
However, I find Qwen-3.8-Flash-Next even more impressive at roughly half the size and seemingly not that far behind in terms of capabilites!
Huge moment for local AI!
I absolutely love how GLM has finally released a smaller model with such fantastic capabilities!
However, I find Qwen-3.8-Flash-Next even more impressive at roughly half the size and seemingly not that far behind in terms of capabilites!
Huge moment for local AI!
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
The next-gen architecture powering Qwen4 is now here! โจ
Get ready for the open release of Qwen3.8-Flash-Next ๐
The countdown starts now! โณ๐ฅhttps://t.co/0Zlesgxnot
@ZixuanLi_@Zai_org Would be great if we could celebrate this moment with a follow-up Air model release :)
I genuinely notice a lot of people still resorting to the GLM 4.5 Air to do stuff to this day!
HUGE if true!
Multimodal support
+
Considering how they've been able to compete with multi-trillion parameter models at ~750B, a value-packed flash model sized under 300B and preferably at ~200B would make it viable for unified memory systems and most prosumer setups!
A new Kimi model, likely K3.1, is now being tested on the Code @arena under the name "korrine"
K3 was tested on the Arena as "kivine" prior to its launch
If anyone's wondering, "Ox Alpha" on OpenRouter is the upcoming GLM 5.3 Flash from fellow Chinese lab Zhipu
Tried the new DFlash2 llama.cpp PR on my 2x RTX 3080 20GB system and I roughly get about 1.4-1.5x decode speedup when acceptance is high against baseline with MTP heads.
Any speedup is a win and the PR hasn't fully matured yet ๐ค
https://t.co/VfCeucI0mG
As someone who has no creative experience, Minimax H3 is such a goated local model!
I've been using it for more than just generating videos. Once you figure out its prompting style (or have an LLM for that), you can also use it to create/edit images, render text, audio and SFX!
Wow! Artificial Analysis Intelligence Index ranks Qwen 3.8 27B right above DeepSeek V4 Flash (max) and Gemini 3.6 Flash!
A model running locally on consumer machines that outperforms Gemini 3.6 Flash!
This also aligns with my private eval results.
https://t.co/w8JbKiaszH
Shocking find: Qwen 3.8 27B (xHigh) vs DeepSeek V4 Flash 0731 (max) is VERY CLOSE in my private eval! I honestly did not see this coming.
Qwen 3.8 27B scores 106/115 and DeepSeek barely edges it with 107/115
15 hours spent. Both expend roughly the same overall tokens but..(1/2)
inclusionAI's Ling 3.0 Flash 124B/A5.1B MoE and Ling 3.0 Tiny 7.9B/A1.3B MoE are now usable with llama.cpp!
GGUFs available at: https://t.co/I6b04YbBJ1
Continuing the list:
- DeepSeek has more knowledge and is noticeably better at creative writing.
- For unified-memory setups that can run DSV4-Flash faster than Qwen, DeepSeek's the better choice
- For setups that rely on CPU-offloading, Qwen runs faster and is ~90-95% as good!
Shocking find: Qwen 3.8 27B (xHigh) vs DeepSeek V4 Flash 0731 (max) is VERY CLOSE in my private eval! I honestly did not see this coming.
Qwen 3.8 27B scores 106/115 and DeepSeek barely edges it with 107/115
15 hours spent. Both expend roughly the same overall tokens but..(1/2)
DeepSeek runs 3x slower for me due to offloading.
Key takeaways:
- Qwen was run at Q6_K_XL and DeepSeek at a custom Q3_K recipe roughly equivalent with UD-Q3_K_XL
- My earlier theory about DeepSeek being time & token-efficient despite being slower is now objectively proven wrong