Why is @deepseek_ai v4 Flash so attractive. Here is a new angle on cache hit rates:
What's the difference between cache hit rate 90% v.s. 95%?
It seems quite small just 5% - why should I, as a user care about this?
It's actually way more significant than what you would think! Assuming that the cache hit / miss price difference is 50x
Then for each token you're paying unit_price x cache_hit_rate + (1-cache_hit_rate)*50*unit_price
Let's simplify unit_price to 1, and only draw a chart with cache_hit_rate from 0.95 to 0.99:
We can easily spot that the 5% cache hit rate different is making your token 3 - 4x more expensive!
Making KV Cache small and save them for longer is what DeepSeek has been doing for ages (well, 2 years), thank you DeepSeek!
This is exactly right. Source code is on the verge of becoming like assembly.
The next step is getting rid of “source code” entirely and just making an efficient binary directly with AI.
For a conventional ICE vehicle, manufacturing represents just 10-20% of its lifetime emissions. For a battery EV, production already accounts for 40-50% of its lifetime emissions on a clean grid and 25-35% on an average grid.
That is to say, the EU-produced ICE vehicle will have much higher lifetime emissions than a China-produced EV. Even charging on a theoretical 100% coal-fired grid, the EV wins.
New GGUF(Q4_K):
huihui-ai/Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF
This is an uncensored version of deepseek-ai/DeepSeek-V4-Flash-0731 created with abliteration.
Note These GGUFs support llama.cpp and ds4 and all expert modules were not ablated.
Q4 used a stronger ablation intensity than Q2, so the rejection rate was smaller.
https://t.co/2ZjYQVSlqC
🚨 BREAKING
We're on Hacker News again 🧇
we figured out how to serve Kimi K3 at 3.8x higher throughput and 71% lower cost on @AMD MI355x vs B200
more here: https://t.co/jHnWGkWAVo
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: https://t.co/iwjmQxLBof
✅ Qwen Studio: https://t.co/4V2pFvDovG
⚡ API: https://t.co/gAGqaLQGbN
@ArenaTiziano@antirez I downloaded DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf, and somehow it seems to be working on my macbook as well. So I wonder what the differences are between the two files. Thanks.
Here is my AI investing guide.
Sitting here August 2026, my current best thoughts are as follows:
1. LPS (Land Power Shell) is still the most obvious and fastest path to cash on cash returns. Lots of value can be assembled and traded quickly at this layer. And as data centers get more pushback, energized land can explode in value. Very bullish here.
I’ve stepped into this layer very aggressively. My partner @anitavlallian and I have acquired almost 6GW coming online in a ramp from today thru 2029 of grid power and behind the meter.
2. Silicon - I helped get @GroqInc off the ground in 2015 and we licensed it to @nvidia for $20B Dec2025. I won’t invest or incubate anything in this layer now. The perf demands of the chips are too high, manufacturing precision is too complex and supply chain influence to get adjacent components like memory isn’t possible for a startup anymore. Lots of capital will be wasted here chasing Groq and Cerebras’ success. Note that both startups made sense a decade ago when these constraints were much more modest.
3. Clouds - Clouds are very very lucrative but very hard to build and very expensive and technically complicated to maintain. And as alignment becomes a more important issue, I expect the clouds will be asked to build robust KYC and attest to it. This makes the risk:reward ratio skewed. I don’t want to be responsible when the USG says a cloud allowed a bad actor to do something bad because of poor KYC.
4. Models are complicated. The big open question is how much of the revenue being generated by them today is because of tokenmaxxing and poor model behavior. If it’s a lot, then the annualized revenues will diminish meaningfully even as token consumption inflects upwards. This is the big economic question at this layer.
5. Harnesses are where the action is and why I started @8090solutions two years ago. In a nutshell, the harness helps enterprises owns their proprietary context (what Alex Karp calls their ‘alpha’). This is an enterprise’s data, workflows, evals, and business rules. A harness that gives this to an enterprise is what creates very low model-agnostic switching costs, which further reinforces my views of #4 above.
6. Applications will be another long term winner along with harnesses. This is where the differentiation between “off the shelf” and “custom time and materials” melts away. Every company, with the right harness, can now imbue their alpha into the software that runs their company. I expect this to mean that “off the shelf” is largely replaced with custom software creating a huge opportunity to write these solutions for companies. Build once and sell repeatedly is a laggard GTM motion for a SaaS world that isn’t needed here. Think custom by design, alpha embedded, proprietary by nature.
Fin.
Good luck to all the players!
DwarfStar branchk "ds4f-mxfp4" now can run the lossless MXFP4 DeepSeek v4 Flash GGUF I published on my Hugging Face account. It rocks even with SSD streaming in 128GB systems at > 20 t/s in case you want to try the *actual* DS4F weights released without any quantization.
let’s put this 80% price drop for 5.6🌙 in perspective
three weeks ago, gpt-5.5 xhigh was our frontier.
it completed 67% of tasks on DeepSWE.
today, luna max matches that score at about $0.12 per task instead of $7.23.
same score.
roughly 60x cheaper.
three weeks later
Wow @deepseek_ai rocks.
DeepSeek-V4-Flash (284B total / 13B active MoE) received a targeted post-training refresh on 31 July. Architecture/size are unchanged from the April preview: hybrid Compressed Sparse Attention + Heavily Compressed Attention, Manifold-Constrained Hyper-Connections, + FP4 experts. At 1M context the design still delivers ~10% of V3.2’s single-token FLOPs & ~7% of its KV-cache footprint.
The 0731 checkpoint moves the needle on agent benchmarks w/out touching activated parameter count:
•Terminal Bench 2.1 → 82.7
•Cybergym → 76.7
•Toolathlon-Verified → 70.3
•DSBench-FullStack / Hard → 68.7 / 59.6
Reasoning modes (Non-think / High / Max) remain, with Max still requiring ≥384k context. API pricing stays at $0.14 / $0.28 per million tokens (cache hit input $0.0028). Insanely cheap!! Native Responses API and Codex adaptation are now present.
For any workload that keeps long contexts resident - multi-step coding agents, repository-scale analysis, tool chains - the efficiency delta is the practical variable. 13B active parameters at M-token length alters both memory hierarchy pressure & the power/cooling envelope of the serving cluster.
Conclusion: this magic 👇 at cabbage 🥬 price. 🔥🔥🔥⚡️⚡️⚡️💙💙💙
The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code:
- 61k in Kimi Code
- 67k in Hermes
- 340k in Claude Code
At K3's $3/M input rate (input tokens make up roughly 95% of agentic workloads), the average cost per task was:
- $0.22 in Kimi Code
- $0.28 in Hermes
- $2.00 in Claude Code