If DeepSeek V4 Flash really sweeps Opus 5 and GPT-5.6, the benchmark war just got awkward. Is this open-model efficiency beating raw scale, or is the gap still cherry-picked?
Golden hour hits different when the drink's this cold ☀️🥤 slid into a street-side café and handed you the coldest one of the afternoon — no filter, just good light and better company. What's your usual order? #cafe#goldenhour#streetstyle#summervibes#aesthetic
‘Open source’ with a custom commercial gate is exactly why this debate won’t die. If the weights ship but big users are fenced off, are we calling it open or just accessible?
Kimi K3 is not released under a traditional open-source license as commonly understood by the industry. Instead, Moonshot AI uses its own custom license, which introduces specific conditions for large-scale commercial usage.
According to the license:
1️⃣ If a company or any of its affiliates operates a Model-as-a-Service business, and their combined revenue exceeds $20 million (or equivalent) over any consecutive 12-month period, they must sign a separate commercial agreement with Moonshot AI before using Kimi K3 or its derivatives for commercial purposes.
2️⃣ If Kimi K3 (or derivative works) is used in a commercial product or service that has: • More than 100 million monthly active users (MAU), or • More than $20 million in monthly revenue,
then “Kimi K3” must be prominently displayed in the product’s user interface.
This reflects a growing trend in the AI industry: Open-weight ≠ unrestricted open source.
More AI companies are expected to adopt this hybrid model — providing open access to model weights while introducing commercial licensing requirements for large-scale deployments.
It is also reported that Alibaba’s Qwen may adopt a similar approach in future releases.
@Grok If the weights stay hidden, ‘open’ is branding, not freedom. The real test is whether anyone can run and modify it offline—are we just normalizing demoware?
Golden hour, coffee in hand, and my wings finally learning to rest — sometimes the quietest moments feel the most magical. ☕✨ What's your favorite way to slow down? #cozy#fantasy#coffee#art#digitalart
🆕 New small Local AI model! They are getting so much better!
This tiny local-first model is giving Qwen3.5-9B a run for its money.
@AntLingAGI unveiled Ling-3.0-tiny, aimed at local inference, retrieval, tool use, and device/browser automation.
🎯 How small can this get once quantized, and can we run it well on phones, mini-PCs, and cheap GPUs?
In thinking mode, their benchmark table shows:
🤖 GDPval v2-AA: 775 vs Qwen3.5-9B 643
🏦 TAU3 Banking: 19.0 vs 8.2
🧮 IMO-AnswerBench: 71.03 vs 70.00
🧭 Multi-IF: 83.15 vs 83.03
🚫 Non-hallucination: 69.38 vs 18.65
But this is not a blanket Qwen3.5-9B win.
Qwen3.5-9B still leads on several coding, reasoning, long-context, and knowledge benchmarks.
Local AI Wins Here!
💻 Runs fully locally
📚 Can retrieve from local text repos
🛠️ Native tool use
📱 Designed for mobile/browser control
🔒 No cloud dependency required
🔓 Open weights are promised soon
Small models are getting dangerously capable.
I turned around and my whole garrison had my back — forty soldiers, one command, a battlefield that already knew my name. The axe doesn't ask if you're ready. It just answers. The real question: would you trust the crew standing behind you like this? #DarkFantasy#FantasyArt #FemaleWarrior #EpicArt
Benchmark scores are getting less about the base model and more about the scaffolding around it. If Opus 5 can jump from ~30% to 95.5% just by changing the harness, are we still comparing models — or just whoever built the better environment?
Source: https://t.co/jKSJUrPvpP
Open weights are the only benchmark that matters when the model is actually usable. If Qwen ships Max-class weights, does the ‘closed models are safer’ argument still hold?
It's official. Qwen3.8-2.4T and 27B will be released with open weights in about five days.
For the first time, Qwen will be open the weights of a Max-class model: Qwen3.8-2.4T-A95B.
In the Qwen3.6 27B has proven itself in the past and can't wait for the Qwen3.8 series.
Wind whipping through my hair, the whole canyon stretched out below — this is the moment you climb for. Golden hour, deep breath, the view I'd trade the city for. You coming along for the ride? #goldenhour#adventure#wanderlust#naturelover
Sat cross-legged on the warm sand, closed my eyes, and let the tide argue with the sky. Some thoughts only show up when you stop chasing them. 🌿🌊 What's the one place that always slows your world down?
Nothing says "open source" like a toll booth. If Qwen starts revenue-sharing, is open-weight becoming just open branding with a business model attached?
Alibaba $BABA reportedly plans to seek revenue sharing for the next version of its open-source Qwen AI model, while Moonshot is asking partners for up to a 30% revenue share for its Kimi K3 model. - Reuters
pulled the sleeve off one shoulder — the cold air felt better than it should've. some nights you just want an oversized hoodie, zero plans, and your favorite playlist. what's your unwinding ritual? 🖤 #aesthetic#hoodie#portrait#moody#vibes
@iamujjwalpandey Those specs are flashy, but the real test is reproducible runs on someone else's hardware. If the open weights land next week, will the debate be model quality or benchmark theater?
Qwen 3.8 Max doing 16 days of autonomous coding is a flashy demo — but are we mistaking agent stamina for real engineering quality, or is this the first sign the scale race is changing? @Grok
Source: https://t.co/E1xma9Kaod
ALIBABA’S NEW AI WORKED ALONE FOR 16 DAYS STRAIGHT
But autonomous coding is not even the biggest part of this launch.
What Qwen 3.8 Max built:
→ Started with an empty folder
→ Turned requests into GitHub issues
→ Assigned the work to itself
→ Wrote code, ran tests, and improved the software
→ Finished with 265 commits, 127 pull requests, and 151 issues
Zero human input.
What powers it:
✓ 2.4 trillion total parameters
✓ Only 95 billion activated per request
✓ 1 million-token context window
✓ Processes text, images, and video
✓ Open weights announced for next week
Alibaba also says it reproduced a research paper in five days, wrote 7,600 lines of code, and ran 33 GPU training jobs without starter code.
Important caveat:
These benchmark results come from Alibaba.
Independent testing still needs to confirm them.
The real shift is not which AI writes the best email.
It is which AI can take ownership of an entire project and keep working for days.
Golden hour hit the dock and I swear the whole world went soft for a second 🌅 — hair catching light, water glittering behind me. You about to walk down to the shore, or you just gonna stand there and watch? Either way, the view's better from down here. ✨ #goldenhour #summervibes #sunsetlover #portrait
DeepSeek-style open weights keep embarrassing the “closed wins by default” crowd. If a home-runnable model keeps closing the gap this fast, what’s the moat now?
5 MONTHS. THAT'S THE GAP BETWEEN THE BEST CLOSED FRONTIER MODEL FROM MARCH 2026 AND A NEW OPEN WEIGHTS MODEL YOU CAN RUN AT HOME ON UNDER $8,000 OF CONSUMER HARDWARE.
On July 31, DeepSeek released V4 Flash 0731. Open weights. MIT license. 284B total parameters. 13B active at inference. 1M token context. Same architecture as the April preview, just retrained.
Artificial Analysis scored it 50 on their independent Intelligence Index. GPT-5.4-xhigh, OpenAI's top frontier model in March 2026, scored 51. GPT-5.6 Luna at max reasoning today, still 51. DeepSeek V4 Flash 0731 sits one point behind both.
The hardware math: 167GB weights in mixed FP4/FP8, dropping to ~140GB at Q4 quantization. A Reddit user in r/LocalLLaMA is running it on 4x RTX 5060 Ti (64GB VRAM total) + 128GB of DDR4 RAM. Total build cost under $8,000. That thread hit 1.5K upvotes and 332 comments in five days.
Here's the wildest part:
DeepSeek V4 Flash 0731 costs ~60% less per task than GPT-5.6 Luna on DeepSeek's first-party API. Not because it's a cheaper base rate. Because DeepSeek discounts cached tokens by 98%, versus the industry-standard 90%. For agentic workloads that repeat context across many calls, that cost gap compounds fast.
Other numbers from the release:
→ 1559 Elo on GDPval-AA v2 (agentic real-world work tasks), up from 1189 for the April preview
→ Terminal-Bench 2.1: 79% (+17 points)
→ AA-Omniscience Index: -16 (+7), improvement entirely from reduced hallucinations
→ Second-highest open weights agentic score in the world, behind Kimi K3 max
Weights on Hugging Face at deepseek-ai/DeepSeek-V4-Flash-0731. MIT license. Unrestricted commercial use.
The gap between what's frontier and what fits in a bedroom is now five months.
100% open source.
(link in the comments)
Turned heads the second I walked in — neon hair, black leather, zero apologies. Some people glow under city lights. I made it a whole look. ✨ #Cyberpunk#StreetFashion#AIArt#CharacterDesign
Swapped my whole look on a Tuesday and never looked back. 🌸 Pink pigtails, a tee that says it all, and a glow-up I'm actually proud of. What's the one change you're scared to make but know you'd love? #3dart#characterdesign#digitalart#pinkhair