Tested GLM-5.2 FP8 and NVFP4 on 8× PRO 6000. NVFP4 is the version that makes local deployment truly usable.
Compared with the fp8 stack, nvfp4 gave us:
• ~1.03M-token KV capacity vs ~200K (fp8)
• successful ~960K-context single-request run
• up to ~281 tok/s aggregate throughput at 16 concurrent 2K-context requests
Tried running DeepSeek-V4-Pro on a single Pro 6000 with fastllm
Results here👇
For local inference, that’s a fairly solid sign.
#LLM#LocalAI#Inference#DeepSeek
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report: https://t.co/drlDrxkYtp
🤗 Open Weights: https://t.co/T13Y8i7SDM
1/n
ran MiniMax-M2.7 on 4× Pro6000 so you don't have to🧪
• Prefill holds strong at ~17K tokens/s from c2 through c32
• Decode scales linearly the entire way: 104 → 633 tokens/s with no sign of hitting a wall
full charts below 👇
#MiniMax#M27#LLMBenchmark#AIInfra
We're delighted to announce that MiniMax M2.7 is now officially open source.
With SOTA performance in SWE-Pro (56.22%) and Terminal Bench 2 (57.0%).
You can find it on Hugging Face now. Enjoy!🤗
huggingface:https://t.co/ApWrahIl3o
Blog: https://t.co/gAxeFsNdW4
MiniMax API: https://t.co/1dgbMx0Q7K
The new Qwen 3.5 by @Alibaba_Qwen running on-device on iPhone 17 Pro.
Qwen 3.5 beats models 4 times its size, has strong visual understanding, and can toggle reasoning on or off.
The 2B 6-bit model here is running with MLX optimized for Apple Silicon.
We’re seeing a massive shift toward Private Cloud and Local LLMs.
Enterprises simply cannot risk their IP becoming public training material🙃The future of AI is local.
30 Concurrent Tasks. 30 Enterprise Scenarios. One Beast of a Machine. 🦾
Check out the HY NV4-6000 in action, powering @MiniMax_AI -M2.5 across 4x Pro 6000 GPUs🔥
From code to legal, we’re redefining what enterprise-ready AI infrastructure looks like.
Real-time dashboard, real-time power.
#AI #Innovation #NVIDIA #Pro6000 #MiniMax #HighPerformanceComputing
3 updates dropped today: Claude Code remote control, Claude Cowork plugins, Notion custom agents..and they're all saying the same thing: AI is becoming a headcount you manage.
Each product is now asking you to define a role, give it a scope, and let it run.
We're moving from prompting to delegating. You don't sit with the AI anymore. You check in on it. You approve things from your phone. You review its output like you'd review a junior's work.
Prompting is becoming delegating. The next interface will be a roster of agents with job descriptions, running async, embedded in your existing tools.
The org chart of a 50-person company might quietly have 200 agents on it within a year. Most standups will be about what the agents did, not what the humans did🙌
OpenClaw ships native Kimi video provider with cache-ttl fully wired up. Practically means you can throw a clip at your local agent and it actually understands what's happening frame by frame. The real play here is what comes next. Once vision+video stabilizes over the next few months, you get agents that watch your screen, watch your camera feed, and act on what they see.
Kimi is the cheapest high-quality multimodal for Chinese contexts by a wide margin, and OpenClaw just made it a first-class citizen. Long-video streaming and Lobster workflow integration are the obvious next steps.
We're living in the golden age of open-source AI and I don't think people realize it yet🙌
Kimi K2.5, GLM 5, DeepSeek...Chinese labs are shipping absolute monsters back to back. Coding, reasoning, long-context… the gap with closed-source is vanishing in real time.
The most important part, you can run them on your own hardware. No surprises.
#OpenSource #OpenSourceAI #KimiK2 #DeepSeek #GLM5 #AI