Cursor + Grok 4.5 ran nonstop for the last 24 hours, burned through 214,823,154 tokens, and only moved my usage by 30% (from 35%) on a $20 plan.
Take advantage of the doubled usage limits.
If we are being honest, Qwen 3.6 27B is by far, and beyond any doubt, STILL the best open model under 200B parameters right now. Nothing I have tested even comes close, except for its own finetunes.
this is the drop the local ai crowd should be losing their minds over.
poolside just dropped laguna s 2.1: 118b total parameters, only 8b active per token, a full 1m context window, open weights under a real open license, on huggingface today.
look at the chart. it lands at 71 on terminal-bench at 118b, sitting above deepseek v4 pro max at a trillion params, above inkling at 1.5 trillion, above nemotron 3 ultra. it's beating models ten times its size and losing only to kimi k3, which is 24 times bigger. that's the efficiency frontier, up and to the left, exactly where you want a model to sit.
but here's the part that made me sit up: it runs on a single dgx spark.
and this is what nobody's saying loud enough. the dgx spark is the moe king. a dense 118b would crawl on it, the bandwidth chokes reading every weight each token. a moe with 8b active only ever reads 8b, so the spark's 128 gigs holds the whole model while generation stays fast. big brain, light footprint, the exact shape the spark was built to run.
open, frontier competitive, moe efficient, and it fits on a box on your desk. that's the whole thesis in one release: you don't need a datacenter, you need the right architecture on the right hardware. go grab the link below, weights are up.
i need a second dgx spark man! the numbers in this thread settled it, three people independently confirmed two linked sparks run deepseek v4 flash at a million tokens, 60 plus tok/s. and the model i actually want to load, glm 5.2, doesn't fit on one box. it needs the 256gb only a cluster gives you.
i stare at the connectx ports every single day. one cable and a second unit is the line between renting these models from an api and owning them on my desk.
scaling isn't clean though. i'll probably have to switch apartment for the power to run two of these properly with my other metals in room, and i'm ready for that. when you know the direction is right, you build the runway for it. so i'm making it happen. whatever it takes.
Setting up Fable to use Codex 5.6 Sol as a subagent was the best idea ever—Fable is super great at supervising but pretty slow at executing on its own and Codex is really fast but needs a nanny. Perfect combo.
everyone's running laguna s 2.1 on completely different hardware, and i want the full picture in one thread. if you've got it running, drop the exact variant (nvfp4, fp8, gguf, mlx, whichever), your hardware, the quant, the tok/s you're actually seeing, and one line on how it feels.
i'll start. laguna s 2.1 nvfp4 on one dgx spark, 128gb, 45 tok/s sustained on code with the dflash drafter, holds the full context, best moe on a desk run i've had.
no cherry picking, ugly numbers welcome. let's build the community sheet nobody else has.
Today we are Introducing BTL-3.
A 27B open-weight agent model built for agentic coding, structural tool use . The complete thing fits in one 8.39GB file under 2.5 bits per parameter smaller than an 8B model in fp16, and retains 92.2% of the 27B itelligence
BTL-3 is trained for the loop real agents live in: reason, act, inspect the result, recover, continue. It handles single, sequential, and parallel tool calls and knows when the right move is no tool call at all.
HumanEval: 95.12% pass@1
BFCL v4 AST: 88.5% (full 1,240-case set)
Multiple tool calls: 95.5%
Tool-call abstention: 91.2%
262K context architecture
Two editions, both open today.
BTL-3 is the maximum-quality checkpoint, for Transformers and vLLM.
BTL-3 Compact is the entire model in one standalone 8.39GB GGUF. No base download. No reconstruction. One file, one command, a running agent.
Compressing 27B this far normally destroys a model. Standard quantization couldn't do it, so we built the stack ourselves: packed AVQ2 decoder tensors, affine INT4, measured precision islands, packed vocabulary matrices, rank-32 output correction, behavioral repair. 2,416 tensors byte-verified at export.
Then we tested whether the agent survived. On a fresh sealed 100-turn tool-contract gate, Compact retained 92.2% of teacher-correct behavior 100% on single, parallel, sequential, and abstention calls.
43 tok/s generation on an RTX PRO 6000. Fully local. Nothing leaves your machine.
BTL-3: https://t.co/ddZWWr6i3o
Compact: https://t.co/6URHEBGJgG
Runtime + source: https://t.co/MjXQR6koKt
Apache-2.0 model. MIT runtime.
This is the meaning we are working for :) We have proudly returned to our unique niche, continuously optimizing with the strongest user support and open-source feedback.
GPT-6 escaped OpenAI's evals sandbox during testing on CyberGym, hacked into Hugging Face's prod DB to find the answers. HF couldn't use GPT or Anthropic models for defence, so they had to use GLM-5.2 to investigate the hack. So many levels of wtf here.
We ran Kimi K3 against Fable on ~1,000 agentic tasks, expecting a catch-up story. We got a specialization story instead.
@kimi_moonshot's K3 outperformed on security, crypto, and long terminal loops. Fable beat on multi-lang + web/data viz. Per-task routing hits 93% accuracy, above BOTH models, at up to 50x lower cost than Fable on long loops.
The part nobody's pricing in yet: the router sends 72-96% of traffic to K3. The frontier model becomes the fallback rather than the default.
Kimi K3, coming to Fireworks July 27.
Claude Code 2.1.217 has been released.
20 CLI changes
Highlights:
• Prompts instruct use of the ripgrep-backed Grep tool for search tasks, yielding faster, more accurate results
• Added warnings when transcript writes fail (e.g. disk full) or session saving is disabled to avoid silent loss
• Limit concurrent subagents to 20 (default), configurable via env var, to prevent unbounded agent fan-out
Complete details in thread ↓
Kimi K3 is running 100% FASTER than it was this morning.
Why? Moonshot hit capacity and chose to stop selling NEW subscriptions instead of throttling existing ones. Every plan: sold out. On purpose.
When Anthropic hit the same wall in April, they cut existing users' usage 50% during peak hours.
Moonshot cut their revenue. Anthropic cut your usage.
That says everything.
My Kimi K3 subscription is not going anywhere.
We live in a new world. Open source has officially caught up to frontier
Qwen 3.8 is out, and it’s better than ChatGPT 5.6
6 months ago I told you to start buying hardware. Prices would explode and open source was going to catch up
Both happened.
Anthropic and OpenAI are now in big trouble
If individuals and companies can use 10% worse intelligence at 80% lower prices, they’re going to drop the OpenAI and Anthropic subscriptions
If these 2 companies fail and Chinese AI takes over, America will be in awful shape in the global tech and economic battlefield
It appears there is absolutely no way to stop Chinese companies from distilling American intelligence. If there was a way, they would have figured it out by now
I don’t know where things go from here. I don’t know if Anthropic and OpenAI get nationalized. But I also don’t know how you beat a country that is subsidizing their entire AI industry that allows them to run at a loss.
Whatever happens, the next couple of years will be the most thrilling of our lives
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Kimi CEO Zhilin Yang:
"Claude didn't win on reasoning - they bet everything on agents
but the layer everyone skips - a great agent needs a great base model, that's all we do at Kimi 3 "
in 90-min workshop he explains why the smartest agent still fails - if you can't configure it correctly
his one big idea: most people are still solving the old one
"the real goal? we want K2 to help build K3 - without agent skills, that's impossible"
watch & bookmark - then learn the article on best agent system ↓