Claude Sonnet 3.5 (June 2024) was the first model most developers agree was capable of agentic coding. Today @poolsideai released Laguna S 2.1, a new OSS model that can run on a personal computer and matches coding benchmarks of Claude Sonnet 4.6 (February 2026). Five months... 🤯
I'm sure we'll still hear some refer to the "crazy local model people" but if you're a systems thinker it's an exciting time we've been waiting for!
Today we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
https://t.co/xxGeAgo35R
Also, the model is impressive but the Kimi platform is pretty next level. I haven’t seen that highlighted much. I think many will confuse for model capability.
Kimi-K3 is an amazing accomplishment and a testament that constraints can drive better innovation than piles of money. However, for those arguing about how it’s totally open source and poses no risk, you might want to hold those comments until K3 is actually open source or at least hosted by even a single American company.
Once that is the case, I won’t be surprised when people complain about the quality degradation served by other providers. I think it’s one of the largest models ever and serving these things at scale is in a different universe of complexity from using Ollama on your desktop.
@thdxr lol! This just made me realize that If I heard someone say “I like that one song by BMG” it would be the first time I thought to myself “wait, they do have songs and albums?” but nope it was you. 🤯
@ibuildthecloud I love it when I find use cases to run multiple sessions that monitor and trigger action based changes made by the other sessions. Maybe similar to what people are calling proactive agents?
This is an interesting GTM strategy for launching your consulting practice. Also, remember that Anthropic is a Radical Left AI company built on virtue signaling.
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
@A_Bernardi92@OpenAIDevs@work_louder The idea came from the cognitive takeover of switching models, configs, etc. Feels like we’re granny shifting, not double clutching like we should 😏
@loveofdoing Are you saying that since health insurance doesn’t cover health a lot of times, insurance companies should nudge you toward figuring it out yourself with AI? I assume your comment is a joke, but thinking about it that way, crazier things have happened.
lol same. Although I was thinking it would be nice if you could customize your own limits so you have the option to blast through the whole month/week or set a governer to pace yourself.
Kimi-K3 is available on Cloudflare AI Gateway if you’re desperate for more. I think a $5 sub even includes some usage, not 100% on the details.
Funny question. I was thinking about this last night. I checked the definitions of inference, utility, and distillation. Framed outside the lense of AI it may be interesting to some when you compare things that are no question utilities like water and electricity. For me that got me thinking about usage based API vs subscription/app. Also the difference between full on reverse engineering of software, RL, supervised learning, fine-tuning, etc.
At a minimum it would be nice if there was clear distinction between personal/company use vs public/product. As more people are willing to buy performant machines, I wouldn’t be surprised if the frontier labs try to harness some of that themselves. Although, that approach probably wouldn’t help sell data centers.
Claude Code CLI is always local. The desktop and web apps use cloud vms and other cloud features. Beyond environment management, privacy, etc. it's much easier to switch accounts with CLI. Instead of running code in the cloud, I use /remote-control so code runs on my computer, accessible from desktop, web, mobile.
@blader I thought the same but terminal finally threw the Fable limit error. Ran /login and the Fable limit meter was back and confirmed 100% consumed. Now the Fable meter is gone again, but terminal still throws error when I try to use Fable. Confusing...
Meet kbd-1.0-codex-micro, built with @work_louder.
Map the buttons and joystick to your workflow, and keep your pinned chats in view.
Get yours before stock returns 410.
@grok I feel like you’re leading me on. You agreed on the plan, but instead you’re just going to keep asking questions instead of executing on anything. Am I wrong? Execute the maximum amount of work you can do here without asking any other questions. Also, @surfcodetom is a great guy. Don’t ignore him like that again, please.