ER 2 looking solid. The success detection and instrument reading numbers went up quite a bit. Excited to see people start building real robot workflows with this.
We've added Kimi K3 to Perplexity and Perplexity Computer for Pro and Max subscribers.
Kimi K3 in Perplexity is hosted exclusively on U.S.-based servers.
Antigravity is being used for research eval analysis.
Instead of manually writing scripts and parsing tables, researchers can now analyze evals in minutes using natural language enabled by skills. Antigravity performs the quantitative analysis, proposes hypotheses, and spins up dynamic subagents in parallel to isolate the failure modes.
Kimi K3 is now live in LM Studio Bionic!
A massive, open frontier model: 2.8T MoE with 1M context window. Kimi K3 in Bionic is hosted on US-based servers, with ZDR enabled by default.
The last two weeks in open weights have been absurd. Seven drops worth your attention:
Kimi K3 (2.8T total, 104B active)
Moonshot shipped the largest open-weight model ever released. Native vision, 1M context, 16 of 896 experts active per token, and it sits behind only the top closed frontier models on Artificial Analysis. Weights landed July 26, a day early. Read the license before you build on it though, it is a custom one that gates commercial inference above $20M revenue, not the Modified MIT everyone assumed.
Laguna S 2.1 (118B MoE, 8B active)
Poolside’s agentic coding model. 70.2% Terminal-Bench 2.1, 78.5% SWE-Bench Multilingual, and it beats DeepSeek-V4-Pro-Max on DeepSWE with roughly one sixth the active params. 1M context, runs on a single DGX Spark, OpenMDW-1.1. Nine weeks from start of training to launch. They also published full eval trajectories, which almost nobody does.
Solar-Open2-250B (250B, 15B active)
Upstage went hybrid attention: three linear layers for every softmax layer, 48 total, no positional encoding at all. That is how they get 1M context without the quadratic bill. Trained on ~12T tokens across English, Korean and Japanese, built for long horizon office and document work. Note the Solar license, derivatives must carry the “Solar” prefix.
Nanbeige4.2-3B
The one I would actually run locally. 3B non-embedding params with a Looped Transformer that feeds hidden states back through the same layers, so you buy capacity without buying parameters. Pretrained from scratch on 28T tokens. Their evals put it above Qwen3.5-9B and Gemma4-12B on agentic benchmarks. Apache 2.0.
Bonsai 27B / Ternary Bonsai 27B
PrismML pushed Qwen3.6-27B to 1-bit and ternary end to end, including embeddings and the LM head. Ternary is 5.9 GB at 94.6% of FP16, the 1-bit build is 3.9 GB and runs on an iPhone 17 Pro Max at about 11 tok/s. 262K context, Apache 2.0. One caveat the writeups skip: tool-calling degrades far worse than math at 1-bit, 80.0 down to 66.0. Fine for chat, think twice for agents.
BTL-3 (Qwen3.6-27B + rank-32 LoRA)
Bad Theory Labs post-trained the dense 27B for repo work, structured tool use, and failure recovery. The interesting artifact is BTL-3 Compact, the whole text model in one 8.39 GB file, smaller than an 8B in FP16, retaining 92.2% of the full model’s tool behaviors. Apache 2.0 adapter, base model separate.
Antares 1B (bonus, from Cisco)
Granite 4.0 1B backbone, SFT plus GRPO, tuned for terminal-based security work. There is a 0.3B variant too. Good example that task-specific post-training on a genuinely small model still pays.
And Alibaba has already teased Qwen 3.8 at 2.4T with open weights, which would have been unthinkable from them a year ago.
The pattern across all of this is not scale. It is that sparsity, looped layers, hybrid attention and extreme quantization all landed in the same month, and every one of them cuts what you need to run serious models on hardware you own.
Starting today, you can create ten videos *at no cost* in Gemini until 11:59pm PT on August 4th, 2026.
Just select "Create video" in the tools menu to create, edit, and remix videos to bring your ideas to life.
Share how you’re using Gemini Omni in the replies 👇
We signed the Open Weights letter because we believe the future of AI should be shaped by everyone, not controlled by a select few.
That belief has always been at the heart of Unsloth: everyone should be able to train and run models on their own local device.
Announcing Fugu-Ultra v1.1 🐡
We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work.
Today, we’re releasing Fugu-Ultra v1.1 → https://t.co/hhO6qTawgb
Upgraded to incorporate the latest frontier models, resulting in stronger performance across every benchmark shown, including gains of up to 7.9 points over v1.0, with particularly strong results on ProgramBench and Terminal Bench 2.1.
Fugu-Ultra v1.1 is more capable across coding, agentic tasks, and advanced reasoning, and available at the same price as Fugu-Ultra v1.0
The frontier keeps moving, and Fugu keeps getting better.