Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
10x more tokens per megawatt.
CoreWeave has the first measured performance of NVIDIA Vera Rubin NVL72, showing 10x improvement in tokens per second per megawatt on DeepSeek-R1 compared to Blackwell.
Today we are Introducing BTL-3.
A 27B open-weight agent model built for agentic coding, structural tool use . The complete thing fits in one 8.39GB file under 2.5 bits per parameter smaller than an 8B model in fp16, and retains 92.2% of the 27B itelligence
BTL-3 is trained for the loop real agents live in: reason, act, inspect the result, recover, continue. It handles single, sequential, and parallel tool calls and knows when the right move is no tool call at all.
HumanEval: 95.12% pass@1
BFCL v4 AST: 88.5% (full 1,240-case set)
Multiple tool calls: 95.5%
Tool-call abstention: 91.2%
262K context architecture
Two editions, both open today.
BTL-3 is the maximum-quality checkpoint, for Transformers and vLLM.
BTL-3 Compact is the entire model in one standalone 8.39GB GGUF. No base download. No reconstruction. One file, one command, a running agent.
Compressing 27B this far normally destroys a model. Standard quantization couldn't do it, so we built the stack ourselves: packed AVQ2 decoder tensors, affine INT4, measured precision islands, packed vocabulary matrices, rank-32 output correction, behavioral repair. 2,416 tensors byte-verified at export.
Then we tested whether the agent survived. On a fresh sealed 100-turn tool-contract gate, Compact retained 92.2% of teacher-correct behavior 100% on single, parallel, sequential, and abstention calls.
43 tok/s generation on an RTX PRO 6000. Fully local. Nothing leaves your machine.
BTL-3: https://t.co/ddZWWr6i3o
Compact: https://t.co/6URHEBGJgG
Runtime + source: https://t.co/MjXQR6koKt
Apache-2.0 model. MIT runtime.
We released Nanbeige4.2-3B, its Looped Transformer increases model capacity without adding parameters, delivering a capable 3B agent.
Nanbeige4.5 is training with LoopSplit, mHC+depth attention & concatenated n-gram embeddings,already in the modeling code.
https://t.co/SnxGOyi3xj
Today we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
https://t.co/xxGeAgo35R
Today we are releasing Laguna S 2.1.
At 118B total parameters, with 8B active per token, it does the work of models several times its size on agentic coding. It is remarkably persistent across long-horizon tasks. And it is small enough to run on a single NVIDIA DGX Spark.
It is far more capable than anything we have created before, and I think it redefines what a model in its weight class can do.
Laguna S 2.1 is an important model for Poolside. What it represents is even more important.
If, five years ago, I had read a book that said that by 2030 everything economically valuable, scientifically interesting, and personally meaningful would be built on intelligence contracted from three or four companies, I would have called it dystopian science fiction.
We are at a fork in the road of what kind of world we can have.
I believe intelligence should and will become a commodity. The question is whether that intelligence comes from three companies, or from many people who can build it, own it, and shape it.
The open ecosystem will not win by being the best in its own category. No one cares who is king of the open-source kingdom. People want the best intelligence for the task they are trying to do, with the right balance of quality, speed, cost, and control.
If we want a different future, open models have to be on par with, or better than, their closed equivalents.
Laguna S 2.1 is a meaningful step in that direction: capable enough to compete far above its weight class, efficient enough to run on hardware you can own, and open-weight so anyone can build on it.
Open-weighting our models is the contribution we can make today toward a world where intelligence can be built and owned by many. And we will keep doing it.
I am very proud of this team’s work. A big shout out to everyone at Poolside who made this possible, from infrastructure and data to architecture, pretraining, post-training, evaluations, and inference.
Laguna S 2.1 is available today under the OpenMDW-1.1 license, with weights on Hugging Face and access through OpenRouter and our API.
We are building toward a future where the most capable intelligence in the world can be owned and shaped by anyone. Laguna S 2.1 is one step. We are going to keep building until that future exists.
https://t.co/eTV7iNATrB
Introducing Gemini 3.5 Flash Lite and Gemini 3.6 Flash🔥
We incorporated all the feedback from the last couple of weeks and... 🥁
- More token efficiency, fewer unneeded tool calls, less eagerness
- Reduced Gemini 3.6 Flash pricing
- Quality improved
- And a lot more!
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model
Key results:
➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation.
➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality.
➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params).
➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04)
➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores.
➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities
Other model details:
Context window: 1M
Size: 2.8T total parameters
Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
Modality: Native multimodal input supports text and images, and the model remains text-only for output.
Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions!
Here is a breakdown of what’s being fixed and updated in this release: 🧵👇
Today, we’re announcing Bonsai 27B: the first 27B-class model to run on a phone.
Bonsai 27B is the new multimodal flagship of the Bonsai family. Based on Qwen3.6 27B, it brings a new capability tier to local AI: multi-step reasoning, structured tool use, long-context workflows, and coherent agentic loops.
Until now, models in this class have been impractical to deploy locally. A 27B model occupies roughly 54 GB in 16-bit precision, and even a strong 4-bit build is around 18GB - too large for a phone and for most laptops.
Bonsai 27B changes that.
It comes in two variants:
• Ternary Bonsai 27B: 5.9 GB, 1.71 effective bits per weight, optimized for laptop-class quality.
• 1-bit Bonsai 27B: 3.9 GB, 1.125 effective bits per weight, optimized for phone-class footprint.
Everything is open-sourced today under the Apache 2.0 license.
GPT-5.6 Sol comes close second to Claude Fable 5 in the Artificial Analysis Intelligence Index at one third of the cost, and leads the Artificial Analysis Coding Agent Index in OpenAI’s Codex harness
We supported @OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and Luna. GPT-5.6 Sol (max) scores 1 point below Claude Fable 5 (max) in the Artificial Analysis Intelligence Index at 59 points, at approximately one third of the cost. GPT-5.6 Terra (max) and Luna (max) score 55 and 51 respectively in the Intelligence Index, at ~50% and ~80% lower Cost per Task than Sol.
GPT-5.6 Sol (max) leads the Artificial Analysis Coding Agent Index at 80 points.
Congratulations @OpenAI and @sama on the launch!
Key takeaways:
➤ One third of the cost of Claude Fable 5: On max reasoning effort, GPT-5.6 Sol costs $1.04 per task in the Artificial Analysis Intelligence Index - offering a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. Reasoning levels across GPT-5.6 Sol and Luna offer a range of options at the Pareto frontier of Intelligence vs Cost per Task. For example, GPT-5.6 Luna (max) matches or exceeds the intelligence of GLM-5.2 (max) and Gemini 3.5 Flash at a lower cost. GPT-5.6 Terra (max) and Luna (max) cost $0.55 and $0.21 per Intelligence Index task, ~50% and ~80% less than Sol. Across reasoning efforts, each new GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra. This means that for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or as intelligent at lower cost.
➤ Leading in all Coding Agent evaluations: The new Artificial Analysis Coding Agent Index pairs models with agentic harnesses and features three frontier coding evaluations - DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. GPT-5.6 Sol (max) in Codex scores 80 in the Index, leading in all three evaluations (tying Grok 4.5 in Grok Build for SWE-Atlas-QnA). In addition to scoring higher, its per task cost is ~40% and ~10% cheaper than Claude Fable 5 (max) and Opus 4.8 (max) respectively in Claude Code. GPT-5.6 Terra (max) and Luna (max) score 77 and 75 in the Coding Agent Index respectively, with ~60% and ~80% per-task cost reductions compared to Sol.
➤ Highest Presentation Elo in AA-Briefcase: GPT-5.6 Sol (max) ranks second only to Claude Fable 5 (max) in AA-Briefcase, and has the highest Presentation Elo of any model. AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. GPT-5.6 Sol (max) has the highest recorded Presentation Elo - its outputs across various file types, including PowerPoint and Excel, are the most visually attractive of any model. Fable 5 (max) still leads AA-Briefcase, largely due to its Rubric Score of 56% vs 42% for GPT-5.6 Sol (max). Fable 5 (max) also scores 1764 in Analytical Quality Elo vs GPT-5.6 Sol (max) at 1592.
➤ First OpenAI models with cache-write pricing: GPT-5.6 introduces cache-write pricing for the first time at OpenAI. Sol, Terra, and Luna are priced at $5/$30, $2.5/$15, and $1/$6 respectively per million input/output tokens. OpenAI has retained its previous discount of 90% for cache reads, but joins Anthropic in introducing a cost premium for cache writes, at 1.25x the price of input tokens. Cache writes occur when input tokens are committed to memory. Charging for a cache write more accurately reflects the model’s cost to serve, as cached tokens occupy memory whether or not they are reused. Also in line with Anthropic's models, GPT-5.6 introduces a max reasoning effort level.
➤ Low token use: GPT-5.6 Sol (max) uses fewer output tokens than most models of comparable intelligence, and defines a new Pareto frontier of Intelligence vs Output Tokens per Task. GPT-5.6 Sol (max) offers a slight improvement in token efficiency with 15k tokens per Intelligence Index task, vs GPT-5.5 at 16k. Notably, it uses fewer tokens and is more intelligent than Claude Opus 4.8 (max), GLM-5.2 (max), and Gemini 3.5 Flash (high).
Today we’re releasing the weights for Laguna M.1,
our most capable model to date, with a 256K context length.
Both base and post-trained checkpoints are now available on Hugging Face under Apache 2.0.
Introducing GLM-5.2: Frontier Intelligence, Open Weights
- Significant improvements in coding and agentic tasks
- Strong long-horizon capabilities with a 1M context window
- Two levels of reasoning effort: GLM-5.2 (max) pushes the limits, while GLM-5.2 (high) strikes a strong balance between performance and token efficiency
- MIT-licensed open weights
- Same API pricing as GLM-5.1
Tech Blog: https://t.co/LAsxUdN0JZ
Weights: https://t.co/g0A1C4UWx4
API: https://t.co/Kc3E22cbN7
Coding Plan: https://t.co/Nk8Y98HNhU
Chat: https://t.co/WCqWT0qCQb
Meet DiffusionGemma!
An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license.
Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
Second big release from us today: Nemotron-3.5-ASR-Streaming!
🌎40 languages
⚡️80ms - 1s controllable latency
🔥240 - 2400 concurrent streams on 1xH100
🧱FastConformer Cache-Aware RNN-T architecture
https://t.co/lxmcAnKeOl