Kimi k3 is an incredible model. It is not an incredible value. In most tasks, it comes out to roughly the same cost as GPT-5.6 Sol.
K3 is half the price of 5.6 Sol per token. GPT-5.6 uses half as many tokens. Price evens out.
GPT-5.6 is 2x faster TPS, so it gets work done ~4x faster than K3 at roughly the same price.
I still love K3 and will be using it for a TON of stuff. I'm just tired of people pretending it's way cheaper when it's not.
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
https://t.co/wDdSEjBe4F
Introducing autoresearch for arXiv papers
Change 'arxiv' to 'autoarxiv' in any paper URL
An agent deploys to resolve setup issues on the codebase, run a minimal reproduction, and estimate full replication cost. Read more below
@asanwal That feels like the main reason to stick with it. Although it feels too nice, as instead of saying No sometimes, it hallucinates and gives incorrect solutions
Yesterday, AWS shook the data lake world by releasing two new S3 features that will forever cement its place there. 👑
Every data engineer must become familiar with them.
A short thread on these game-changers 🧵 (2 minute read)