We had a team of agents rebuild SQLite from its 835-page manual.
It created a replica in Rust which passed 100% of a held-out test suite.
Interestingly, cost varied 15x depending on which model mix we used.
extremely unprofessional. if kimi wants to make it as a frontier lab, they need to act like one: perhaps silently route people to worse models, and maybe write a blog post about the collapse of humanity
I tried Grok Build for the first time today.
I've been using claude code for over a year now, which is crazy to think about.
I'm not one of those people that's complaining about anthropic stringing us along, or any of the other things happening. I just legitimately love trying different options/harnesses.
It's hard for me to ascertain whether any harness is "better", but one thing I will say - speed really matters, and grok is FAST.
I also think that the TUI is much better than CC, codex, etc. Not that this is as important, but it's a nice-to-have.
Overall, pretty impressed. Will continue to test available options as time moves on, but it seems like grok is a strong contender in the ever-growing list of coding harnesses
Meet kbd-1.0-codex-micro, built with @work_louder.
Map the buttons and joystick to your workflow, and keep your pinned chats in view.
Get yours before stock returns 410.
I was clearly wrong about Anthropic. They are obviously currently the leader in AI. No company has released a model as good as Mythos/Fable and they will undoubtedly have Mythos 2 ready soon.
And I would never cut them off in a way that hurt them badly, even as a competitor. That’s not my style.
Tesla open sourced its patents and we made the Supercharger network available to all competitors, even though we could have made it a walled garden.
SpaceX launches competing satellite systems with no increase in price or use of unfair terms.
Even my worst enemies can attack me on this platform.
…