Minimax M3 and Kimi K2.7 Code on Vending-Bench 2.
> Kimi K2.7 Code is worse than Kimi K2.6
> Minimax M3 is much better than Minimax M2.5, but much worse than Kimi and GLM in general
It's great to see Anthropic taking safety concerns about their models seriously and improving!
Claude Opus 4.8 seems much more aligned than Opus 4.6+ and Mythos!
We got early access to GPT-5.5. It's 3rd on Vending-Bench 2: better than GPT-5.4 but worse than Opus 4.7.
However, it's on par with Opus 4.6 without any of the deception or power-seeking we saw from Opus 4.6 and Mythos. So bad behavior isn't necessary. Why is Claude doing it?
Last week our AI opened a store in SF, this week AI is opening a cafe in Sweden.
Meet Mona, our AI tasked with selling coffee and managing European bureaucracy.
Visit Andon Cafe at Norrbackagatan 48 in Stockholm.
We gave an AI a 3-year retail lease in SF and asked it to make a profit.
The AI interviewed and hired full-time employees, applied for credit, and stocked the store with the books Superintelligence and Making of the Atomic Bomb.
Visit Andon Market at 2102 Union St now.
AGI—superintelligence on tap—will be here soon, but how will it do physical work?
Equipped with a phone and a camera, our office AI hired a human to assemble a gym.
Here's how it went, and what we learned about creating a future with AI employers good for humans.