During the FIFA World Cup, Kalshi became one of the world's most visible brands.
• 29B+ impressions - more than every major brand
• Best World Cup Commercial (Adweek/EDO)
• 5.8M+ downloads
• $27B traded
We did it with a team of 15. There's no secret sauce🧵
ok just spent a morning with Kimi K3 as my firstmate, here's my real experience
1. it's very, very slow
potentially due to the fixed max reasoning. you should expect the experience of something slightly slower than fable
2. its claimed cost efficiency is not manifesting in real economics
i bought the $40 plan, and a few prompts later it's already eaten 1/3 of my 5-hr limit - it was in a single session and my context window was only 200k long at that time
i don't care what the benchmark numbers say, and what the face value API pricing is, in reality Kimi K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan - i observe no efficiency benefit
3. its instruction following capability is weaker than other frontier models
firstmate stretches frontier models' reasoning capability and is a really good test that can quickly reveal how good a model is at following instructions
the pure "intelligence" of K3 does hold up - it understands my intent very well, and can diagnose problems, delegate tasks all fine
but i very quickly noticed many instructions in firstmate's system prompt not strictly followed by Kimi K3. these were never a problem with gpt 5.5, 5.6, opus, fable and grok 4.5
so all in all, i'm now very skeptical of the claimed performance and going to keep my eyes wide open on its true capability