HOW DID WE MISS THAT
🚨"im-a-good-gpt2-chatbot" MUCH Better Than gpt4o and On Par With SONNET 3.5?! 🚨
Context: I replicated the entire reasoning benchmark of https://t.co/va0igVxeWr (by @ylecun among others) but for the "im-a-good-gpt2-chatbot"
new livebench - next to old
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!
5.6 sol growth is insane.
the inference team has done heroic work to be able to support demand.
we are going to move mountains to continue to scale, but it is possible there are some hiccups soon.
@thsottiaux I was already running a long task while you started those changes and it's unusually slow (been running for 1.5hours). Could it be related to the changes you made to Sol?
Morning. The last 48 hours of Codex and ChatGPT Work have been intense! Three important updates:
- Temporarily removing the 5 hour usage limit restriction for all Plus, Business and Pro plans
- Rolling out changes that will make GPT 5.6 Sol more efficient across the board and that will be reflected in less usage being used so that it can take you further. Exact impact to be quantified and shared
- We hit 6M active users, and are landing a usage reset in the next hour
Go do things
@threepointone a lot of people follow this logic to a standardization that just doesn’t work. models are all different with different personalities. it’s jarring to have them auto switched underneath you
@BiotechMongoose dude what are you saying it says phd even in computational biology plus python and ML
that's MANY people compared to the available job positions out there for such roles