@jhong@rehan_shei maybe it was just the time it was transitioning between the different platforms, but it was stuck on rick & morty for quite a while (if you saw the multiple scenes of rick shooting robots)
Grok 4.6 is now #2 on EEbench.
It breaks Anthropic’s hold on the top of our electrical-engineering benchmark and makes @xAI only the second lab to cross the 50% mark.
Hardware agents are getting real, fast.
@zachdive Glad you're enjoying it!! Let me know where you're finding issues and i'll fix it over the next few weeks.
Just getting started on CAD capabilities!
One of the coolest parts of the AI Evals for Engineers and PMs course was seeing @sh_reya and @HamelHusain build a custom eval workflow tool using Claude Code. Highly recommend it for PMs who want to truly make their LLM-powered app work https://t.co/zKg55iqawq