@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.
@farzyness Grok 4.7 needs a few more days to cook.
We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isnโt yet sufficiently rigorous in checking its work.
@farzyness Grok 4.7 needs a few more days to cook.
We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isnโt yet sufficiently rigorous in checking its work.
As a former journalist with masters degree and 20 years experience... Neither this documentary or review qualify. It's just 2 butt hurt, egotistical, attention seekers that some moron gave a platform. Walter Isaacson spent time with him and interviewed everyone. This is a 4 hour piece of historical fiction parading (the dangerous part) as journalism. It's not. And we should trash it. Publicly. And as much as possible.
I really think it would be a good look for the big AI companies to stop passive-aggressive shade-throwing at the others. Just let your models speak for themselves, which, in this case, lately has not been a good story for Anthropic. They should focus on improving their model so that developers start using it again
I am pleased to see that OpenAIโs new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. Nice work!
Evaluating and naming other labs turns out to be a great way to encourage them to train more aligned models. We will continue to do this until other labs pay more attention to safety. This is good for everyone and there is a lot of room left to go!
We solved prompt injection in practice for Claude models about two months ago. But prompt injection is a significant security risk no matter what model you use, and it is important that the industry similarly spends more effort to train their models to be resistant to prompt injection, among other elements of model alignment.
As models become more capable and central to businesses and economies, the risks only increase. We should be taking them seriously, and doing the right thing for our customers and the world.