So opus 5 is verbose. Bossy and doesn’t listen. We all saw the shift. But the performance never left, it just needs better instructions.
Opus is genuinely a powerhouse, and anyone that’s say anything else has a skill issue.
@HexxRL Big part of the spike is due to codex being included in the desktop app and no longer separate download. So any person that clicked on it and ran one prompt counts as a new user. But most were already paid
@sama ”Its getting too expensive to create new models and our investors are losing faith in us since we have been spamming the word agi for 3 years with nothing to show for ”
We’re extending the 50% increase to weekly Claude Code limits through August 31.
We hope to make this a permanent change to our plans, but strong demand for our models means that capacity may be tight over the coming weeks. We’ll keep you posted as things develop.
People have to understand that a human subjected rating is not based onskillset and potential. Only on the average output. Which if you don’t know what you’re doing they perform similarly.
Now do a cinematic comparison and it’s not even close for Sd2.5
BREAKING: Seedance 2.5 by @BytePlusGlobal places 4th overall on Video Arena with an Elo of 1302.
It is now ByteDance’s highest-ranked video model, with four Seedance models placing in the top 10.
Congratulations to the team on the release!
@s_lale@DesignArena@BytePlusGlobal Higher elo means higher winrate not higher capabilities. Omni is great for simple prompts but lacks for multi image references and long form. Which is Seedance MOAT. Most use cases are simple = random sample becomes google favoured.
@Ianbasecoat@Its_Nova1012 Opus on xhigh does less mistakes for me, similar benchmarks and 20-30% cheaper. High also does fine for spamm tasks.
Medium hallucinates. But I would recommend xhigh over max generally.
@MarcosHernanz Im pretty sure any company could brute force that speed by just throwing a huge amount of compute, but its only to make headlines. This is Definetly not an efficient way of running a model
@jun_song They are saving their share price. First negative cash float since 2000. And going above and beyond would mean a free fall in their stock as they notice the cash inflow into ai is somewhat stagnant.
@makeamarkery Why would they compare a flash model to a pro one?
The complaint is that Google is taking 6 business months to release a new pro model.
@GoogleDeepMind pull out the whip
@learn22438@beffjezos Perhaps a grok / grok build destillation would make sense. One for the people and one for the businesses. Like how gpt has codex and the normal models. @elonmusk figure it out.