@AestherML Lol, yep, I was right about benchmaxxxing - https://t.co/QZamp6siHU
"While Gemini 4 has performed well on benchmarks widely used to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort."
@TokenGremlin Nah, AA score is 53 for it, so it is a decent model, but still behind the MOGus 5.5. Plus Google is always benchmaxxxing hard. Still hyped - if Pro is able to write as well as Flash and is only a bit dumber than frontier, then it might be the best creative writing model out there
@Ostrognienpi@gimchel7@MaZhavorsiAnni Why Ukrainians dislike Poland? I don't get it. I heard they also worship a Nazi collaborator who killed many poles during WW2.
@TokenGremlin Flop, too much hype for nothing - 6.1 sol is decent, but it should have been launched as 6 sol because Opus/Sonnet 5.5 just totally killed any hype 6.1 could have generated. The rest of the stuff are just small QoL upgrades or toys. Anthropic's razor focus on models is paying off
@robinebers Lol. Are you celebrating launching of $500 sub that has x25 limits, when a few weeks ago you had x20 for $200? Ultrafast - x8 for x6 price, wow, very exciting. 6.1 Sol is most likely not close to Opus 5.5 as they skimmed way too quickly. Dots is okay but just yet another toy.