The Capgemini employee who exposed toddlers being abused at the creche on the company campus was fired first. Now she has been arrested for sharing the videos.
This is deeply disturbing.
Fable 5 silently fell back to Sonnet 5 without informing me.
I suspected the output wasnโt at the expected intelligence level, so I asked which model was actually being used. Thatโs when I discovered the fallback had happened behind the scenes.
Not great.
#anthropic#claude
@KaiXCreator I saw some Chinese characters in my output with Opus 4.8 yesterday, similar to how you see them with some of the open source Chinese models
We are entering the final era of benchmarks
Cognition just released FrontierCode, a coding eval built around real maintainer grade software tasks across major open source repos.
It measures merge-ability, so basically an anti slop benchmark - meaning correctness, test quality, style, scope discipline, and whether a real maintainer would actually accept the PR.
A model that saturates this bench would be capable of turning concise human intent into production grade changes across large codebases - so no more slop code.
Top score on Diamond is still only 13.4/100.
Excited to see what 5.6 and Mythos do here.
@asaio87 Agent loop, workflow, ralph all seem to have one goal and that's to increase token usage... I had a "deep research" trigger 101 agents, doing mostly superficial work ๐.. ppl need to watch out to not fall for the latest "influencer" driven narrative
Today weโre releasing DeepSWE, a new standard for agentic coding benchmarks.
On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.