I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
My first blogpost is out: https://t.co/GA6V36chR8
Like many of you, my feed has been dominated recently by Anthropic's Optimization team take-home.
TL;DR: They retired this "notoriously difficult" exam because Claude Opus 4.5 effectively solved it. So, they released the task, and everyone started the grind.
Yes, AI beat human candidates here, but AI-generated solutions aren't easy to follow if you aren't familiar with the domain. And they carry 0 educational value.
I didn't just want to see the solution; I wanted to understand the mechanics under the hood.
So I spent the weekend digging into the task. I started with a naive Python baseline that took 147k cycles. After three rounds of specific optimizations, I got it down to ~2,200 cycles.
That’s a 65x speedup—and likely would have passed the hiring bar half a year ago.
I’ve just published a full breakdown of the solution. I explain everything in detail and depth, but I kept it accessible. Hope you'll love the visuals!
Enjoy!
This is cool & all but what are these tasks? I've decided to look deeper into it and made some plots that OpenAI folks didn't include in the paper :(
I was trying to figure out why, on GDPval, Opus beats GPT-5. My main hypothesis - which I still buy - is that the model is bigger, knows the small quirks better, and does more compute per token, so it performs better. No surprises there. The alternative idea was that it just renders visuals better: for a long time OpenAI models were weaker at working with web pages, etc.
I noticed some of the benchmark tasks could be sensitive to that—things like laying out a presentation or a PDF brochure. I wasn’t about to manually skim 220 half-page prompts, so I ran them through an LLM and had it classify them.
Prompt: "the importance of good visuals and the style in the resulting work.
For the context of this field we may consider the performer to be sloppy but very smart, or messy but highly intelligent.
They don't care about the visuals, and the result will be slapdash.
How significant will the effect on the result be in this case?"
In 58% of tasks, according to GPT-5-high, there’s no effect or it’s negligible. In 8% of tasks, it’s very important. In theory that could explain the benchmark gap, but I wouldn’t call it compelling evidence.
More findings in the thread 👇
Мой любимый анекдот про математиков )
Трое математиков и трое физиков собираются ехать на поезде в другой город на конференцию. Они встречаются перед кассой на вокзале. Первой подходит очередь физиков и они, как все нормальные люди покупают по билету на человека. Математики же покупают один билет на всех. «Как же так?» — удивляются физики — «Ведь в поезде контроллер, вас же без билетов оттуда выгонят!». «Не волнуйтесь» — отвечают математики — «У нас есть МЕТОД».
Перед отправкой поезда физики рассаживаются по вагонам, но стараются проследить за применением загадочного «метода». Математики же все набиваются в один туалет. Когда контроллер подходит к туалету и стучит, дверь приотворяется, оттуда высовывается рука с билетом. Контроллер забирает билет и дальше все они без проблем едут в пункт назначения.
После конференции те же вновь встречаются на вокзале. Физики, воодушевившись примером математиков, покупают один билет. Математики не берут ни одного. — А что же вы покажете контроллеру? — У нас есть МЕТОД.
В поезде физики набиваются в один туалет, математики — в другой. Незадолго до отправления, один из математиков подходит к туалету, где прячутся физики. Стучит. Высовывается рука с билетом. Математик забирает билет и возвращается к коллегам.
Поэтому не пользуйтесь методом, если не знаете, как он работает )
@ar10rka@bunopus Uplifting Vision (UV) — система поддержки и наставничества, которая помогает сотруднику расти без угрозы увольнения.
Как тебе такая аналогия? Или у тебя была другая идея? 😃
@m_fetis Тут соль не в её ответе, человека в выходной можно понять. Тут в общей культуре и привычке многих именно звонить по ерунде (без предупреждения), когда можно написать. Уволила - не понравился ответ, и с такой начальницей всё равно работать, думаю, не комфортно
@m_fetis@AlexanderErok Вот точно! Это вообще олдовая привычка звонить по кейсу - когда можно в чате решить парой вопрос-ответ) Надо отучать людей звонить без предупреждения. У многих заметил привычку эту
Сегодня отмечается Международный день «Бросай свою ненавистную работу»
В этот день предлагается задуматься, приносит ли вам ваша РАБота радость и удовлетворение? Или у вас от неё только стресс и никакого счастья.