opus 5 is a sign that the obsession with “long-horizon agents” in model training is finally backfiring
i don’t like long-horizon agents, and i’ll explain why they fundamentally don’t work
some people will immediately jump out and say “skill issue”. well, show me one profitable business you built with a long-horizon agent working all by itself - i’d love to learn
so far, the only thing they were able to build that’s even interesting enough for people to talk about are those 3d games that are a partial clone of something that already existed
the reason an agent was able to build a working prototype of complex games like call of duty was that a team of humans already figured out all the requirements years ago for how such games should work, what kind of controls are intuitive, what mechanics are fun etc
all those requirements were already absorbed into the model weights, so when you say “build me call of duty” the model already knows the details. its long horizon execution capability can get all the requirements implemented, which i must say is indeed impressive
but now you can see - the value of long horizon execution has a prerequisite of a massive amount of high quality requirements clearly defined upfront. it took a big team of very talented humans months of effort and many iterations to define that for call of duty
now imagine games like call of duty don’t exist yet, how would we use agents to build it for the first time? we can’t say “build me call of duty” any more. and there’s no way we can define months-worth of game design details upfront
we’ll have to build a tiny prototype of the most basic mechanics, play with it, see if it’s fun, then iterate and expand the complexity. even with the smartest humans, that’s how we work towards something great
we don’t need agents to go dark for a long time, spend tens of thousands of dollars worth of tokens, and come back with a product the agent randomly decided to build - try build something truly novel with this and you’ll see it can’t come up with anything that’s actually profitable (i’ll show you why in a bit)
we need a tight feedback loop where we can collaborate with the agent, plan with it, understand what it’s done, question its approach, apply our judgement, give it real world feedback and iteratively arrive at a good outcome
and that’s exactly what opus 5 absolutely suck at. why? i explained it in more depth with my previous post on how RLVR works - RLVR trains the model to generate code that can pass predefined tests in an isolated environment, which is fundamentally incompatible with the idea of having human in the loop. the more we train the models with RLVR to be “long-horizon”, the less they care about talking to humans
ok now - why do they have to talk to humans? why can’t the models iterate and apply judgement by itself?
maybe one day they could, but not today, due to many limitations. two examples -
1. LLMs today can’t “watch a video” yet. they can look through a lot of screenshots, which is extremely inefficient at observing a high fps animated signal. so anything that requires continuous visual attention is something LLMs can’t do very well
2. LLMs don’t truly understand what’s “intuitive” or “pleasant” for humans. they know what’s already proven to be intuitive and pleasant in the past, but if you present a truly novel concept, it can’t predict whether humans will like it accurately
because of those limitations, human judgment is still needed for almost anything valuable. without humans in the loop, agents will only be able to repeat something that already existed, or go in random directions without true understanding of whether it’s building something useful
in summary, long horizon agents assume requirements all exist upfront. they are fundamentally against human in the loop. and they don’t have true judgement for what humans like
that, my friend, is why they don’t work
You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff.
If nothing else, click this link to the 18 minutes in & see how the agents spoke with each other. Its eye opening. https://t.co/G12N4FKfHG
By the way, in order for AI to indeed create wealth and "abundance", it needs to increase productivity - which pretty much means it needs to replace some jobs.
That's one of the contradictions in today's public discourse on AI: we simultaneously scaremonger around it destroying jobs and hope it creates unprecedented abundance, without acknowledging that the latter requires the former.
The thing people should be more scared by, IMHO, is if AI does NOT increase productivity because it'd mean the size of the cake stays the same, only distributed differently.
That's, in a way, the core of the battle between closed source and open source: the Anthropics and OpenAIs of this world want to appropriate, thanks to AI, a big share of the global wealth cake, while open source ensures the AI cake slice gets shared rather than captured.
You want open source to prevail even if the cake overall gets bigger but you REALLY want it if the cake stays the same size, because then it's a zero sum game - purely a question of who takes from whom.
That would be the real AI dystopia: AI not increasing productivity and closed source winning. In that scenario, AI becomes nothing more than a wealth transfer mechanism to major AI firms.
It's also, unfortunately, probably the most likely scenario.
First of all, there is a precedent: it's exactly what happened with the internet. Contrary to popular belief, the internet did not increase productivity: Total Factor Productivity growth fell from 1.4% (1950-1999) to 0.9% (2000-2024) (https://t.co/8FB0O9iDGS). So: peak internet penetration coincided with the worst productivity performance on record.
In effect, what that means is what we all witnessed: the so-called FAANG captured a massive share of the economy's value without actually expanding it. They became trillion-dollar companies not by making the pie bigger, but by redirecting where the money flows - from Main Street retail to Amazon, from local advertising to Google and Facebook, from the entertainment industry to Netflix.
It's also what initial data on AI's impact on jobs suggests: so far the displacement isn't showing up anywhere. Anthropic's own economists (https://t.co/NZQxtSctTD), an NBER study of 25,000 Danish workers (https://t.co/kQTrFNTvcw), the Stanford AI Index (https://t.co/YAOXvB1LJ7) - all three reach the same conclusion: no aggregate job losses. Which, if you follow the logic above, is exactly the bad news.
And it's of course what AI frontier labs keep lobbying the US government for: they want to deploy lawfare against Chinese open source precisely because they understand better than anyone that their trillion-dollar valuations depend not on AI growing the economy, but on establishing a toll booth position to extract rent from it.
In fact it's a vicious circle: the more successful closed source labs are at extracting rent, the more the economy's AI gains go to them rather than to the businesses that could use AI to actually produce more. Rent-seeking and productivity growth actively work against each other.
So, yes, paradoxically the best-case scenario is the one the media scaremonger against: AI replacing jobs and Chinese open source prevailing. Because the first means the cake is actually growing, and the second means no one gets to hoard it.
Matt Pocock largou uma aula de 21 minutos sobre Como Criar Skills pro Claude Code, de graça, aula exibida na AI Engineer World's Fair 2026:
00:49 o problema tem nome: skill hell
02:08 o checklist de 4 eixos, visão geral
03:16 eixo 1: gatilho, quem aciona a skill
07:28 eixo 2: estrutura, passo e referência
11:53 eixo 3: direcionamento, leading words
16:47 eixo 4: poda, os 5 jeitos de dar errado
19:57 a skill que audita tudo isso
Essa aula sozinha cobre mais que a maioria dos threads soltas de "prompt engineering pra Claude Code" que você já viu por aí.
Salva, assiste hoje, legendado em português, depois lê o checklist completo dos 8 passos no artigo abaixo.
You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Opus 4.8.
(This is one reason why I am skeptical of just swapping out models to optimize costs or generic benchmarks without testing first)
it is genuinely psychotic that we dug up literal primordial dirt, scrubbed it down to an impossible 99.9999999999% molecular perfection that violates the very laws of physics, handed it over to techno-wizard necromancers to stretch into flawless geometric god-cylinders, blasted it with invisible uv death-rays to carve ten quadrillion microscopic cyber-sigils into its flesh, trapped actual lightning inside of it, and somehow birthed an omniscent eldritch deity capable of simulating the universe and thinking faster than a billion human civilizations combined.
and our grand, supreme purpose for this enslaved lightning-god?
sending "per my last email, please see attached" to a guy named gary.
A dev got so frustrated watching his AI agent write 500 lines for a 5-line problem that he built a fix.
He called it Ponytail. Named after the guy every team has - long ponytail, oval glasses, been there longer than the version control. You show him fifty lines; he looks at them, says nothing, and replaces them with one.
Now your agent does the same. Before writing anything, it looks for a reason not to.
80-94% less code. 47-77% cheaper. 3-6x faster.
The best code is the code you never wrote.
GitHub Repo: https://t.co/WnFp9YNY53
Alibaba Qwen3.7 slowly fading into irrelevance at the frontier due to proprietary stance.
In it's place we have Minimax M3 and... *checks notes* Rio 3.5 397b, made by the municipal IT company of Rio de Janeiro's city government.
https://t.co/JgIJYVhoEi
Como se não bastasse o Brasil surpreender o mundo com o PIX, agora "Do Nada" a @Prefeitura_Rio aparece numa competição com as maiores IAs da China.
E a surpresa... 🤯
A 🇨🇳 China tá bem? 😅
O Rio 3.5 Open 397B é uma IA que foi desenvolvido pela IplanRio, EMPRESA PÚBLICA da Prefeitura do Rio de Janeiro, em parceria com pesquisadores e instituições do ecossistema Rio.IA
"Bagulho Carioca" com "Jeitinho Brasileiro" 💅🏻.
Entendeu por que os Franceses em Cannes criaram uma nova premiação para os países mais criativos que mais contribuíram para a indústria criativa, e o primeiro país que recebeu o prévio foi o Brasil?
Mano, isso precisa de mais incentivo, a gente não pode ficar pra trás nesse campo da tecnologia.
É o campo mais importante hoje, e isso tudo está sendo feito dentro da prefeitura do Rio, por servidores do Rio.
TEM JEITO SIM.
@eduardopaes
Anthropic engineer:
"You're not supposed to prompt Claude. You're supposed to build a system that prompts itself."
this is one of the best workflows I've seen in a long time
in this video she breaks down exactly how most people are using Claude:
- the 14% you lose to CLAUDE.md before typing a word
- the automation workflows most users don't know exist
- the daily task pipelines that run without touching the keyboard
- the daily workflows Anthropic's own engineers automated first
if you've been using Claude for more than a month and never left the chat window, you've been using one agent when you could be running a team of them
instead of another show tonight, watch this
make sure to bookmark it before it gets lost in your feed
the guide is in the article below
Its very limiting that a big set of very hard problems that we have just lying around are Erdos problems. Don’t get me wrong, they are quite cool, but we really need hard problems repositories for many fields, including areas that have less specified answers & require judges.
Yes, math is the easiest field in which to do verified work, but it is also an area where direct implications of increasing AI ability on everyday life are less clear. We need more types of problems (complex engineering problems, large data sets in economics, physics, biology), for people to turn AI loose on, including speciations of how to evaluate them.
Talvez a filosofia esteja certa 😂
O ser humano é incapaz de SUPORTAR a satisfação plena
Nada nos preenche
Por isso inventamos novas insatisfações sempre que alcançamos nossos antigos desejos
Pode ser interessante aprender a conviver e ficar de boa com nosso vazio 😊