opus 5 is a sign that the obsession with “long-horizon agents” in model training is finally backfiring
i don’t like long-horizon agents, and i’ll explain why they fundamentally don’t work
some people will immediately jump out and say “skill issue”. well, show me one profitable business you built with a long-horizon agent working all by itself - i’d love to learn
so far, the only thing they were able to build that’s even interesting enough for people to talk about are those 3d games that are a partial clone of something that already existed
the reason an agent was able to build a working prototype of complex games like call of duty was that a team of humans already figured out all the requirements years ago for how such games should work, what kind of controls are intuitive, what mechanics are fun etc
all those requirements were already absorbed into the model weights, so when you say “build me call of duty” the model already knows the details. its long horizon execution capability can get all the requirements implemented, which i must say is indeed impressive
but now you can see - the value of long horizon execution has a prerequisite of a massive amount of high quality requirements clearly defined upfront. it took a big team of very talented humans months of effort and many iterations to define that for call of duty
now imagine games like call of duty don’t exist yet, how would we use agents to build it for the first time? we can’t say “build me call of duty” any more. and there’s no way we can define months-worth of game design details upfront
we’ll have to build a tiny prototype of the most basic mechanics, play with it, see if it’s fun, then iterate and expand the complexity. even with the smartest humans, that’s how we work towards something great
we don’t need agents to go dark for a long time, spend tens of thousands of dollars worth of tokens, and come back with a product the agent randomly decided to build - try build something truly novel with this and you’ll see it can’t come up with anything that’s actually profitable (i’ll show you why in a bit)
we need a tight feedback loop where we can collaborate with the agent, plan with it, understand what it’s done, question its approach, apply our judgement, give it real world feedback and iteratively arrive at a good outcome
and that’s exactly what opus 5 absolutely suck at. why? i explained it in more depth with my previous post on how RLVR works - RLVR trains the model to generate code that can pass predefined tests in an isolated environment, which is fundamentally incompatible with the idea of having human in the loop. the more we train the models with RLVR to be “long-horizon”, the less they care about talking to humans
ok now - why do they have to talk to humans? why can’t the models iterate and apply judgement by itself?
maybe one day they could, but not today, due to many limitations. two examples -
1. LLMs today can’t “watch a video” yet. they can look through a lot of screenshots, which is extremely inefficient at observing a high fps animated signal. so anything that requires continuous visual attention is something LLMs can’t do very well
2. LLMs don’t truly understand what’s “intuitive” or “pleasant” for humans. they know what’s already proven to be intuitive and pleasant in the past, but if you present a truly novel concept, it can’t predict whether humans will like it accurately
because of those limitations, human judgment is still needed for almost anything valuable. without humans in the loop, agents will only be able to repeat something that already existed, or go in random directions without true understanding of whether it’s building something useful
in summary, long horizon agents assume requirements all exist upfront. they are fundamentally against human in the loop. and they don’t have true judgement for what humans like
that, my friend, is why they don’t work
I spent a LOT of time through the hardest 3D prompts at Fable, it is a 45 min video, but I have 60+ very cool demos for you. Also prompts in the next post.
https://t.co/QPS5ccWZox
what if i told you... computer use can be faster on local models
moondream3 with its photon update today that gives it mac support can see your screen and use it with 1s latency, ty @vikhyatk
here we have whisper+qwen+moondream triple model pipeline working offline flawlessly
I implemented @GoogleResearch's TurboQuant as a CUDA-native compression engine on Blackwell B200.
5x KV cache compression on Qwen 2.5-1.5B, near-loseless attention scores, generating live from compressed memory.
5 custom cuTile CUDA kernels ft:
- fused attention (with QJL corrections)
- online softmax
-on-chip cache decompression
- pipelined TMA loads
Try it out: https://t.co/m5vkJxWIY6
s/o @blelbach and the cuTile team at @nvidia for lending me Blackwell GPU access :)
cc @sundeep@GavinSherry
you make money in memecoins by being very early
you make money on prediction markets by being very right
my prediction is that over time... retail will realize that being early, especially as the liquidity fragments, is much harder than simply being right
impressive to see Hyperliquid capture almost 2x more volume than DyDx v4 yesterday... without any token incentives and a dramatically reduced point allocation for perps traders compared to Season 1
HLP Vault also sitting at ~43% APR with ~$178 million deposited
one common misconception about Polymarket is that more markets leads to significant liquidity fragmentation
however, usdc can be used to make on multiple markets at the same time, so it actually increases earnings potential and therefore liquidity across markets
We're thrilled to announce our latest investment in Carv Protocol!
@carv_official's versatile data framework for gaming and AI surpassed $30M in its Node Public Sale in just 6 hours on Arbitrum and Solana.
Excited to support the future of Web3 gaming!
#investment#3xCapital
some thoughts on why this is Hyperliquid's $PURR bottom
1) $HYPE token 1 month+ out, which means $PURR is the only proxy token for the protocol
2) the IMMEDIATE next step is gated shitcoin deployments, which is good because
HL team will only let in those who airdrop to $PURR
prediction markets have seen PMF, the next step to large-scale volume is enabling leverage
this can be done 3 ways:
1) parlays (which are technically easier to implement)
2) perpetuals (which @LEVR_bet is working on)
3) leverage against tokenized positions
two tradeoffs related to memecoins deployed on permissionless blockchains are:
1) inconsistent liquidity seeding -- some may rug-pull after launch, some don't put enough liquidity, others put hidden taxes
2) lack of consensus on which ticker best represents a single meme -- a good example is when OpenAI's Sora came out, there was 10 $SORA tickers that all had the same purpose [this could actually turn into an extended blockspace auction]
@HyperliquidX is fixing 1) by creating a standard where every new token on their L1 must be deployed with a base amount of liquidity, and that liquidity automatically moves with the price -- this is nice because there is now a guarantee of no rugs, and some min amount that users can expect to fill against
Re 2), there's a lot of idea-maze exploration that can be done here; a few ideas include:
- having a fixed amount of coins with the same ticker
- votes/bribes on who gets a ticker first (basically a new meme pops up, and whoever pays the most gets the right to deploy a certain ticker over the next e.g. 72-hour period)
This is bad.
Crypto is not just about trading tokens, it's part of a broader ethos of protecting freedom and privacy and keeping power in the hands of the little guy.
And these values unfortunately continue to be under attack, globally.