Agents can write code faster than teams can review, deploy, and maintain it. Today we’re introducing the Agent Development Lifecycle and the Cloudflare primitives that underpin it. https://t.co/ms4ifso5on
now that the dust is settling around the last wave of model releases, and i've had enough time to use all these models in practice, let me share a more complete set of thoughts
1. grok 4.5 is probably the single most significant event during the last couple of weeks
i've been talking with many heavy users across model families, and it's pretty much a consensus that grok 4.5 is the most "pleasant" frontier model to work with day to day. it's incredible how precisely the team behind it found this perfect sweet spot and created a model that's so fast, efficient and capable
it proved spacexai is now a 3rd real player in addition to anthropic and openai. they have real time data from the biggest public townsquare of humans, they have acquired a popular agent harness, and now they have proven they can build great models. and if you look closely, they are designing their own chips, they have their own data centers, they can send GPUs into the space and create tokens out of sunshine
holy shit
2. speaking of data, human usage over a harness proved to be extremely important for training good models. grok 4.5 was the first model that incorporated cursor's data and it made a massive difference compared to previous generations of grok
this explained why amazon mandates employee usage of kiro, why meta installed mass surveillance over employee devices, why anthropic bans 3rd party harnesses, and why google is still struggling with gemini - because they don't have a popular harness with mass adoption to collect the data
this is part of why i don't think the subsidized LLM subscriptions will end any time soon, because a wide consumer adoption is the best source of data collection. we're paying the subsidized tokens by teaching their models how to get work done
3. opus 5 flopped. almost no one likes it. the only people who like it seem to be using it to one-shot 3d games that look impressive but no one will ever buy
anything that AI can one-shot is just the definition of garbage, because if you can one-shot this thing with a quick prompt, you should know that it means billions of other people can also do it - you will not create anything of value this way
and this is just a symptom of a more fundamental problem that model training is heading down a slippery slope where machine verifiable outcome is dominating over human feedback
the latest training process rewards the agent for running for a long time and finishing a complex project, yet no longer seems to care about how the agent talks to its human
jargons, walls of text, "an honest mistake" - opus 5 showed us that we need AI that's more human friendly. let's not build a world where we end up working with robotic a**holes all day
4. fable 5 remains undefeated as the upperbound
i talked about this in my previous post about wisdom vs diligence. the benchmarks blend both together so it's not easy to see, but fable 5 is the GOAT on the "wisdom" dimension despite it not winning on every benchmark. if you used it meaningfully, you know what i'm talking about
that said, it seems anthropic is extremely paranoid about other players, including open models, reaching the same level of intelligence, which is an indication that the moat is not strong. kimi k3 is just a preview of what it looks like
at the same time, fable is the first time token cost is becoming a very real problem. it's the only model so far that i can't afford to keep using all day. i suspect this will remain true for a while, that we have to pick and choose what tasks to give to fable-tier models, not using them as a daily driver
5. openai is in an interesting position
gpt models have been great at efficiency, but now grok is also very competitive. gpt also haven't quite reached the same wisdom upperbound where fable is yet - although gpt 6 may change that
i think there are two angles for openai to pursue:
- continue to bet on efficiency, and go after enterprise adoption while claude is too expensive and grok has a brand tax to pay there. this is a very viable strategy
- or.. compete head on with fable on wisdom and win against them on the "human friendly" aspect. although traditionally this hasn't been the strength of openai models either, so this feels like a low-ROI option
alright, that's a bit of a long post but the landscape is just becoming increasingly complex. hope these thoughts are helpful in terms of providing a reference for how to rationalize everything happening
Introducing Nimbus - docs for the agentic web 🌧️
Nimbus is an Astro framework for building docs, made for human and agentic workflows.
- own every file, no theme to fork
- agent-native: llms.txt, .md twins, MCP
- features install as recipes your agent applies
- support for agentic authoring, with a built-in prose linter
- open source
We've dogfooded it on our Cloudflare Docs this week (which has 8,700+ pages) and cut build times from ~19m to ~7m now.
Your docs are product zero, they deserve the same care as the product itself.
It's still really early and a bit of an experiment, so I'm excited to see where this ends up.
Have you ever wanted to make your own font? Well I've written something that will hopefully dissuade you from that idea.
https://t.co/UC8l6JMlWm
It's free to read for the next 48hrs. Enjoy.
Durable Objects now supports apac-ne and apac-se location hints. Use these to fine-tune placement for users in Northeast or Southeast Asia and lower latency.
https://t.co/xyezdNrnYt
we relaunched @cloudflare's startups website and made the review process much faster.
https://t.co/OHLy9BM7zK
up to $350k in credits. apply plz, it's time to build
Cloudflare has integrated with Anthropic's Claude Managed Agents to provide a fast, isolated execution environment for autonomous code delivery. This means builders can scale agent workflows globally while strictly controlling access to private backends and easily customizing their agent’s tools and runtimes. https://t.co/HZ1kD0qVmO
https://t.co/gHoAUUfHAR - boot once, run everywhere.
A MicroVM that runs on hardware you already own.
Close your laptop and it hands off to another host.
Works across macOS, Linux, and Raspberry Pi. (aarch64)
Announcing a preview of the next edition of the Agents SDK — from lightweight primitives to a batteries-included platform for AI agents that think, act, and persist: https://t.co/jqwJXMo6tb
An experimental voice pipeline for the Agents SDK enables real-time voice interactions over WebSockets. Developers can now build agents with continuous STT and TTS in just ~30 lines of server-side code. https://t.co/fsXTBZzs3x