it depends onwho uses it and what they’re using it for. users who don't work in/on AI might find it much easier to play with, more intuitive, familiar UX like an app they already use, so no need to "learn" a lot to make it work or master it. (from my friends work in hedge funds) and it's always "on"
tried it for an hour today, fwiw, the "cloud" part is the direction to go, I think
what're the best state management products/tools/projects for agents right now?
my use cases
- cross-session / persistent memory
- state restore & sync across different agents and devices
- shared states so agents team can work together.
any recommendations?
@claudeai respect protecting the model from distillation. but it's a little chilling to think we're drifting toward a future where one or two companies hold this level of intelligence, and can decide who else gets to build it, or even access it.
loop engineering is judgment compression.
a good loop doesn’t just automate the work. it compresses the judgment: what counts as done, what can be verified, when to retry, when to escalate.
my 3R rule:
- repetitive
- rubric-able
- return-bearing
if a flow has all 3, it wants a loop.
vertical agents just collapsed into a one-click install. codex got the form right: you don't build the app, you install the role - its skills, tools and apps, onto a client you already live in.
apps don't own the front door anymore, agent clients do. the question isn't how to bridge an agent into your app, it's whether your app is even reachable in the agent clients where your users now live. go in, get the reach, but you don't own the user anymore. you're the tenant now.
We’re making Codex more useful for your work by expanding plugins beyond individual tools.
These plugins turn Codex into a specialist for a specific role with a single install, no coding required.
Codex can access 62 popular apps and 110 skills for work across sales, data analytics, creative production, product design, and public equity investing.
https://t.co/nunrYP2uMI
codex sites isn't an app builder. it's the browser getting pulled inside the agent.
we spent two years putting agents into the browser: dia, extensions, copilots. codex and cursor are reversing it: the agent client browses the web with built-in browsers (flagged last week) and now builds it too with sites. both ends.
the agent client isn't replacing the web. it's replacing the browser. it reads the web like the browser did, and writes it like the browser never could.
Building apps has never been easier.
With Sites, Codex can turn your work, ideas, and plans into an interactive website or app your team can explore, use, and share with a URL.
Rolling out to Business and Enterprise plans, before expanding more broadly.
your job title is already a lie.
you're not an IC. you're not a manager. you're an individual framer.
you still do the work, just through a team of agents now. the job stopped being contributing. it became framing.
and framing is the only thing that separates people now.
opus 4.8, finally a win after six weeks of regressions.
4.7 fought its own harness, couldn't leave claude code to finish unattended, and like a lot of people I made codex my main driver. 4.8 won that trust back.
I never left for the model, claude still has more taste. I left for the app. the codex app IS the best agent client out there right now, claude code is still a terminal bolted to a fragile app.
the model recovers the work. the app wins the user.
still behind on the half that made me leave.
the trend here: agent clients shipping browsers as the surface for 3p "agent extensions".
cursor agent + built-in browser. codex + built-in browser. same direction. anything that works with agents becomes usable inside the agent client.
(thought this was https://t.co/XNDk8TqYz3 at first glance
html is NEVER the new markdown.
format follows audience. one isn't replacing the other, they were never in the same category.
pre-agent era, software optimized for one kind of communication: human to human. agent era, we use file types optimized for different communication types: agent to agent, agent to human.
in my work today, three file types serve three jobs.
markdown is still the most effective for agent to agent. text-based information that compounds. no tag overhead, agent focuses on what matters.
html for ANYTHING that needs visualization, mainly agent → human. denser, interactive. slides, posters, dynamic reports.
json for structured data, machine boundaries.
thariq named the real pain. being in the loop matters, and long markdown plans go unread. that's why html surfaced. it's an agent to human communication.
the act of planning is the key, not the plan itself. that stays. the trend that emerged from it ("replace markdown with html") doesn't.
Soooo @trq212 has straight up changed my life with these 5 words:
"HTML is the new markdown."
It's so obvious in hindsight: while .md is simple to write and agents can read it well, it's a total slog to eyeball as a human.
In this special ep recorded live at Code with Claude, Thariq walks me through how he:
- uses HTML artifacts as interactive specs
- builds throwaway micro-UIs
- maintains a living design system in HTML
- prompts Claude with "whatever is needed" to give it room to actually think
He also tells us what comes after the SWE + PM role 👀
As always, a huge TY to our amazing sponsors:
🔀 @celigoinc - Intelligent automation built for AI: https://t.co/wXLyXeR53y
🪪 @withpersona - trusted identity verification for any use case: https://t.co/l6yr4t26Un
Watch now on YT: https://t.co/svRSIIBTPq
good thing: now I get $200 credit per month
bad things:
- forced to go back to terminal for non-coding works.
- need to reconstruct my setup and workflows to make most use of my "subscription"
isn't the priority right now to make token consumption a more joyful thing?
Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage.
The credit covers usage of:
- Claude Agent SDK
- claude -p
- Claude Code GitHub Actions
- Third-party apps built on the Agent SDK
every 4-6 weeks, I become the bottleneck of the team (me and my agents). not just new model releases, but also skills, filesystems, new infra, harness upgrades.
my agents are sitting idle right now, waiting on my input. that's the signal, it's time to upgrade the "team" workspace again.
skills, tools, filesystems, folders as agents, remote machines as workspaces.
the acceleration sometimes creates anxiety. the upgrades always bring the excitement of becoming more capable.
it's the best of times.
it's 10pm on a tuesday
guess I'm officially taking the QA role for anthropic
2-3 builds a day, but I'm locked in 🫡
time to prove I'm better than agents!
same reason why i got a 5x codex plan as backup. hard move since cc is my fav since the first day it released.
it's not "dumb", it just doesn't click sometimes. my guess is: 4.7 might requires calibration on almost everything, skill, prompting and flow. I think they released the model before finish tuning cc perfectly.
been testing codex with GPT 5.5 last few hours
and the question comes again: which sub stays, which goes. every time a "great" model release or a "breakthrough" harness update. why can't my agents just get upgrades for free whenever anything in the stack improves?
I have 20x claude max and 5x codex pro, roughly 2:1 in actual use. sometimes mix them, design on claude code, spin up codex for execution.
it works, but it's painful. I hate it. spent hours wiring things up, tweaking settings, moving context around. I "had to" get the codex plan two months ago.
ideally, one plan should be enough, my agents pick the right harness and model for the task.
today, agent = harness. claude code, codex, amp, droid, and a dozen others, each one running as an individual agent. fine when there were three. there are 20+ now. and the best one changes every week.
I don't think this is the right way for "agent" to evolve. from a user's standpoint, an agent shouldn't be bound to a specific harness. it should navigate and use them on demand. harness stays with its model, aiming to make the most of it. identity, skills, persona travel with my agents.
how I want my designer engineer agent to work: same persona, same identity, same skills. uses claude code for design, switches to codex for execution. I never think about it, never pick. so when a new model ships, I'm using it on the same day. a better harness feature lands, I'm routed there.
stack improves, my agents and I benefit. no new sub. no new setup.
20 harnesses is not variety. it's a tax, and I'm just tired of paying it.