Sitting in the barbershop waiting for my turn, quietly working on my projects through Codex Remote.
Everyone probably thinks I'm just scrolling.
They have no idea what I'm using, ๐
Built and shipped my own Android movement app with Codex in a surprisingly short time.
I was tired of Play Store apps asking for subscriptions or locking useful features and also too much ads, so I made exactly what I wanted.
No account. No subscription. No locked features.
Just software built for meโand already part of my day.
@arena How do you separate model quality from harness quality here ?
In long-running repository work, compaction, tool behavior and failure recovery can affect the result as much as the model.
I think a useful leaderboard should expose both
Your next website visitor may not be human.
GEO helps AI agents discover and understand your product.
WebMCP helps them use it.
Web apps can't be designed only for humans who search and click anymore. They need to be agent-ready too.
AI-agent requests grew 45% in one quarter.
Weโre adding support for WebMCP in the ChatGPT desktop appโs built-in browser and ChatGPT Sites.
When you visit a compatible website, ChatGPT or Codex can automatically use it to complete your task.
Update to the latest version of the ChatGPT desktop app, then just ask Codex to create a WebMCP-enabled app and deploy it to Sites.
Updating Codex on Windows is a loop:
Click Update โ app disappears โ never comes back โ reopen it โ update still isn't installed โ repeat.
Does anyone else have this issue ?
This is why I chose Hermes as source rather than as the final product. Its real value is that the harness is yours. I'm stripping it into a small, narrow runtime built around my workflow.
Control also means being able to remove almost everything.
Just finished the first very early v0 stage, and still a lot to do
https://t.co/ADsRcazx6f
@kimmonismus The interactive workload gains matter most for agents. In practice, I think long-context latency and tool-call speed matter more than peak tokens per second.
@DanDr1s How is trustworthiness measured here? In production-agent work, it includes detecting incomplete changes, recovering from tool failures, and preserving context across long runs.
Does the score test any of that ?