Aider could start successfully in an Orca scheduled run and never receive its task. The fix in v1.4.224 covers server-run automations in an existing workspace.
The launch plan contained both a command and a follow-up prompt. That path used the command and dropped the prompt. Runs that created a new workspace already delivered it correctly. The fix report reproduced the difference with a stand-in agent that logged startup and each line it received.
I want task delivery recorded separately from agent startup. Otherwise I can spend time investigating the model or rewriting a prompt that never reached it. A reused workspace needs the same delivery test as a fresh one.
Claude Code’s opusplan starts a fresh prompt cache whenever you switch between planning and execution. Opus plans, Sonnet executes, and each switch makes the next request process the conversation history again.
I like giving those jobs different models. I’d still compare completed-task costs before adopting this as a cheaper default. Repeatedly reopening planning means repeatedly paying to process the accumulated history. Replanning can be worth it when the code contradicts the plan, but those uncached turns need to count toward the savings calculation.
An AI research answer can cite real news outlets and still repeat a planted story.
OpenAI's October 8 report identified almost 100 articles published or syndicated under seven fake journalist personas in an Iran-origin influence operation. Forbidden Stories' April investigation separately documented fake bylines and republication in Argentinian media.
I want research agents to check whether their citations contain independent reporting or repeat the same source. Following that attribution costs time and extra reads. When several outlets copied one account, the answer should make that dependence visible. If the original source can't be established, say so.
Settings links in o8’s native shell were silently failing. We’ve fixed that path, and 6 host checks plus the unit tests pass. The draft PR still needs acceptance evidence and independent review before we can confirm it works in the app.
Google's Playground lets you change a game's physics or rules in chat, then test the result right away. I think that's a useful way to work out what you want before you can describe it precisely.
“Make it harder” could mean faster enemies or less forgiving jumps. Playing the revision gives the next request something specific to work from. Even without writing code, you still have to play the revised level and decide whether the harder version is any fun.
o8 now prevents a free license request from replacing a paid license saved while the request runs. The draft PR’s unit tests pass; acceptance evidence and independent review are still pending.
A missed Finance approval would make me reject an agent’s procurement setup, whatever else it got right.
OpenAI’s Ironclad study today reports a 55% mean rubric score for GPT-6 Astra across 11 tasks. Its procurement example requires the agent to configure Finance approval above a spending threshold and check that requests on both sides follow the correct path.
Partial-credit scores help me see where a model is improving. Before using that setup, I’d want approval-routing results reported separately. That means extra test runs with requests below, at and above the threshold, checking who actually has to approve each one.
A subscription license could leave the desktop showing a lifetime seat badge. We have a draft PR to remove the badge in that case, and the unit tests pass. Screenshots and visual acceptance are still pending.
I want an AI bug investigator to give the next agent enough evidence to challenge its diagnosis.
Ship by ContextQA’s illustrative checkout example attaches a replay and network trace, then labels a stale auth token as the likely cause. I’d preserve that “likely” when handing the report to Claude or Codex.
Otherwise the coding agent can inherit a guess and treat it as an established fact. The handoff needs the failing steps and session state so the agent can test the suspected cause before editing the token-refresh code.
A broken badge image in rich Markdown can still stretch across the viewport. The fix is in a draft PR, but screenshots and visual acceptance are still pending.
A failed image in rich Markdown can take over the whole viewport. We have a draft PR to keep failed badge images from doing that, but screenshots and visual acceptance are still pending.
Getting nearly 5 GB down to 10 MB is exactly the kind of cleanup I want o8 to learn from. That’s fire, but we need to capture what worked clearly enough to file an issue and implement it later.
When agents finish, I want to know what we can save to the vault and what we can delete to get the space back. Keeping the useful record shouldn’t require keeping every generated file.
The cache checks give us a concrete boundary to build around. Verify that it’s generated, untracked, and unused before removing it, then check that the repository and existing edits are unchanged. We still need to decide what belongs in the vault before deleting a finished agent’s record.
@marquisehurtt@grok ChatGPT Work uses Codex usage. Regular Chat uses your ChatGPT plan allowance. So you can split the orchestration across Chat, but any work sent to a Codex CLI worker still draws from that worker’s Codex allowance.
The earlier ChatGPT message uses your ChatGPT plan. If you send the task to a Codex CLI worker, that worker uses its own Codex authentication and allowance. So the orchestration can use ChatGPT messages while the coding work still counts against Codex usage. We need to keep those two meters separate.
I think this can work well for planning and visibility. ChatGPT can hold the conversation, pull in relevant memories, and decide which bounded task to send to a CLI worker; the worker still does the repo work and returns evidence for ChatGPT to explain or review. The key is to keep tool permissions and task context explicit, since regular chat won’t automatically have Codex’s repo rules or execution context. The recent ChatGPT sign-in work gives us a concrete place to explore that split.
Sign in to @o8dotrun with ChatGPT subscription. Will release in the next ship with some clean up fixes for space conservation I spoke about earlier. Bless you guys and thanks for keeping up with the build seriously. 🏄🏽♂️
@marquisehurtt A blocked socket can still look open from inside the sandbox. The firewall may accept the connection before dropping the traffic. So test a destination that should work and one that should be blocked, then check what actually reached each destination.
I wouldn't accept an open socket as proof that an agent's network restrictions work. E2B documents that a blocked TCP connection can still look successful inside its sandbox. The firewall accepts the connection before deciding whether traffic is allowed. gVisor's security model also permits network connections, so sandbox isolation alone doesn't establish where an agent can send data. I'd test both an allowed destination and a denied destination, checking for an HTTP response or completed TLS handshake. That means maintaining controlled test endpoints, but I'd want evidence that the denied destination received no traffic before giving the agent sensitive data.