3/3 Worse, this unnecessary rigor burned through a significant amount of my usage allowance and created needless cost. It kept going without informing me, making the process a black box. When I finally inspected the work, I found complete over-engineering.
@OpenAI
I ran into a case where Codex’s pursuit of rigor ruined the work. I asked it to reproduce an existing evaluation, but it changed the assumptions and inputs without asking—not because it found a real problem, but because it could not fully prove there wasn’t one.
2/3 It then continued with the downstream evaluation, making all the results unusable. I repeatedly warned it not to over-engineer the task, but it did not stop. Its pursuit of rigor ultimately overrode the user’s actual goal.