The prompt is usually the first thing people fix.
But in real agent systems, the failure is often somewhere else: evidence selection, context flow, tool calls, routing, workflow structure, or evaluation.
That is why agent optimization cannot stop at the prompt.
This feels right.
AI pilots don’t usually fail because the demo looked bad. They stall because teams can’t measure quality, trust, cost, and risk well enough to put agents into real work.
Evals are the starting point. The post-deployment loop is what turns those evals into continuous improvement.
Is your agent actually driving business outcomes?
And do you have a loop that learns from real production data and improves itself to hit those outcomes better?
If this is something you’re running into with production agents, I’d love to connect.
Interesting read. I agree that the real work starts after the trace.
Curious what felt different from other observability tools like LangSmith or Arize in actual use.
Also wondering whether fixing prompts and code is enough, or if the whole agent system needs to improve together.
More context on MEGA Optimus here.
We don’t assume the prompt is always the problem. The goal is to use evaluation to decide which system change should actually be kept.
https://t.co/i3RjdxmBuF
The prompt is usually the first thing people fix.
But in real agent systems, the failure is often somewhere else: evidence selection, context flow, tool calls, routing, workflow structure, or evaluation.
That is why agent optimization cannot stop at the prompt.
Building AI agents is getting easier.
Improving them is still hard.
Looking to connect with founders and builders working on reliability, evals, optimization, token efficiency, or AgentOps.
Curious what you’re building and what you’re struggling with.
Let’s connect, learn from each other, and build better agents together!