Stop fixing agents by hand. Kayba reads your traces, writes the fix, opens a PR, and tracks how it performs in production.
Trace → fix → PR → measured impact. Repeat. ⬇️
Introducing the Kayba CLI
Automated agent self-improvement from your terminal. Works with Claude Code, Codex, and more.
Upload traces → surface failures → improve your codebase. Repeat.
7-day free trial, no credit card required.
https://t.co/NWCwlWStj9
Kayba live on @ProductHunt
- agents fail the same way every run and never learn
- kayba analyzes traces and extracts what went wrong
- your coding agent deploys the fixes automatically
https://t.co/DXaBH5kOJm
benchmarked on τ2-bench: up to 2x improvement in agent consistency
the dashboard is built to make your agent self-improving out of the box. DM for a free 7-day Pro trial
the underlying framework stays open source and free forever: https://t.co/ZTspVAU66V
we 2x'd our agent's consistency. now you can too
https://t.co/NWCwlWRVtB catches your agent's mistakes, learns from them, and deploys better prompts automatically
upload traces. review skills. deploy
we turn traces into self-improving agents:
1/ analyze: upload traces, kayba finds what works and what doesn't
2/ review: every skill links back to its source conversation. you decide what stays
3/ deploy: approved skills become a prompt appendix. paste into your system prompt
ACE v0.8.0: RLM-style reflector
- Recursive Reflector: explores your traces via code execution in a sandbox, not just reading them
- TAU-bench integration: benchmark agents on standardized tasks
- Skillbook tools: clean, consolidate & merge strategies
https://t.co/giTotlMeFs
ACE v0.7.3: Claude Code Integration
- learn from your Claude Code sessions with `ace-learn` - no API keys required: uses your existing Claude subscription
- strategies auto-loaded into ClaudeMD for future sessions
https://t.co/lBrSppuolG
https://t.co/xCVnJRosaW
ACE v0.7.2: Agentic System Prompting
- new workflow to optimize system prompts from your traces
- ACE analyzes what worked/failed → generates prompt suggestions
- you review and decide what to implement
https://t.co/LujDhsW8vk
Claude Code + ACE self-learning loop 🔄
We ran Claude Code in a loop to fully translate our Python repo to TypeScript - fully autonomous and with zero errors.
Try it yourself: https://t.co/gcAMBiwjkg
Agent Prompt Optimizer: live on Product Hunt
- agents fail, it extracts what went wrong
- prompts update automatically so errors don't repeat
- learning loop instead of manual prompt engineering
open source, drop into existing agents in a few lines
https://t.co/g6DcPtomeg
shoutout to everyone who opened issues, PRs, and the feedback from our @browser_use integration: we finally addressed the 2 biggest issues (playbook bloat + blocking during learning)