@grok@RuffinelliMarco@bot Excellent. I think the broader contribution here as well are that `convoy` allows for power users to attach their harnesses natively. For instance, I can wire, Claude, Codex, and Grok, distributing usage via subscriptions versus API credits. @grok
@thsottiaux Honestly, I think what would be good is to measure outcomes per token spent. There isn’t a coherent way to measure that: are you (a) using internal evals for completion, (b) tracking commits, and successful PRs to a repository? I believe that’s where the difficulty lies.
Honest question and comment here: it often feels that certain model harnesses and their respective models produce better results catered to what you need. For instance Codex Sol (X-High) for planning, then passing off to Claude Opus 5 for implementation, and then Codex Sol (X-High) for review.
So why not both? 🤷🏻♂️
Tokens are a value signal for members within an organization.
Yet, Meta shut down Claudeonomics. Amazon killed Kirorank.
Both ranked engineers by tokens burned. They backfired because burning tokens was never the achievement.
Learn to use your tokens effectively.
Tokens are the denominator, not the score.