@kimmonismus You can use CodexBar to compare token usage side by side across the models you use. And right now, Opus definitely gives you more than GPT-6.1 Sol — while also being the smarter model. They’re simply not in the same tier.
@kaiostephens Can you actually access the Dot itself from a terminal/CLI? Astra already works in Codex CLI, but that uses the Work/Codex allowance. I’m wondering whether the Dot can be used directly from the terminal with its separate usage allowance.
We do the equivalent: each shard gets its own Supabase slot restored from a schema dump, reused warm. But DB tests aren't our bulk (3 shards × ~7 min in parallel): it's static guards, typecheck, and above all reds from state leaking between files within a shard. How do you isolate files: a DB per file, or transaction rollback?
We run 10+ AI agents in parallel (Claude Code, Codex) on our accounting ERP in production. They write code far faster than we can integrate it: the bottleneck is qualification before main.
Anyone solved this cleanly? Real-world experience welcome 🧵
@OHCAYGO We don't mark them stale: a verdict names the exact base+head pair, and the queue only accepts one naming the batch's own pair. An earlier green just doesn't match. Checks that read outside state (schema drift, advisories) always rerun at merge.
@OHCAYGO Good point. We already emit a verdict naming the exact base/head pair, but it only lives in CI logs that expire. Adding a durable per-batch record: tree hash that landed (rebase rewrites SHAs, the tree doesn't), base, and every check that ran. Thanks!
@zan_sheum Taking your advice: fixtures first, then the full suite once per batch, one PR per agent. Our data agrees: of our last 21 red queue runs, only 8 were real batch defects, incl. migrations clashing with another agent's work; 5 were flaky tests. Will report back!
@NabuSuite Classified our last 21 red queue runs: only 8 were real batch defects (≈1% of commits), several of them composition bugs: two agents' commits green alone, red together. The other 13: flaky/non-hermetic tests (5), runner env (3), judge/queue bugs (5)
@zan_sheum Thanks, that's where we're heading: catch the red on the agent's own branch so only that agent pays. How long is your e2e suite? And does the queue re-run it on landing? If not, how often do two green branches go red once combined? Our gate is ~45 min on one box.
Next we'll try one PR per agent so the queue can batch + bisect. Did that work for you at 10+ agents? Or did you go for:
– more CI runners?
– impact-based test selection?
– flaky-test quarantine?
Anything battle-tested, I'm all ears 🙏
Already tried: a homemade integration "train" (6 weeks of tooling, never worked end to end), local pre-qualification before the queue, batches capped at 100 commits. It limits the damage, but it doesn't hold up as we add agents.
@hokazuya Honestly, hard to complain
RTX 6000 Pro prices nearly tripled and GB10s roughly doubled. Worst case, you accidentally made a great resellable investment.
Sell the stack and get the H200s!
@elonmusk AI progress here is clearly impressive!
But the 37% CPA score needs context: 4 deliberately tricky benchmark tasks, no prior client context, and no clarification allowed.
Also, what was the actual level/seniority of the CPAs who scored 37%?
@superalesha Try Qwen3.8-Flash-Next: 125B/6B active (+51B n-gram) vs 320B/18B for GLM-5.3-Flash.
NVFP4 + Marlin on 4×3090, n-gram embeddings in RAM. We get ~105 tok/s decode, ~212 tok/s at 4 streams, ~2.8k tok/s prefill.
GLM wins on DeepSWE (63.4 vs 58.7) but Qwen holds up in prod.
@GoogleDeepMind By the time Google actually releases this model to everyone, this ranking will probably already be outdated.
That’s the frustrating part of Google’s AI strategy: great models, but top-tier releases often arrive too slowly and too inconsistently.
@daniel_mac8 I totally disagree. Low effort doesn’t make the controls and bug findings. I just audited the low effort code produced and catched a lot. It is fast but very limited