@BIGBULLapp@trycaliber That's the harder half, and I'd be lying if I said it was fully solved. Automating the halt was the easy part. Getting from "it tripped" to "here's exactly what it was doing" without someone digging through logs is still mostly unsolved on our end too.
@catonooka Fair, that's a genuinely bad ratio then. What was the task, something touching a big chunk of the codebase, or a large context/file getting pulled in? That's usually what turns a normal-looking prompt into a session-ending one.
@KevinDonahuePRO The 35% number tracks with what I keep hearing directly, most finance leaders can't even get a clean cost breakdown, let alone tie it to ROI. Attribution has to come before the ROI conversation is even possible.
@BIGBULLapp@trycaliber Fair, and honestly true. A hard cap contains the bleed, it doesn't fix why the loop happened. That's a separate problem, the cap buys you the time to go find the cause instead of finding it via the invoice. Not claiming it replaces actually debugging the workflow.
@b1rkh0ff I see! Four different models across four stages plus feature work on top, that's a lot of places for the 50% to have actually gone. Any sense of which stage ate the most, or does it all just land as one number once it's done?
@MJHire The overage-invoice moment is the tell, by then the drift already happened weeks ago unwatched. Built Guardrail by NEAT for that gap on the coding-agent side. Are your clients seeing this mostly on dev tools, or across departments?
@trycaliber@BIGBULLapp The ownership question matters, but even with perfect attribution you're still finding out after the loop already ran. Built Guardrail by NEAT to catch it earlier, hard cap on the coding-agent side that blocks the next request before the budget's gone, not just after.
@jtwaleson@cursor_ai That's the scary part of "auto" routing, the vendor can change what's cheap and expensive under you with zero warning. Did you catch this from the bill after the fact, or did someone notice the model choice shifting in real time before the $800 landed?
@heman10x "Recording cost after execution is an autopsy" is the best one-liner Ive read on this. Same architecture on our end, block before dispatch, not log after. I Built Guardrail by NEAT for Claude Code/Codex specifically. Your leaky bucket approach is sharper for burst detection thou
Every AI cost horror story I've read this week has a different shape:
An orchestrator polling every 30s to check if it's done.
Tokens burned idle waiting on a script.
3 parallel sessions, no idea which one's the problem.
What's yours? Genuinely collecting these.
@ollie_thoughts@thsottiaux@victornunez@samarthap_@sama@paw_lean@Kon A mix of large and small with no clear single culprit is the hardest version of this to diagnose by feel. I built Guardrail by NEAT to break that down per session so you'd see exactly which one ate the 90% instead of guessing across the whole day.
@tanujbuilds The file-touch proxy is a clean way to get a reference graph for free out of data you're already storing. Would guess it also catches something document frequency misses entirely, an episode that got touched a lot but used completely ordinary language every time.
@vr000m That's the exact blind spot, per-session tokens without a model or effort-level breakdown means you can't tell Fable-instead-of-Opus from a thinking-effort bump, even though they'd drain very differently. Sounds like the TUI needs one more dimension, not more data.
@LegacyNL@thsottiaux Down to 0% usually means something ran longer or looped harder than the task needed, not that normal use is actually this expensive. Was it one session that did it, or does it creep down steadily across the day without one obvious culprit?
@My_Ai_Bi That's basically what Guardrail by NEAT gives you, live spend showing as you work instead of tabbing over to check. Might save you the Chrome window habit.
@Elenkova_dxb hat exact feature doesn't need to wait for an upstream release. Guardrail by NEAT sits in front of the whole fleet as a proxy and enforces a global cost cap across every agent, no version dependency on Claude Code itself. Might be worth a look.
@emerginghope_ This is the whole thesis, said better than most people say it. "No work starts without an estimate, a budget, and someone accountable" is exactly what's missing from every agent framework by default. Built the same guard into Guardrail by NEAT.