Predict what an agent loop will cost before you run it, including the case that produces the alarming invoices: an agent retrying a failing step until something stops it.
Predicts the runaway-loop bill before it lands. Plug in your worst-case agent loop assumptions. The calc rolls forward, multiplies tokens by iterations, applies your caching hit-rate, and tells you when you blow the budget.
A per-task budget hook at 1,838,000 tokens covers this scenario with a 25% margin. Drop this into .claude/settings.json:
{
"hooks": {
"PreToolUse": [
{
"matcher": ".*",
"hooks": [
{ "type": "command", "command": ".claude/hooks/budget-guard.sh 1838000" }
]
}
]
}
}The failure mode this exists to model is specific: an agent attempts a task, the task fails in a way the agent cannot detect as terminal, and it tries again with slightly more context each time. Every iteration costs more than the last because context has grown. Without a hard iteration cap, that curve does not flatten. It accelerates, and it usually runs overnight.
The controls are unglamorous and effective. A hard iteration cap, a spend ceiling that halts the run, context pruning so growth is bounded, and prompt caching to cut the cost of the repeated prefix. Set the cap from a modelled number rather than a guess, and set an alert well below your pain threshold.