What jumped out at me is how little of this is really about Claude specifically. You could swap in most agent frameworks and the same warning still holds: if your loop treats the model like an honest coworker instead of an untrusted request generator, you are asking for trouble. That part I buy immediately.
The strongest idea here is also the least glamorous one: stop conditions are not a UX detail, they are the control plane. A demo agent can “keep going until it feels done.” A production loop cannot. You need a hard boundary for completion, waiting on a human, explicit refusal, retry exhaustion, and resource budget. That feels obvious once stated, but in a lot of agent code I’ve seen, it’s exactly the thing that gets left fuzzy because everyone is focused on tool prompts and chain-of-thought theater instead of the boring loop mechanics.
I also like the article’s insistence that schema validation is not business validation. People still confuse “valid JSON” with “safe to execute,” which is how you end up letting a well-formed but nonsensical payment request through. The distinction between recoverable errors and terminal ones is useful too. If the model gets back AMOUNT_MISMATCH, maybe let it repair itself once. If it hits a permission failure or policy denial, stop pretending a second try will save you. That sounds simple, but it forces you to think about which errors are actually model-fixable and which are just a hard no.
The part I’d push on a bit is the confidence around allowlisting tools. “Just don’t give the agent the tool” is good security advice, but in real systems the messy part is not whether a tool exists. It’s whether the surrounding data, side effects, and escalation paths effectively reintroduce the same capability through the back door. So yes, tool minimization is smart. No, it is not the whole story.
The other thing I’d actually want to implement first is the trace. Not because it sounds principled, but because every serious agent incident becomes impossible to debug without a clean record of model version, inputs, tool calls, results, and approvals. If you cannot reconstruct the run, you cannot tell whether the agent failed, the prompt failed, or your own guardrails did.
This is one of those pieces that reads like “basic hygiene,” which is usually a good sign. The problem is that basic hygiene is exactly what production agent systems keep getting wrong.
Reference: Five Stop Conditions Every Production Claude Tool Loop Needs