Use cases
Three problems, one endpoint. Each of these is the same risk check read from a different angle.
A budget written into a system prompt is a suggestion. The agent that overspends is usually not the one that was told to; it is the one that looped, or that was handed a task whose cost nobody estimated, or whose credential ended up somewhere else.
Card fraud models are trained on how people buy. An agent does not buy like a person: it transacts at machine speed, at odd hours, in tiny amounts, with counterparties it found seconds ago. The behaviour that looks anomalous for a human is Tuesday for an agent, and the behaviour that should alarm you (a retry loop, a quote overrun, a mandate that outgrew its scope) has no equivalent on a card rail at all.
Disputing an agent-initiated payment means answering questions nobody logged the answers to. What did the agent know? What was it authorised to spend? Was anything unusual about the payment at the time, or does it only look unusual now that you know how it ended?