Spend caps for AI in production
Shipping AI features means shipping a metered cost with no natural ceiling. Before the clever prompt, you need a kill switch, token caps and per-tenant quotas — or one loop will bill you all night.
Shipping an AI feature is different from shipping a normal feature in one uncomfortable way: it has a variable cost with no natural ceiling. A slow endpoint is a performance problem. An AI endpoint with no guardrails is a billing problem, and billing problems compound while you sleep. Before you write the clever prompt, you write the controls that stop it from bankrupting you. If that ordering feels backwards, it is because most teams learn it the expensive way.
Here is the layered approach Propel uses, from the bluntest instrument to the most granular.
The first control is a kill switch. Every AI path checks whether AI is enabled before it does anything, so if a model provider melts down, a prompt injection starts driving unexpected calls, or costs spike for any reason at all, there is one switch that stops all of it immediately. This is the control you hope never to use and are grateful exists the one time you need it. It has to be a single, boring flag — not a deploy, not a config migration — because the moment you need it you do not have time for anything slower.
The second control is a per-request token cap. Any single call to the model is bounded. This matters because the failure mode of a language model is not usually a small overrun — it is a pathological input that produces a runaway generation, or a loop that keeps feeding context back in until a request that should cost a fraction of a cent costs orders of magnitude more. A hard per-request ceiling turns a potential catastrophe into a truncated response, which is a far better outcome to debug.
The third control is quota, and it operates at two levels: per-organization and per-user. Each tenant has an AI-credit allowance, and individual users draw against it. This is what stops one enthusiastic user, or one misconfigured integration inside a single customer, from consuming everything and either running up a bill or starving everyone else on the account. Per-request caps limit the size of one action; quotas limit the aggregate over a period. You need both, because they fail differently — one bad request versus a thousand ordinary ones.
Underneath all of it is usage logging. Every AI call is recorded — tokens in, tokens out, which tenant, which user. Without this you are flying blind: you cannot enforce a quota you cannot measure, you cannot attribute a cost spike to a cause, and you cannot tell a customer why their usage looks the way it does. Metering is not an afterthought bolted on for billing. It is the substrate every other control stands on. A cap you cannot measure is a cap you cannot trust.
There is one more decision that shapes the whole cost profile: whose models, and whose bill. By default Propel uses hosted models, which keeps the feature working out of the box. But an organization can bring its own key — OpenAI or Anthropic — and when they do, the usage runs against their account and their spend controls, not ours. This is not only a cost lever. It is a governance lever, because a customer with strict AI policies often wants their AI traffic on infrastructure they already govern, and the token accounting works the same on either path so nothing downstream cares which one is in use.
The reason to build all of this before the feature is nice to demo is simple. An AI feature without spend controls is not a lean version of the real thing. It is a liability wearing the costume of a feature. The controls are not the boring scaffolding around the interesting part. For anything running in production and touching a metered API, they are the interesting part — and the difference between an AI capability you can sleep next to and one you cannot.