Estimates, Not Invoices
USD cost uses catalog rates (in, out, cache, reasoning) times the token buckets Pangolin recorded for the call. Cache-write tokens use the input rate. Missing reasoning rates use the output rate.
If the model ID is not in the catalog, the request has no USD amount and does not count toward USD budgets. Token budgets still sum tokens for that call.
OpenRouter, Vercel AI Gateway, and Custom may match a catalog ID with approximate pricing.
If the upstream omits usage, Pangolin estimates tokens from the request and response.
The check uses prior recorded usage. A request that would cross the cap can still complete, so usage can slightly overshoot. Reset periods are trailing windows from now, not calendar months. Daily means the last 24 hours.
How Usage Is Calculated
After each call, Pangolin records prompt, cache-read, cache-write, completion, and reasoning tokens.- Token budgets sum those counts.
- USD budgets multiply the same counts by catalog rates, then sum dollars.
Where to Set a Budget
Each budget has exactly one scope. You can add more than one budget on the same scope when the unit or period differs, for example a daily USD cap and a monthly token cap on the same provider.
All matching scopes apply together. A call can hit a provider budget, a resource budget, a role budget, and a key budget at once. Exceeding any of them blocks the request.
Fields
The editor uses Spend Type, Maximum Spend, and Reset Period.
Lifetime covers all recorded usage for that scope.
Over Budget
The gateway returns HTTP 429 with the messageAI usage budget exceeded for this request. The JSON envelope matches the API the client is calling:

