Set a cost-based or token-based spending cap on a workspace's LLM integration to prevent unexpected charges and control flex credit consumption.
| Where Can I Use This? | What Do I Need? |
- Prisma AIRS AI Gateway (Americas region)
|
- AI Gateway activated
- LLM integration in progress or saved
|
Budget limits cap the total spending for a workspace on a specific LLM integration,
either in USD cost or in token count. AI Gateway tracks consumption per request and stops
routing traffic from the workspace once the limit is reached. Budget limits complement
rate limits: rate limits control request velocity, while budget limits control total
spending over a period. When a workspace exhausts its budget, AI Gateway returns HTTP 412
(Precondition Failed) until the budget resets or is raised.
AI Gateway tracks flex credit consumption at 1 flex credit per 4 characters of LLM
input or output. When you set a token-based budget, the limit is applied to the total
token count across all model calls from that workspace.