Set per-workspace request-per-minute and request-per-day ceilings on an LLM integration to prevent runaway usage and ensure equitable access across teams.
| Where Can I Use This? | What Do I Need? |
- Prisma AIRS AI Gateway (Americas region)
|
- AI Gateway activated
- LLM integration in progress or saved
|
Rate limits control the volume of requests a workspace can make to an LLM integration
within a given time window. They protect against application bugs or unexpected traffic
spikes that could exhaust your LLM provider quota or drive unexpected costs. Rate limits
apply to the number of API calls, not to token consumption — use budget limits to cap
spending. When a workspace exceeds a rate limit, AI Gateway returns HTTP 429.