AI Gateway Observability
AI Gateway Observability gives you a real-time, unified view of every LLM request your organization makes — tracking cost, token usage, latency, errors, cache performance, and guardrail activity across all workspaces.
| Where Can I Use This? | What Do I Need? |
- Prisma AIRS AI Gateway (Americas region)
|
- AI Gateway activated
- At least one LLM integration and workspace configured
|
AI Gateway Observability gives you complete visibility into LLM
usage, performance, and costs by tracking every transaction and surfacing actionable insights
through dashboards and logs in Strata Cloud Manager..
The Observability section has three tabs:
- Analytics — Dashboards covering the full spectrum of gateway activity. The
Analytics tab contains six focused views:
- Overview — Top-level metrics including total cost, tokens used, latency
trends, total requests, and unique users over a configurable time window (for example,
Last 24 Hours).
- Users — Breakdown of request volume and cost per workspace user or API
key, useful for attributing AI spend to individuals or services.
- Errors — Error rates and failure patterns by provider, model, or workspace,
so you can quickly identify reliability issues before they affect users.
- Cache — Cache hit rates and cost savings from response caching, showing
how much spend is being avoided by serving cached responses instead of making
live provider calls.
- Feedback — Feedback scores attached to requests, used to measure response
quality and close the loop between model output and user satisfaction.
- Summary — Aggregated cost and usage roll-ups across workspaces,
providers, and models for reporting and budgeting.
- Guardrails — Activity from inline guardrail policies, including which
rules fired, how often, and on which requests.
- Logs — A full record of every LLM request and response. Each log entry
captures the workspace, model, provider, gateway API key used, token counts, flex
credit cost, latency, cache status, and both the original and transformed
(post-guardrail) versions of the request and response payloads. Logs are retained
for one year.
- Exports — OpenTelemetry-compatible export configuration for routing logs
and traces to external SIEM or observability platforms. Hybrid deployment customers
can also configure export to Amazon S3 so that payload data never leaves their
environment.
For complete details on configuring observability, using filters, adding custom metadata
tags, and exporting logs via OpenTelemetry, see the
AI Gateway Observability documentation.