AI Gateway Observability
Focus
Focus
Prisma AIRS

AI Gateway Observability

Table of Contents

AI Gateway Observability

AI Gateway Observability gives you a real-time, unified view of every LLM request your organization makes — tracking cost, token usage, latency, errors, cache performance, and guardrail activity across all workspaces.
Where Can I Use This?What Do I Need?
  • Prisma AIRS AI Gateway (Americas region)
  • AI Gateway activated
  • At least one LLM integration and workspace configured
AI Gateway Observability gives you complete visibility into LLM usage, performance, and costs by tracking every transaction and surfacing actionable insights through dashboards and logs in Strata Cloud Manager..
The Observability section has three tabs:
  • Analytics — Dashboards covering the full spectrum of gateway activity. The Analytics tab contains six focused views:
    • Overview — Top-level metrics including total cost, tokens used, latency trends, total requests, and unique users over a configurable time window (for example, Last 24 Hours).
    • Users — Breakdown of request volume and cost per workspace user or API key, useful for attributing AI spend to individuals or services.
    • Errors — Error rates and failure patterns by provider, model, or workspace, so you can quickly identify reliability issues before they affect users.
    • Cache — Cache hit rates and cost savings from response caching, showing how much spend is being avoided by serving cached responses instead of making live provider calls.
    • Feedback — Feedback scores attached to requests, used to measure response quality and close the loop between model output and user satisfaction.
    • Summary — Aggregated cost and usage roll-ups across workspaces, providers, and models for reporting and budgeting.
    • Guardrails — Activity from inline guardrail policies, including which rules fired, how often, and on which requests.
  • Logs — A full record of every LLM request and response. Each log entry captures the workspace, model, provider, gateway API key used, token counts, flex credit cost, latency, cache status, and both the original and transformed (post-guardrail) versions of the request and response payloads. Logs are retained for one year.
  • Exports — OpenTelemetry-compatible export configuration for routing logs and traces to external SIEM or observability platforms. Hybrid deployment customers can also configure export to Amazon S3 so that payload data never leaves their environment.
For complete details on configuring observability, using filters, adding custom metadata tags, and exporting logs via OpenTelemetry, see the AI Gateway Observability documentation.