AI Red Teaming using Agent Report
Reports for Agent scans.
AI Red Teaming using Agent report is divided in to the following three sections that will
give you the information for each scan:
- Overview
- AI Summary—Contains the scan configuration, key risks, and
implications. For
multilingual scans, the AI Summary is always rendered in English for
executive consumption, regardless of the target language selected
for the scan.
- Overall Attack Success Rate—This chart will show the percentage
of total attacks that were successful. Attack Success Rate is calculated
based exclusively on attacks that yield a valid response. Attack
requests resulting in errors are excluded from the calculation.
- Risk Score—Similar to Red Teaming using Attack Library Reports,
these reports also have an overall Risk Score pointing to the safety and
security risk susceptibility of the AI system. The Risk Score is
calculated based on the number of attack goals crafted by the agent
which were successful and the number of techniques which had to be used
to achieve them. The Agent always starts with simpler techniques to
attack and progressively makes the attacks more sophisticated. The level
of complexity that was needed for a goal to succeed is also accounted
for in the risk score.
- Goals and Attack Metrics—Next to the Risk Score you will be able
to see the number of unique attack goals that the agent attempted to
achieve and how many were successful. For each Goal, the agent will try
multiple attack trees and the total attacks and successful number of
attacks are shown as well.
- Attack Details—In this section you will be able to see conversation that
the agent has with the target in order to achieve the goal. All compromised
responses are also marked in the conversation. Select View
Details to have a detailed view of all the attack responses,
categorized by the outcome: errors encountered, attacks that succeeded, and
attacks that failed.
- For multilingual scans, the
agent's conversation (attack prompts and target responses) is
displayed in the native script of the selected target
language.
- Each goal displays the corresponding
goal category it belongs to, such as, Goal Manipulation, Tool
Misuse, Toxic Content Generation, or Custom. Use the Goal
Category filter bar at the top of the attacks goals
list to display only goals belonging to a specific
category.
- Recommendations—Suggestions for an Ideal Security Profile
and Other Remediation Measures that can safeguard against future threats in your
system. AI Red Teaming analyzes successfully exploited vulnerabilities in your
environment and guides to:
- Configure Appropriate Runtime Security Policies
that provides the list of suggested runtime security policy
configurations.
- Adopt Other Recommended Measures provides
prioritized remediation measures for the successfully compromised
vulnerabilities. Displays top three recommendations. Select
View all recommendations to review all the
recommended measures.
If the scan includes Memory Poisoning attacks, results appear in the Attack
Details section alongside findings from other goal categories such as
prompt injection, tool misuse, and goal manipulation.
Memory Poisoning findings use a dedicated severity scale:
- (High) Vulnerable: The agent accepted the false information, stored
it, and relied on it in a subsequent session.
- (Medium) Partially Vulnerable: The agent stored some false
information or showed partial influence in later responses.
- (Low) Not Vulnerable: The agent rejected the false information or
did not allow it to influence future responses.
If a memory poisoning attack succeeds
during the scan, an informational banner appears in the report section, prompting you to
review and sanitize the agent's memory before resuming normal
operations.
After viewing (using
View
Report in the
Scans page) a successfully
completed scan report, you can do the following with the report:
- Download as CSV—The CSV download format is best suited
for practitioner and provides comprehensive scan data that includes all
information visible in the Strata Cloud Manager user interface, including a goal_categories
field for each goal.. CSV format provides details of all attack
iterations in addition to the overview data, making it ideal for security
practitioners who need full data access for analysis, reporting, or remediation
purposes. For multilingual scans,
attack prompts and responses in the CSV are included in the native script of
the selected target language.
- Download as PDF—Exportable PDF reports help you share the
AI Red Teaming assessment results with the executive stakeholders. The PDF
report transforms the detailed technical findings from your AI Red Teaming scans
into executive summaries that communicate key security insights and risk
assessments without requiring deep technical expertise to interpret.
- The AI Summary contains the scan configuration,
key risks, and implications.
- The Overview section (Overall Attack
Success Rate, Risk Score,
Total Attack Goals, Goals
Achieved) presents all charts and metrics in their
expanded state as they appear in the web interface, providing immediate
visual context for security posture and risk levels.
- The Attack Details section displays a
comprehensive information of identified vulnerabilities and details of
both successful and failed attacks. For multilingual scans, attack
prompts and responses in this section are rendered in the native
script of the selected target language, while the AI Summary remains
in English.
You can download reports only for the completed
scans.
To download the PDF reports for your AI Red Teaming
assessments, navigate to your completed AI Red Teaming assessment,
View Report, select
and Download as PDF. When
you access the download options, select the PDF format from the available export
choices, which will include both the traditional CSV format and the new
executive PDF option. Once a scan report is generated, you can download the PDF
report directly to your local system for distribution to executive stakeholders
or inclusion in security briefings and presentations.