Shadow Data Discovery
Shadow Data Discovery enables Enterprise Data Loss Prevention (E-DLP) to detect and categorize
shadow data within your organization using predefined and custom groups.
| Where Can I Use This? | What Do I Need? |
| Strata Cloud Manager |
- Data Security license
Enterprise DLP license
Or any of the following licenses that include the Enterprise DLP and Data Security licenses
- Prisma Access CASB license
- Next-Generation
CASB for Prisma Access and NGFW (CASB-X) license
- Data Security license
|
Contact Palo Alto Networks to enable Shadow Data Discovery on your
tenant.
Shadow data refers to sensitive information that exists within your organization but
remains unidentified and unprotected by your current data loss prevention systems. This
data includes dynamically and rapidly generated unstructured content such as:
Research and development data with prototypes, designs, and patents.
Email communications between investment bankers using jargon or coded language to
exchange insider trading information.
Device logs or custom telemetry data.
Financial documents from mergers and acquisitions.
Confidential documents concerning private partnerships, earnings reports, or code
repositories.
Feedback forms containing customer complaints or potential issues that could
compromise your organization.
You can't configure pattern-based matching for all types of sensitive data because each
definition lacks the complete contextual understanding of a document or payload. Trainable
classifiers can grasp contextual nuances, but they depend on manual uploads of known
sensitive documents that might not always be available or sufficient for training. These
limitations create gaps in your data security posture, allowing sensitive information to
go undetected.
Shadow Data Discovery enables
Enterprise Data Loss Prevention (E-DLP) to analyze your organization's
data and identify patterns without requiring you to predefine what constitutes sensitive
information.
Enterprise DLP runs summarization and categorization on your
file-based assets, collecting summaries for up to 100,000 documents and organizing them
into
categories, which are AI-generated groupings of similar documents
based on their content and context.
Enterprise DLP then maps these categories to
groups, which are broader classifications that align with common data types
such as source code, personally identifiable information (PII), or financial records.
Groups
make it easier to understand your data landscape and author data protection policies
against meaningful classifications rather than individual clusters.
Enterprise DLP maps categories to a set of predefined groups automatically, and you
can customize those mappings or create your own groups to reflect how your organization
structures its data. This approach helps you discover and protect sensitive data that
would otherwise remain hidden in your environment.
Shadow Data Discovery supports file-based scanning for the following SaaS apps onboarded
to Data Security (SaaS API).
- Amazon Simple Storage Service (S3)
- Atlassian Confluence
- Azure Disk Storage
- Bitbucket
- Box
- Citrix ShareFile
- Confluence Data Center
- Dropbox
- GitHub
- Google Base
- Google Cloud Storage
- Google Drive
- Office 365
- Quip
- Workday HCM
|
|
Start a Shadow Data Discovery scan so Enterprise DLP can
analyze documents at rest in apps you onboarded to Data Security. Enterprise DLP uses machine learning to
discover and categorize documents into groups based on their content
and context.
|
|
|
After Enterprise DLP scans your organization's shadow data,
analyze the results to understand the AI-generated groups and
categories and what types of data exist in your organization.
|
|
|
Customize how Enterprise DLP organizes your shadow data by
moving categories between predefined groups, creating custom groups,
and managing group-to-category mappings to reflect your
organization's data structure.
|
|
|
After you analyze the discovered shadow data, take remediation action
to create a group-based custom document
type from discovered files to prevent exfiltration of
sensitive data on subsequent scans.
|