Enable optical character recognition (OCR) to scan images for sensitive information
with Enterprise Data Loss Prevention (E-DLP).
On
May 7, 2025,
Palo Alto Networks is introducing new
Evidence Storage and
Syslog Forwarding service IP
addresses to improve performance and expand availability for these services
globally.
| Where Can I Use This? | What Do I Need? |
- NGFW (Managed by Panorama or Strata Cloud Manager)
- Prisma Access (Managed by Panorama or Strata Cloud Manager)
Prisma Browser
|
Or any of the following licenses that include the Enterprise DLP license
- Prisma Access CASB license
- Next-Generation
CASB for Prisma Access and NGFW (CASB-X) license
- Data Security license
|
Strengthen your security posture to prevent accidental data misuse, loss, or theft by
enabling optical character recognition (OCR) for Enterprise Data Loss Prevention (E-DLP). Enabling
OCR allows Enterprise DLP to scan image files for sensitive information that
matches your Enterprise DLP data profiles.
Enterprise DLP uses two approaches to detect sensitive data in images:
- Custom and Predefined Regex Data
Patterns—OCR extracts text from the image and compares the extracted text
against regex patterns.
- Predefined ML-based Data
Patterns—Enterprise DLP inspects the image directly to identify
sensitive data patterns without extracting text.
Enterprise DLP supports detection in images containing alphanumeric English
characters, including non-English languages written in alphanumeric English
characters. Enterprise DLP also supports inspection of images that include
handwritten text and images scanned by another device.
Image quality affects detection accuracy for both approaches. For regex-based
patterns, image quality determines how accurately Enterprise DLP extracts
text. For ML-based patterns, image quality directly affects pattern identification.
Poor image quality can result in false positive and false negative detections.
Detection works best when images have:
High Resolution and Pixel Density (DPI)—Low image clarity can
prevent accurate character identification.
Example—cl being interpreted as
d or the letters
o and O being
interpreted as a zero (0).
No Image Noise and Artifacts—Random speckles, dust, or digital grain
found in scanned documents or compressed JPEGs can distort how Enterprise DLP interprets characters in images.
Example—Black speckles or dust next to a
P character causing it to look like an
R or a period
(.) looking like a comma
(,).
High Contrast and No Background Interference—Detection works best
with high-contrast images where text clearly stands out from its background.
Enterprise DLP can't effectively scan image text when a colored
background, watermarks, or poor lighting is present.
Example—Background and text color are too similar. Shadows and glare
obfuscate text in an image.
No Image Skew and Orientation Distortion—OCR scans in horizontal and
vertical rows. Image skews and orientation distortions can cause row
misalignment that disrupts the logical flow of data within the image. This
can result in text being cut and combined in unexpected ways. Skews and
distortions greater than 15° typically result in detection issues.
Example—Image of a credit card at a 45° angle might prevent Enterprise DLP from accurately detecting the full credit card
number.
| OCR Support |
|
Optimal Image Resolution
|
50 x 50 - 1,800 x 1,800 pixels
|
|
File Inspection Limitations
|
Enterprise DLP inspects 15
images per forwarded file at random.
|
|
Supported Image File Types
|
Supports the OCR inspection for all supported
Image File Types.
|
|
Supported File Types
|
Enterprise DLP supports extraction of supported image file
types from all supported
File Types.
|
|
Supported Languages
|
English, Japanese
|
Enterprise DLP does not support OCR for Microsoft Visio XML drawing (.vdx)
files that need rendering to display. For example, OCR can't inspect a .vdx file
if the XML is the drawing representation.