Microsoft 365 sensitive data discovery
Thousands of your files hold card numbers or IDs. Most were never classified.
In a typical tenant, sensitive data discovery finds thousands of files with payment cards, national IDs, health records or credentials on the first scan - and 20-40% of them carry no label at all. 1Security runs 300+ detectors plus OCR on standard licenses, imports what Purview already found, and puts a who-can-reach-it count next to every detection type.
- 300+detection types, plus OCR for scanned documents and images - on a Business Basic license
- 2engines kept apart: Purview detections imported, 1Security scan run independently
- 6reach counts on every detection type: files, emails, sites, groups, users, apps
The problem
A count of sensitive files is not an answer. Reach is.
Knowing that 4,000 files contain card numbers does not tell you whether a single one of them is exposed.
Classification tools stop at the count. An auditor, a DPO or an incident commander needs the next step: which sites hold those files, which of them sit behind an anyone link, which third-party apps have a standing read into them, and how many external guests are one click away. Without that join, "4,000 files with card numbers" is a number you cannot act on and a disclosure you cannot scope.
The second problem is coverage. Many tenants have never run content-level detection across every site and mailbox, and hope the sharing policy covers it. In practice it does not: it is common to find hundreds of files with personal data behind links that anyone on the internet can open, and dozens of consented apps with tenant-wide read into all of it. 1Security detects on the license you already have.
The third is scanned paper. Card numbers and passport scans live in JPGs and PDFs of contracts, screenshots and photographed forms - which text-only scanners cannot read. Discovery that skips images misses the files most likely to end up in a breach notification.
What you get
One row per detection type, with who can reach it
The Sensitive Info screen is a live table: not just "4,120 files contain IBANs", but "in 2 sites, reachable by 1,900 users and 3 apps".
Two engines, labelled
Purview Sensitive Information Type detections are imported on day one, and the 1Security scan reads content independently. Every detection records which engine found it, so you can compare them.
300+ detectors on standard licenses
Payment cards, IBANs, national IDs for dozens of countries, health identifiers, API keys and passwords - detected by 1Security on a Business Basic license, no premium license and no prior classification setup.
OCR for scans and images
JPG, PNG, TIFF, scanned PDFs and up to 10 embedded images per Office document go through OCR, because that is where card numbers and passport copies actually sit.
Reach across six resource types
Every detection type shows how many files, emails, sites, groups, users and apps it touches. Sort by the apps column to see which applications can read personal data.
Framework and confidence filters
Scope the view to GDPR, HIPAA, PCI-DSS, SOX, CCPA or FERPA, and to high, medium or low confidence - so audit figures come from high-confidence matches only.
Pivot to exposure
From any type, jump to Files with the same filter plus "anyone link", "external users" or "no label" - the exact files to fix, ready for a staged automation.
How deep it goes
GDPR framework, sorted by the apps column
One filter and one sort produce the list most DPOs have never seen: every app that can read personal data.
Filter Sensitive Info to the GDPR framework and sort by apps. Each row is a data-processing relationship. Most tenants find several apps here that were consented with a click, never assessed and never documented - a survey tool with Files.Read.All, a leaver's browser extension, an AI note-taker. It is common to see 5-15 apps with a standing read into personal data.
The second pattern is concentration. Pick the riskiest type you hold - payment cards, health data - and read the sites count. Two sites holding 90% of it means two remediation projects instead of two hundred. That reframing is usually the difference between a plan and a backlog.
The third moves from detection to exposure. From the type, pivot into Files with "anyone link" or "external users" added. In a typical mid-size tenant that list is a few hundred files - short enough to fix this week, long enough that nobody would have found them by hand.
In practice
From first scan to fixed files
How a tenant goes from "we have no idea" to a scoped, defensible number.
- 01
Connect read-only, import Purview
If Purview Sensitive Information Types are in place, their detections appear in Sensitive Info the same day. If not, nothing is missing - the 1Security scan starts on the same connection.
- 02
Run the 1Security scan
The scan streams file and email content through 300+ detectors and OCR, stores the detections and discards the content. Coverage builds continuously; a mid-size tenant has a usable map within days.
- 03
Read reach, not just counts
Sort by sites to find concentration, by apps to find undocumented processors, by users to see how many accounts can open payment data. Filter to high confidence before quoting anything.
- 04
Fix the exposed slice
Pivot to Files with exposure filters and stage a cleanup - expire anyone links, remove external access, request labels - behind the 72-hour review window. Point Purview DLP at what remains.
The difference
What the join to reach adds
Purview finds content. Everything below is the join to who can reach it.
- Which sites concentrate a given sensitive information type, ranked
- Which third-party apps and AI agents can read regulated data
- How many external guests and anyone links are one click from it
- Detections from two independent engines, each labelled with its source
- OCR results from scanned documents and images, not just text files
- Detection on Business Basic - no premium add-on required
- Framework filters that hand each auditor only the data their regulation covers
- The pivot from a detection type straight into sharing and exposure filters
Licensing and privacy
Standard licenses, no stored content
Everything 1Security detects itself runs on a Business Basic license. Content is streamed into the analysis process and discarded when it finishes - only the detection, its confidence and its location are stored, and no third-party AI service sees your files in the default configuration.
- 6compliance frameworks as scoping filters: GDPR, HIPAA, PCI-DSS, SOX, CCPA, FERPA
- 3confidence levels on every detection, so audit figures and backlog stay apart
- 10embedded images per document run through OCR - where card numbers hide
Related
Where this fits
Discovery answers where the data is. Label coverage answers whether it is protected, and file permissions answer who can open it. An audit asks all three.
Label coverage
The 20-40% of sensitive files that carry no label, per site, with a gap list you can act on.
See labels →File permissions
Who can open the files these detections live in - direct grants, groups, links and inheritance resolved.
See files →App governance
The 30-50 consented apps in a typical tenant, and which of them hold a standing read into regulated data.
See apps →
FAQ
Questions teams ask first
Do we need a premium license or an existing Purview setup?
No. The 1Security engine detects on its own and runs on a Business Basic license. If you have Purview, its detections are imported alongside ours - existing investment shows up immediately and gets an independent second opinion.
Does it read file contents, and does it keep them?
It reads them - that is how content detection works. Content is streamed into the analysis, run through the detectors and OCR, and discarded; only the detection type, confidence and location are stored. No third-party AI service sees your files in the default configuration.
How long until we have a usable map?
Purview detections show the same day you connect. The 1Security scan builds coverage continuously - a mid-size tenant has a usable map within days, and the largest tenants keep scanning in the background while you already work with what is found.
Can we scope a view to one regulation?
Yes. Framework filters cover GDPR, HIPAA, PCI-DSS, SOX, CCPA and FERPA, so each auditor gets a view with exactly the data their regulation covers, and nothing else.
Where does remediation happen?
Two places. Exposure fixes - expire links, remove external access, request labels - are staged in 1Security behind the review window. DLP and auto-labelling stay in Purview; this tells you which types, in which sites, reachable by which apps, so the policy you write there is aimed at something real.
Find your card numbers, IDs and health data this week.
Connect read-only and detections start landing with reach attached. The apps column is usually where the first uncomfortable conversation starts.
Or keep assuming the sharing policy covers it.