Microsoft data exposure
Thousands of files hold card numbers or IDs. Which ones can anyone open?
In a typical tenant, 1Security's first scan finds thousands of files with card numbers, national IDs, credentials or health records - and a meaningful share of them sit behind anyone links, organization-wide sharing or external users. Knowing a file holds a card number is a report row. Knowing that 214 people, two OAuth apps and one AI agent can read it is a decision. 1Security detects 300+ sensitive information types with its own engine and lands every finding on the permission graph, so the map shows exposure, not just presence.
The problem
You know what the data is. You do not know who can reach it.
Classic data discovery ends with a list: these files contain personal data, these hold credentials, these have health records. What it cannot answer is the question every regulator, auditor and incident commander actually asks: who could read them, today, through which grant? In the tenants we connect to, 20-40% of files with sensitive content carry no label, and an ordinary account can reach hundreds of thousands of files.
The blind spot is wider than text. Card numbers live in scanned contracts, photographed IDs and screenshots pasted into decks - places a text-only scanner never looks. 1Security runs OCR as part of the same scan, so the map covers what an attacker would actually read, on a Business Basic license, with nothing extra to deploy.
And the map stays private: content is streamed into analysis and discarded. Only the detection type, match count and confidence persist - the matched card numbers or IDs themselves are never written to the database.
In practice
From "we have sensitive data somewhere" to a ranked exposure list.
Everything below runs on read-only consent and standard licenses - first findings the same day.
- 01
Let the scan find the data
Connect read-only. 1Security's engine reads file and email content for 300+ sensitive information types - credit cards with checksum validation, national IDs, IBANs, credentials, health data - plus OCR for images and scans, each detection with a confidence level. Purview detections, where present, sync in alongside as a second opinion.
- 02
Read Sensitive Info as a reach table
Every detection type is a row with live counts of the files, emails, sites, groups, users and apps that can reach it. "4,120 files with card numbers, reachable by 1,900 users and 3 apps" is the shape of the answer, and every count is a click into the list behind it.
- 03
Filter to what is exposed
Open Files, filter to with sensitive info + anyone link, then + shared with organization, then + external users. Add a framework filter - GDPR, HIPAA, PCI-DSS, SOX - and a confidence bucket so the number you quote is high-confidence only. In most tenants the anyone-link cut alone is hundreds of files.
- 04
Fix where one action removes the most
Sort by reach and start at the top: expire the anyone links on sensitive files, remove organization-wide sharing on the payroll library, restrict the app or agent nobody remembers consenting. Each is a suggested automation with a live match count, staged behind a 72-hour review window.
What makes it work
Three parts of the platform behind the map.
The map is where two systems meet: the detectors and the permission graph.
Access management
Effective access resolved through nested groups, sharing links and inheritance - the "who can reach it" half of every exposure number.
Explore the feature →Reporting
Regulator-grade numbers with confidence buckets, framework filters, saved views and exports - the map as evidence, not just a screen.
Explore the feature →Copilot security
The AI side of the same map: which sensitive information types sit within each Copilot agent's reach, before a prompt finds out.
Explore the feature →
FAQ
Common questions.
Do we need Purview or a premium license?
No. The 300+ detectors are 1Security's own engine and run on Business Basic upward. If you do run Purview, its SIT detections and labels are synced in and shown side by side, and the files that have sensitive content but no label become one filter.
Does 1Security store our sensitive content?
No. File and email content is streamed into analysis and discarded; what persists is the detection type, match count and confidence bucket. The matched values themselves are never written to the database.
What about images and scanned documents?
Built-in OCR reads jpg, png, tiff and other image formats, plus images embedded in documents - with sensible limits: files up to 50 MB, the first 5 pages of a PDF, up to 10 embedded images per document.
How long until the map is useful?
Reach counts start filling the same day you connect. Large tenants scan in stages for weeks, but the highest-exposure files - anyone links on sensitive content - usually surface within hours, and the automations show their live match count before you grant anything.
See your exposure as a ranked list, not a scan report.
Connect read-only and watch detections land on the permission graph - the anyone links on card numbers and IDs first, the same day.
Or keep answering "who can reach it?" with a shrug.