Agent skill

Pii In Unstructured

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Detects PII in unstructured data including emails, documents, images, and logs using NER-based detection with spaCy and Microsoft Presidio, regex patterns, OCR integration, and confidence scoring.

Apache-2.0Auto-check passedDocuments & Office

Install Pii In Unstructured

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill pii-in-unstructured -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills pii-in-unstructured --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/pii-in-unstructured .claude/skills/pii-in-unstructured && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pii-in-unstructured
GitHub stars
301
Token cost
~2.4k tokens
SKILL.md length
764 words
Files
5 (incl. scripts, references, assets)
Skills in repo
280
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detects PII in unstructured data including emails, documents, images, and logs using NER-based detection with spaCy and Microsoft Presidio, regex patterns, OCR integration, and confidence scoring.

  • Tasks that involve Privacy and GDPR
  • SKILL.md covers Overview, Unstructured Data Sources at…, Detection Architecture and OCR Integration for Scanned…, plus 4 more sections
  • Runs Python scripts from its folder

What it does

Pii In Unstructured is an agent skill from mukul975/Privacy-Data-Protection-Skills. Detects PII in unstructured data including emails, documents, images, and logs using NER-based detection with spaCy and Microsoft Presidio, regex patterns, OCR integration, and confidence scoring. Keywords: PII detection, unstructured data, NER, spaCy, Presidio, OCR, regex, email scanning, document scanning.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Documents & Office, covering Privacy and GDPR. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Privacy and GDPR

Example prompts

  • “Use the pii-in-unstructured skill to detect PII in unstructured data including emails, documents, images, and logs using NER-based detection with…”
  • “/pii-in-unstructured”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pii In Unstructured loads about 2.4k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 764 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 764 words, ~2,422 tokens.

Download SKILL.mdSave it as .claude/skills/pii-in-unstructured/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
pii-in-unstructured
description
Detects PII in unstructured data including emails, documents, images, and logs using NER-based detection with spaCy and Microsoft Presidio, regex patterns, OCR integration, and confidence scoring. Keywords: PII detection, unstructured data, NER, spaCy, Presidio, OCR, regex, email scanning, document scanning.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
data-classification
metadata.tags
pii-detection, unstructured-data, ner, spacy, presidio, ocr, regex

PII Detection in Unstructured Data

Overview

Unstructured data — emails, documents, images, chat logs, call transcripts, and system logs — accounts for an estimated 80% of enterprise data and presents the greatest challenge for privacy compliance. Unlike structured databases where personal data resides in known columns, unstructured data contains PII embedded in free text, attached files, scanned images, and metadata. This skill covers detection approaches using Named Entity Recognition (NER), pattern matching, OCR, and hybrid pipelines, with focus on Microsoft Presidio and spaCy as implementation frameworks.

Unstructured Data Sources at Vanguard Financial Services

SourceVolumePII RiskDetection Challenge
Email (Exchange Online)2.1M messages/monthHIGH — names, account numbers, financial data in body and attachmentsMixed text and attachments; forwarded chains contain accumulated PII
SharePoint documents4.2TB across 1,200 sitesHIGH — contracts, KYC docs, customer correspondenceMultiple formats (docx, pdf, xlsx); embedded images
Teams chat890K messages/monthMEDIUM — casual references to customers, internal discussionsShort messages, abbreviations, context-dependent PII
Application logs50GB/dayMEDIUM — IP addresses, user IDs, error messages with PIIHigh volume, mixed with non-PII technical data
Scanned documents45K pages/monthHIGH — passport scans, signed contracts, medical certificatesRequires OCR; variable image quality
Call transcripts8K transcripts/monthHIGH — customers state names, account numbers, personal detailsSpeech-to-text errors, colloquial language
PDF reports12K documents/monthMEDIUM — financial reports may contain customer listsEmbedded tables, charts with PII labels

Detection Architecture

Microsoft Presidio Pipeline

Presidio is an open-source PII detection and anonymisation SDK developed by Microsoft, designed for integration with enterprise data pipelines.

Input Text/Document
       │
       ▼
┌──────────────────┐
│  Pre-processing  │  Format conversion, encoding normalisation,
│  (text extract)  │  OCR for images/scanned PDFs
└──────┬───────────┘
       │
       ▼
┌──────────────────┐
│  Presidio        │  Multiple recognisers run in parallel:
│  Analyzer        │  - NER model (spaCy/transformers)
│                  │  - Pattern recognisers (regex)
│                  │  - Custom recognisers (org-specific)
│                  │  - Context-aware enhancers
└──────┬───────────┘
       │
       ▼
┌──────────────────┐
│  Confidence      │  Each detection assigned confidence score
│  Scoring &       │  Threshold filtering applied
│  Filtering       │  Context enhancement boosts/reduces scores
└──────┬───────────┘
       │
       ▼
┌──────────────────┐
│  Results         │  PII locations, types, confidence scores
│  (structured)    │  Ready for classification, redaction, or alerting
└──────────────────┘
Component Details

NER Model (spaCy/Transformers):

  • spaCy en_core_web_trf model (transformer-based) for English NER
  • Detects: PERSON, ORG, GPE, DATE, MONEY, CARDINAL entities
  • Custom-trained NER for financial domain entities (IBAN, account numbers)

Pattern Recognisers (Regex):

  • UK National Insurance Number: [A-CEGHJ-PR-TW-Z]{2}\d{6}[A-D]
  • Email address: [a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}
  • UK phone number: (?:0|\+44)\d{10,11}
  • IBAN: [A-Z]{2}\d{2}[A-Z0-9]{4}\d{7}([A-Z0-9]?){0,16}
  • Credit card: \b(?:\d{4}[-\s]?){3}\d{4}\b (with Luhn validation)
  • IP address (v4): \b(?:\d{1,3}\.){3}\d{1,3}\b
  • Date of birth patterns: \b\d{2}[/-]\d{2}[/-]\d{4}\b
  • Vanguard account: VFS-\d{10}
  • ICD-10 code: [A-Z]\d{2}\.\d{1,4}

Context-Aware Enhancement:

  • Keyword proximity: if "National Insurance" appears within 50 characters of a NINO pattern, boost confidence by 15%
  • Section headers: if detected within a section titled "Personal Details" or "Patient Information", boost confidence by 10%
  • Negative context: if pattern appears within code blocks, SQL queries, or configuration files, reduce confidence by 25%

OCR Integration for Scanned Documents

Pipeline for Image-Based PII Detection
Scanned Document / Image
       │
       ▼
┌──────────────────┐
│  Pre-processing  │  Deskew, denoise, contrast enhancement,
│  (image)         │  resolution upscaling (if < 300 DPI)
└──────┬───────────┘
       │
       ▼
┌──────────────────┐
│  OCR Engine      │  Tesseract OCR (open-source) or
│                  │  Azure AI Document Intelligence (cloud)
│                  │  Output: extracted text with bounding boxes
└──────┬───────────┘
       │
       ▼
┌──────────────────┐
│  Presidio        │  Standard NER + pattern detection
│  Analyzer        │  on OCR-extracted text
└──────┬───────────┘
       │
       ▼
┌──────────────────┐
│  Confidence      │  Adjust for OCR quality:
│  Adjustment      │  OCR confidence < 80% → reduce PII confidence by 20%
│                  │  OCR confidence > 95% → no adjustment
└──────────────────┘
Document Type Detection for Vanguard
Document TypeOCR StrategyExpected PII
Passport scanAzure AI Document Intelligence (ID document model)Full name, DOB, nationality, passport number, photo (biometric)
Utility billGeneral OCR + address pattern recognitionFull name, address, account number
Medical certificateGeneral OCR + health NER modelName, diagnosis, doctor name, dates
Signed contractGeneral OCR + contract template matchingNames, addresses, financial terms, signatures
Cheque imageBanking-specific OCR modelName, account number, sort code, amount

Confidence Scoring Framework

Show full SKILL.md (308 more words)Show less
Score Composition
ComponentWeightDescription
Pattern match confidence40%Regex pattern specificity and validation (e.g., Luhn check for credit cards)
NER model confidence30%Model probability score for entity classification
Context enhancement20%Keyword proximity, section header, document type
Source quality10%OCR quality score, document resolution, text extraction confidence
Confidence Thresholds
Confidence LevelScore RangeAction
HIGH85-100%Auto-classify and auto-label; include in discovery report
MEDIUM70-84%Queue for human review; include in discovery report as pending
LOW50-69%Log for audit; do not auto-classify; available for bulk review
BELOW THRESHOLD< 50%Suppress; do not report unless specifically queried

Implementation Patterns

Pattern 1: Email Scanning Pipeline
python
# Conceptual pipeline for Exchange Online email scanning
# 1. Microsoft Graph API retrieves email messages
# 2. Extract body text (HTML → plain text conversion)
# 3. Extract attachment text (document parsing)
# 4. Run Presidio analyzer on combined text
# 5. Map findings to email metadata (sender, recipients, date)
# 6. Apply classification labels via Microsoft Purview
Pattern 2: Log File Scanning

For application logs, specific patterns dominate:

  • IP addresses in access logs
  • User IDs and session tokens in authentication logs
  • Email addresses in error messages
  • Stack traces containing file paths with usernames

Log scanning requires higher false-positive tolerance and volume-optimised processing.

Pattern 3: Chat Message Scanning

Teams/Slack messages present unique challenges:

  • Short messages with minimal context
  • Abbreviations and informal language
  • Customer names mentioned without surrounding keywords
  • Account numbers shared in fragments across messages

Strategy: scan message threads rather than individual messages to capture context.

Accuracy Benchmarks

Expected Performance by Data Source
SourcePrecision TargetRecall TargetKey Challenges
Email body text> 92%> 88%Forwarded chains, signatures, disclaimers
SharePoint documents (Office formats)> 90%> 85%Embedded tables, headers/footers
Scanned documents (OCR)> 85%> 80%OCR errors, handwriting, poor image quality
Application logs> 88%> 82%IP address over-detection, reference number ambiguity
Chat messages> 80%> 75%Short context, informal language, abbreviations
Call transcripts> 82%> 78%Speech-to-text errors, overlapping speech, accents

Integration Points

  • auto-data-discovery: Unstructured PII detection complements structured data discovery platforms
  • data-labeling-system: Detected PII drives automatic sensitivity label application
  • classification-policy: Detection results feed into classification tier assignment
  • data-inventory-mapping: Unstructured PII findings expand the data inventory to include document repositories and email systems

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/pii-in-unstructured of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Pii In Unstructured next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pii In Unstructured compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pii In Unstructured this skillmukul975/Privacy-Data-Protection-Skills301—~2.4kAutomated safety check: PassApache-2.0
Update Legal Pdfsremotion-dev/remotion63k—~287Automated safety check: PassCustom licence
Documenso Data Handlingjeremylongshore/tons-of-skills-marketplace2.8k—~2.2kAutomated safety check: PassMIT
Yc SaaS Drafterlawve-ai/awesome-legal-skills847—~2.4kAutomated safety check: PassMIT
Data Exportgustavscirulis/snapgrid1161 repos~2.9kAutomated safety check: NotesCustom licence
Transfer Impact Assessment Tia Oliver Schmidt Prietzlawve-ai/awesome-legal-skills847—~4kAutomated safety check: PassAGPL-3.0

Similar skills

  • Update Legal Pdfs

    remotion-dev/remotion

    Official

    Regenerate the downloadable PDF copies of Remotion's Terms, Privacy Policy, DPA Statement, and DPIA Statement after editing their docs pages.

    63k GitHub stars~287 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Documenso Data Handling

    jeremylongshore/tons-of-skills-marketplace

    Handle document data, signatures, and PII in Documenso integrations.

    2.8k GitHub stars~2.2k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Yc SaaS Drafter

    lawve-ai/awesome-legal-skills

    Drafts a customized Customer Agreement starting from the Y Combinator standard form SaaS template.

    847 GitHub stars~2.4k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Data Export

    gustavscirulis/snapgrid

    Generates data export/import infrastructure for JSON, CSV, PDF formats with GDPR data portability, share sheet integration, and file import.

    116 GitHub starsUsed in 1 repo~2.9k tokens
    Legal & ComplianceAuto-check: notes
  • GDPR Transfer Impact Assessment for Chapter V transfers under the EDPB Recommendations 01/2020 six-step methodology, the CNIL TIA Guide (January 2025), and EDPB essential guarantees.

    847 GitHub stars~4k tokensUpdated 8 days ago
    Legal & ComplianceAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 280 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    301 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    301 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Dpia

    mukul975/Privacy-Data-Protection-Skills

    Conducts Data Protection Impact Assessments for AI and ML systems per EDPB Guidelines 04/2025 on AI processing.

    301 GitHub stars~3.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    301 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    301 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    301 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Pii In Unstructured

What does Pii In Unstructured do?

Detects PII in unstructured data including emails, documents, images, and logs using NER-based detection with spaCy and Microsoft Presidio, regex patterns, OCR integration, and confidence scoring. Pii In Unstructured is an agent skill from mukul975/Privacy-Data-Protection-Skills. Detects PII in unstructured data including emails, documents, images, and logs using NER-based detection with spaCy and Microsoft Presidio, regex patterns, OCR integration, and confidence scoring.

When should I use Pii In Unstructured?

Pii In Unstructured fits situations like: tasks that involve Privacy and GDPR.

How do I install Pii In Unstructured in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill pii-in-unstructured -a claude-code`. Or copy the skill folder (skills/privacy/pii-in-unstructured in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/pii-in-unstructured in your project. Claude Code loads it when a task matches its description.

How do I install Pii In Unstructured in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill pii-in-unstructured -a codex`. Or copy the skill folder (skills/privacy/pii-in-unstructured in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/pii-in-unstructured in your project. Codex loads it when a task matches its description.

Can I use Pii In Unstructured in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill pii-in-unstructured -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pii-in-unstructured, .gemini/skills/pii-in-unstructured, .github/skills/pii-in-unstructured and .opencode/skills/pii-in-unstructured in your project.

What does Pii In Unstructured need to run?

Going by SKILL.md and its folder, Pii In Unstructured needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Pii In Unstructured access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pii In Unstructured safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pii In Unstructured use?

Pii In Unstructured is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pii In Unstructured use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Pii In Unstructured?

Skills that share tags, products or a category with Pii In Unstructured: Update Legal Pdfs (remotion-dev/remotion, 63k stars), Documenso Data Handling (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Yc SaaS Drafter (lawve-ai/awesome-legal-skills, 847 stars) and Data Export (gustavscirulis/snapgrid, 116 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pii In Unstructured?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 301 GitHub stars. The repository holds 280 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.