Agent skill

Data Labeling System

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Implements data classification labels and tagging systems including metadata tagging, DLP integration, automated label propagation, user-applied labels, and label inheritance rules.

Apache-2.0Auto-check passedData & Analytics

Install Data Labeling System

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-labeling-system -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills data-labeling-system --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/data-labeling-system .claude/skills/data-labeling-system && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-labeling-system
GitHub stars
295
Token cost
~2.8k tokens
SKILL.md length
952 words
Files
5 (incl. scripts, references, assets)
Skills in repo
278
Repo updated
First seen
Licence
Apache-2.0

At a glance

Implements data classification labels and tagging systems including metadata tagging, DLP integration, automated label propagation, user-applied labels, and label inheritance rules.

  • Tasks that involve Data cleaning
  • SKILL.md covers Overview, Label Taxonomy, Microsoft Purview… and User-Applied Labelling, plus 4 more sections
  • Runs Python scripts from its folder
  • Tasks that involve Privacy and GDPR

What it does

Data Labeling System is an agent skill from mukul975/Privacy-Data-Protection-Skills. Implements data classification labels and tagging systems including metadata tagging, DLP integration, automated label propagation, user-applied labels, and label inheritance rules. Covers Microsoft Purview sensitivity labels and enterprise labeling architecture. Keywords: data labeling, sensitivity labels, metadata tagging, DLP integration, label propagation, Purview, classification.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Data & Analytics, covering Data cleaning and Privacy and GDPR. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Data cleaning
  • Tasks that involve Privacy and GDPR

Example prompts

  • “Use the data-labeling-system skill to implement data classification labels and tagging systems including metadata tagging, DLP integration…”
  • “/data-labeling-system”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Labeling System loads about 2.8k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 102 tokens; SKILL.md has 952 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 952 words, ~2,777 tokens.

Download SKILL.mdSave it as .claude/skills/data-labeling-system/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
data-labeling-system
description
Implements data classification labels and tagging systems including metadata tagging, DLP integration, automated label propagation, user-applied labels, and label inheritance rules. Covers Microsoft Purview sensitivity labels and enterprise labeling architecture. Keywords: data labeling, sensitivity labels, metadata tagging, DLP integration, label propagation, Purview, classification.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
data-classification
metadata.tags
data-labeling, sensitivity-labels, metadata-tagging, dlp, label-propagation, purview

Data Classification Labels and Tagging System

Overview

Data classification labels are the operational mechanism through which classification policy is enforced across the enterprise. Labels attach classification metadata to data assets — documents, emails, database records, and cloud resources — enabling automated enforcement of handling requirements through DLP policies, access controls, and encryption. This skill covers the design and implementation of a labelling system using Microsoft Purview Information Protection as the primary platform, with architecture patterns for automated labelling, user-applied labelling, label inheritance, and cross-platform propagation.

Label Taxonomy

Vanguard Financial Services Label Hierarchy
Vanguard Classification Labels
├── Public
│   └── (no sub-labels)
├── Internal
│   └── Internal - Project Confidential
├── Confidential
│   ├── Confidential - Customer Data
│   ├── Confidential - Employee Data
│   ├── Confidential - Financial Data
│   └── Confidential - Legal
└── Restricted
    ├── Restricted - Special Category (Art. 9)
    ├── Restricted - Criminal Data (Art. 10)
    ├── Restricted - AML Investigation
    └── Restricted - Board & Strategy
Label Properties
LabelColourVisual MarkingEncryptionDLP PolicyAuto-Apply
PublicGreenFooter: "Vanguard Financial Services — Public"NoneNoneNo
InternalBlueFooter: "Vanguard Financial Services — Internal Use Only"OptionalWarn on externalNo
ConfidentialAmberHeader + Footer: "CONFIDENTIAL"Azure RMS (AES-256)Warn + audit external; block personal emailYes (when PII detected with >85% confidence)
RestrictedRedHeader + Footer: "RESTRICTED" with red background; watermark on printAzure RMS (AES-256, double key encryption)Block all external; block USB/print; alert DPOYes (when Art. 9/Art. 10 data detected with >85% confidence)

Microsoft Purview Implementation Architecture

Component Architecture
┌─────────────────────────────────────────────────────────────────┐
│                    Microsoft Purview                             │
│                                                                  │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────────────┐  │
│  │ Sensitivity   │  │ Auto-labelling│  │ DLP Policies          │  │
│  │ Labels        │  │ Policies      │  │ (endpoint, email,     │  │
│  │ (definitions) │  │ (rules)       │  │  SharePoint, Teams)   │  │
│  └──────┬───────┘  └──────┬───────┘  └──────────┬───────────┘  │
│         │                  │                      │              │
│         ▼                  ▼                      ▼              │
│  ┌──────────────────────────────────────────────────────────┐   │
│  │              Unified Label Application Engine             │   │
│  └──────────────────────────────────────────────────────────┘   │
│         │                  │                      │              │
└─────────┼──────────────────┼──────────────────────┼──────────────┘
          ▼                  ▼                      ▼
   ┌─────────────┐  ┌──────────────┐  ┌──────────────────────┐
   │ Office Apps  │  │ SharePoint/  │  │ Endpoint DLP         │
   │ (Word, Excel,│  │ OneDrive     │  │ (Windows devices)    │
   │  Outlook)    │  │              │  │                      │
   └─────────────┘  └──────────────┘  └──────────────────────┘
Auto-Labelling Configuration
Policy 1: Confidential — Customer PII Detection
SettingValue
Policy NameVFS-AutoLabel-Confidential-CustomerPII
ScopeAll SharePoint sites, OneDrive accounts, Exchange mailboxes
ConditionsContent contains ANY of: UK National Insurance Number (HIGH confidence), IBAN (HIGH), Vanguard Account Number (HIGH), Credit Card (HIGH + Luhn validated)
Minimum count1 instance of any SIT at HIGH confidence
Label appliedConfidential - Customer Data
Priority2 (overridden by Restricted auto-label)
Policy 2: Restricted — Special Category Detection
SettingValue
Policy NameVFS-AutoLabel-Restricted-SpecialCategory
ScopeAll SharePoint sites, OneDrive accounts, Exchange mailboxes
ConditionsContent contains ANY of: ICD-10 codes (MEDIUM+), health terminology trainable classifier (HIGH), biometric template format detection, genetic marker patterns
Minimum count1 instance at MEDIUM confidence or above
Label appliedRestricted - Special Category (Art. 9)
Priority1 (highest priority — overrides all other auto-labels)
Policy 3: Restricted — Criminal Data Detection
SettingValue
Policy NameVFS-AutoLabel-Restricted-CriminalData
ScopeHR SharePoint sites, Compliance SharePoint sites
ConditionsContent contains: DBS reference patterns, criminal conviction terminology, SAR reference numbers
Label appliedRestricted - Criminal Data (Art. 10)
Priority1
Label Inheritance Rules
RuleDescriptionImplementation
Container inheritanceItems in a labelled SharePoint site inherit the site's label as minimumSite sensitivity label propagates to new items; existing items retain higher label
Email attachment inheritanceAttachments inherit the email's label if attachment label is lowerOutlook plugin checks attachment label vs email label on send
Parent-child inheritanceChild documents inherit parent folder label as minimumSharePoint library policy; items cannot be labelled below folder label
No downgrade without approvalUsers cannot remove or downgrade labels without justificationLabel policy: require justification text for downgrade; audit log entry
Highest label winsWhen documents are merged or combined, the highest label appliesUser training + DLP monitoring for combined documents

User-Applied Labelling

Labelling Responsibilities
ScenarioWho LabelsHow
New document creationAuthorSelect label from Office ribbon (Word, Excel, PowerPoint)
New email compositionSenderSelect label from Outlook toolbar; mandatory before sending external
File upload to SharePointUploaderLabel prompt on upload if no label detected
Data export from systemExporterLabel selection required before export completes
Physical document printingPrinterClassification header/footer printed automatically; user selects tier if not auto-labelled
Show full SKILL.md (413 more words)Show less
Mandatory Labelling Policy
SettingValue
Require label on documentsYes — all Word, Excel, PowerPoint documents must have a label before save
Require label on emailsYes — for emails to external recipients; recommended for internal
Default labelInternal (applied if user does not select; user can override up or down)
Justification for downgradeRequired — user must enter text justification; logged in audit
Justification for removalRequired — DPO-approved exception only

DLP Integration

DLP Policy Matrix
LabelExternal EmailUSB/RemovablePrintScreenshotCloud Upload
PublicAllowAllowAllowAllowAllow
InternalWarnWarnAllowAllowBlock (non-approved cloud)
ConfidentialWarn + auditBlockSecure printAllow (watermarked)Block
RestrictedBlockBlockBlock (unless DPO approved)BlockBlock
DLP Alert Routing
Alert SeverityLabel TriggerRouting
LowInternal label — external email warn overriddenSecurity team email
MediumConfidential label — external sharing attemptedSecurity team + Data Owner
HighRestricted label — any policy triggerDPO + CISO + immediate investigation
CriticalRestricted label — data exfiltration indicatorsDPO + CISO + Incident Response Team + 15-minute SLA

Label Propagation Across Platforms

Supported Platforms
PlatformLabel MethodPropagation
Microsoft 365 (Word, Excel, PowerPoint)Native sensitivity label in file metadataFull support — label travels with file
Outlook / ExchangeEmail header X-MS-Exchange-Organization-ClassificationLabel persists in message store and forwarded copies
SharePoint OnlineDocument library metadata + file metadataDual storage — library and file level
OneDrive for BusinessFile metadataSame as SharePoint
TeamsChannel/chat message metadataLimited — file attachments inherit, message labels in compliance
PDF exportVisual marking (header/footer) + XMP metadataVisual marking persists; metadata depends on PDF viewer
Azure SQL DatabaseColumn-level sensitivity classification (sys.sensitivity_classifications)Native Azure SQL feature; integrates with Purview
AWS S3Object tags (key: classification, value: tier)AWS-native tagging; read by Macie and IAM policies
On-premises file sharesFile metadata (NTFS ADS) via AIP Unified Labelling clientRequires AIP client installed on endpoints

Enforcement Precedents

  • ICO v Interserve Group (2022): GBP 4.4 million — failure to implement adequate data classification and labelling contributed to staff not recognising the sensitivity of data compromised in a breach
  • CNIL v Free Mobile (2022): EUR 300,000 — customer personal data stored without classification or access controls; staff could access all customer data regardless of role or need

Integration Points

  • classification-policy: Labelling system is the technical implementation of the classification policy
  • auto-data-discovery: Discovery results trigger auto-labelling for newly detected PII
  • pii-in-unstructured: PII detection in documents drives label recommendation or auto-application
  • data-inventory-mapping: Labels feed the data inventory with current classification status per asset
  • data-lineage-tracking: Labels propagate through data lineage — transformed data inherits source label

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/data-labeling-system of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Data Labeling System next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Labeling System compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Labeling System this skillmukul975/Privacy-Data-Protection-Skills295—~2.8kAutomated safety check: PassApache-2.0
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 10 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from mukul975/Privacy-Data-Protection-Skills

All 278 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    295 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    295 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    295 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed
  • Retention Schedule

    mukul975/Privacy-Data-Protection-Skills

    Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.

    295 GitHub stars~3.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Data Labeling System

What does Data Labeling System do?

Implements data classification labels and tagging systems including metadata tagging, DLP integration, automated label propagation, user-applied labels, and label inheritance rules. Data Labeling System is an agent skill from mukul975/Privacy-Data-Protection-Skills. Implements data classification labels and tagging systems including metadata tagging, DLP integration, automated label propagation, user-applied labels, and label inheritance rules.

When should I use Data Labeling System?

Data Labeling System fits situations like: tasks that involve Data cleaning; tasks that involve Privacy and GDPR.

How do I install Data Labeling System in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-labeling-system -a claude-code`. Or copy the skill folder (skills/privacy/data-labeling-system in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/data-labeling-system in your project. Claude Code loads it when a task matches its description.

How do I install Data Labeling System in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-labeling-system -a codex`. Or copy the skill folder (skills/privacy/data-labeling-system in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/data-labeling-system in your project. Codex loads it when a task matches its description.

Can I use Data Labeling System in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-labeling-system -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-labeling-system, .gemini/skills/data-labeling-system, .github/skills/data-labeling-system and .opencode/skills/data-labeling-system in your project.

What does Data Labeling System need to run?

Going by SKILL.md and its folder, Data Labeling System needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Data Labeling System access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Labeling System safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Labeling System use?

Data Labeling System is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Labeling System use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Data Labeling System?

Skills that share tags, products or a category with Data Labeling System: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Labeling System?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.