Agent skill

AI Training Data Class

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Classifies sensitive data in AI/ML training datasets including bias detection for Art.

Apache-2.0Auto-check passedLegal & Compliance

Install AI Training Data Class

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-class --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/ai-training-data-class .claude/skills/ai-training-data-class && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-training-data-class
GitHub stars
295
Token cost
~2.9k tokens
SKILL.md length
1,302 words
Files
5 (incl. scripts, references, assets)
Skills in repo
278
Repo updated
First seen
Licence
Apache-2.0

At a glance

Classifies sensitive data in AI/ML training datasets including bias detection for Art.

  • Works in 4 steps: Demographic Representation Analysis → Proxy Variable Identification → Disparate Impact Analysis → …
  • Tasks that involve Privacy and GDPR
  • SKILL.md covers Overview, GDPR and AI Act Intersection, Training Data Classification… and Data Card Documentation, plus 3 more sections
  • Runs Python scripts from its folder

What it does

AI Training Data Class is an agent skill from mukul975/Privacy-Data-Protection-Skills. Classifies sensitive data in AI/ML training datasets including bias detection for Art. 9 categories, data card documentation, provenance tracking, and consent verification for model training. Keywords: AI training data, ML dataset, bias detection, data card, model training, Art 9, consent, GDPR AI.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Legal & Compliance, covering Privacy and GDPR, Fine-tuning and AI governance. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Privacy and GDPR
  • Tasks that involve Fine-tuning
  • Tasks that involve AI governance

Example prompts

  • “Use the ai-training-data-class skill to classify sensitive data in AI/ML training datasets including bias detection for Art”
  • “/ai-training-data-class”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Demographic Representation Analysis
  2. Proxy Variable Identification
  3. Disparate Impact Analysis
  4. AI Act Art. 10(5) Assessment

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Training Data Class loads about 2.9k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 1,302 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 1,302 words, ~2,856 tokens.

Download SKILL.mdSave it as .claude/skills/ai-training-data-class/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
ai-training-data-class
description
Classifies sensitive data in AI/ML training datasets including bias detection for Art. 9 categories, data card documentation, provenance tracking, and consent verification for model training. Keywords: AI training data, ML dataset, bias detection, data card, model training, Art 9, consent, GDPR AI.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
data-classification
metadata.tags
ai-training-data, ml-dataset, bias-detection, data-card, model-training, gdpr-ai

Sensitive Data Classification for AI/ML Training Datasets

Overview

AI and machine learning models trained on personal data raise distinct classification challenges. Training data may contain direct personal data, inferred special categories, proxy variables for protected characteristics, and data whose consent scope does not extend to model training. The EU AI Act (Regulation (EU) 2024/1689) imposes additional requirements for high-risk AI systems, including data governance obligations under Art. 10 that intersect with GDPR classification requirements. This skill provides a framework for classifying training data, detecting bias-relevant features, documenting data provenance, and verifying consent coverage.

GDPR and AI Act Intersection

GDPR Requirements for Training Data
GDPR ArticleApplication to AI Training
Art. 5(1)(b) — Purpose limitationTraining a model is a distinct processing purpose; if data was collected for customer service, using it for ML training requires a compatible purpose assessment or new lawful basis
Art. 5(1)(c) — Data minimisationTraining datasets must not include more personal data than necessary for the model objective
Art. 6 — Lawful basisModel training requires its own lawful basis; legitimate interests (Art. 6(1)(f)) is most common, but requires LIA documentation
Art. 9 — Special categoriesIf training data contains or enables inference of special category data, Art. 9(2) condition required
Art. 22 — Automated decision-makingIf the trained model makes decisions with legal or significant effects, additional safeguards apply
Art. 25 — Data protection by designClassification of training data is a by-design measure enabling appropriate technical protections
Art. 35 — DPIAHigh-risk AI processing (profiling, automated decision-making) requires DPIA
EU AI Act Art. 10 — Data Governance for High-Risk AI

The AI Act Art. 10 requires that training, validation, and testing datasets for high-risk AI systems:

  1. Are subject to appropriate data governance and management practices (Art. 10(2))
  2. Are relevant, sufficiently representative, and to the extent possible free of errors and complete (Art. 10(3))
  3. Take into account the specific geographical, contextual, behavioural, or functional setting within which the AI system is intended to be used (Art. 10(4))
  4. Where special category data processing is strictly necessary for bias detection and correction, this is permitted under Art. 10(5), subject to appropriate safeguards including pseudonymisation and GDPR compliance

Training Data Classification Framework

Classification Tier 1: Personal Data Content
ClassificationDescriptionExample
TRAINING_PII_DIRECTDataset contains direct identifiersCustomer names, email addresses in NLP training corpus
TRAINING_PII_INDIRECTDataset contains indirect identifiersCustomer IDs, transaction patterns enabling singling out
TRAINING_SPECIAL_CATDataset contains Art. 9 special category dataHealth records for medical diagnosis model
TRAINING_CRIMINALDataset contains Art. 10 criminal dataFraud transaction labels derived from criminal investigations
TRAINING_PSEUDONYMISEDPersonal data replaced with tokens but re-identification key existsPseudonymised customer data with mapping held by data team
TRAINING_ANONYMISEDData verified as anonymised per WP29 criteriaAggregated population statistics with k ≥ 10
TRAINING_SYNTHETICArtificially generated data with no real personal dataGAN-generated synthetic transaction data
TRAINING_NON_PERSONALNo personal data contentMarket price data, weather data, product specifications
Classification Tier 2: Bias and Proxy Detection

Even when a dataset does not directly contain Art. 9 special category data, it may contain proxy variables that correlate with protected characteristics:

Proxy VariableCorrelated Protected CharacteristicDetection Method
Postcode/ZIP codeRacial/ethnic origin, socioeconomic statusGeographic demographic analysis
First nameGender, ethnic origin, age cohortName demographics database lookup
Language preferenceEthnic origin, nationalityStatistical correlation analysis
Shopping patternsReligious belief (halal/kosher purchases), health statusPurchase category analysis
Web browsing historyPolitical opinions, sexual orientation, health statusTopic modelling on browsing categories
Employment gap patternsGender (maternity), disability, healthStatistical pattern analysis
Credit scoreRacial/ethnic origin (documented correlation in US/UK studies)Disparate impact analysis
ClassificationDescriptionCompliance Requirement
CONSENT_COVERS_TRAININGOriginal consent explicitly covers AI/ML trainingDocument consent text and verify specificity
CONSENT_DOES_NOT_COVEROriginal consent did not anticipate ML trainingNew consent required or alternative lawful basis needed
LEGITIMATE_INTERESTML training relies on legitimate interests (Art. 6(1)(f))Documented LIA required
CONTRACT_PERFORMANCEML training is necessary for contract performanceNarrow scope — must be genuinely necessary
PUBLIC_DATAData sourced from publicly available sourcesStill requires lawful basis; public availability is not a lawful basis
RESEARCH_EXEMPTIONProcessing under Art. 89(1) research exemptionAppropriate safeguards including pseudonymisation required

Data Card Documentation

A data card is a structured document accompanying each training dataset, providing transparency about its contents, provenance, and limitations. Modelled on the "Datasheets for Datasets" framework (Gebru et al., 2021) and adapted for GDPR compliance.

Show full SKILL.md (576 more words)Show less
Required Data Card Fields for Vanguard Financial Services
SectionFields
1. Dataset IdentityName, version, creation date, owner, purpose
2. Personal Data ClassificationTier 1 classification, data elements present, classification labels
3. Data SubjectsCategories of data subjects, volume, geographic scope
4. ProvenanceOriginal collection purpose, source systems, processing chain from collection to training set
5. Consent/Lawful BasisTier 3 classification, consent text reference or LIA reference, purpose compatibility assessment
6. Special Category AssessmentWhether Art. 9 data is present (direct or inferred), Art. 9(2) condition if applicable
7. Bias AssessmentProxy variables identified, disparate impact analysis results, demographic representation statistics
8. De-identificationTechnique applied (pseudonymisation, anonymisation, synthetic generation), assessment reference
9. RetentionTraining data retention period, model retention period, deletion schedule
10. Access ControlsWho can access the training data, who can access the model, audit logging
11. DPIA ReferenceDPIA document reference if applicable
12. LimitationsKnown biases, geographic limitations, temporal limitations, data quality issues

Bias Detection Methodology

Step 1: Demographic Representation Analysis

For each training dataset, calculate representation statistics:

  • What percentage of records come from each demographic group (to the extent known)?
  • Does the representation match the target population?
  • Are any groups under-represented by more than 20% relative to population?
Step 2: Proxy Variable Identification

Scan all features for proxy correlation with Art. 9 protected characteristics:

  • Calculate correlation coefficient between each feature and known protected characteristics (where available)
  • Flag features with |correlation| > 0.3 as potential proxies
  • Document all proxy variables in the data card
Step 3: Disparate Impact Analysis

For classification or scoring models:

  • Calculate model performance metrics by demographic group
  • Apply the 80% (four-fifths) rule: if the selection rate for any protected group is less than 80% of the rate for the most favoured group, disparate impact may exist
  • Document disparate impact analysis results in the data card
Step 4: AI Act Art. 10(5) Assessment

If bias detection requires processing special category data:

  • Document why processing is "strictly necessary" for bias detection and correction
  • Implement pseudonymisation of the special category data used for bias testing
  • Ensure GDPR Art. 9(2) condition is established (typically Art. 9(2)(g) substantial public interest or Art. 9(2)(j) research)
  • Process in a controlled environment with access restricted to the bias assessment team
  • Delete special category data after bias assessment is complete

Enforcement and Regulatory Precedents

  • Italian Garante — Clearview AI (2022): EUR 20 million fine for processing biometric data scraped from public sources for AI facial recognition training without lawful basis, consent, or transparency. Established that public availability does not provide lawful basis for AI training.
  • Italian Garante — ChatGPT/OpenAI (2023): Temporary ban and subsequent enforcement requiring OpenAI to establish lawful basis for training data collection, implement age verification, and provide opt-out mechanisms. Highlighted that AI training on personal data requires GDPR compliance throughout the data lifecycle.
  • CNIL — Enforcement Notice on AI Training Data (2024): CNIL published guidance sheets on AI training data requiring purpose limitation assessment, proportionality analysis, and specific measures when training data contains special category data.
  • ICO — Generative AI and Data Protection Consultation (2024): ICO's position that legitimate interests is the most likely lawful basis for AI training but requires documented LIA considering data subject reasonable expectations.

Integration Points

  • personal-data-test: Training data must first be classified as personal or non-personal
  • special-category-data: Art. 9 data in training sets requires heightened protections
  • pseudo-vs-anon-data: De-identification of training data must be validated
  • classification-policy: Training data classified under enterprise classification tiers
  • data-lineage-tracking: Full provenance from original collection to model deployment must be tracked

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/ai-training-data-class of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

AI Training Data Class next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Training Data Class compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Training Data Class this skillmukul975/Privacy-Data-Protection-Skills295—~2.9kAutomated safety check: PassApache-2.0
Compliance Testingpetrkindlmann/qa-skills165—~4.6kAutomated safety check: PassMIT
Compliance Osalirezarezvani/claude-skills28k—~3.3kAutomated safety check: PassMIT
Chief AI Officer Advisoralirezarezvani/claude-skills28k—~3.5kAutomated safety check: PassMIT
Ra Qm Skillsalirezarezvani/claude-skills28k—~833Automated safety check: PassMIT
Region Configindranilbanerjee/digital-marketing-pro8551 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    165 GitHub stars~4.6k tokensUpdated 3 mo ago
    Legal & ComplianceAuto-check passed
  • Compliance Os

    alirezarezvani/claude-skills

    Compliance OS — meta-orchestrator that lets compliance teams CONFIGURE which frameworks apply, COMPUTE cross-framework control overlap, SIMULATE internal audits, and CONSOLIDATE evidence across…

    28k GitHub stars~3.3k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Chief AI Officer Advisor

    alirezarezvani/claude-skills

    Chief AI Officer advisory for startups: model build-vs-buy decisions (API vs fine-tune vs in-house), AI risk classification under EU AI Act + US state patchwork, AI cost economics…

    28k GitHub stars~3.5k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Ra Qm Skills

    alirezarezvani/claude-skills

    Router/index for the 15 regulatory & quality-management skills bundled in this plugin (ISO 13485 QMS, EU MDR 2017/745, FDA submissions under QMSR, ISO 14971 risk, CAPA, document control, ISO…

    28k GitHub stars~833 tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Region Config

    indranilbanerjee/digital-marketing-pro

    Configure a brand's regional settings — timezone, languages, currency, compliance regulations (GDPR, CCPA, APPI, LGPD, EU AI Act Article 50, and more), local platforms, business hours, holiday…

    855 GitHub starsUsed in 1 repo~3.5k tokens
    Legal & ComplianceAuto-check passed
  • Analyzes how multiple regulations interact for a specific product, service, or business model.

    836 GitHub stars~3.1k tokensUpdated 5 days ago
    Legal & ComplianceAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 278 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    295 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    295 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    295 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed
  • Retention Schedule

    mukul975/Privacy-Data-Protection-Skills

    Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.

    295 GitHub stars~3.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about AI Training Data Class

What does AI Training Data Class do?

Classifies sensitive data in AI/ML training datasets including bias detection for Art. AI Training Data Class is an agent skill from mukul975/Privacy-Data-Protection-Skills. Classifies sensitive data in AI/ML training datasets including bias detection for Art.

When should I use AI Training Data Class?

AI Training Data Class fits situations like: tasks that involve Privacy and GDPR; tasks that involve Fine-tuning; tasks that involve AI governance.

How do I install AI Training Data Class in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a claude-code`. Or copy the skill folder (skills/privacy/ai-training-data-class in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/ai-training-data-class in your project. Claude Code loads it when a task matches its description.

How do I install AI Training Data Class in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a codex`. Or copy the skill folder (skills/privacy/ai-training-data-class in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/ai-training-data-class in your project. Codex loads it when a task matches its description.

Can I use AI Training Data Class in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-training-data-class, .gemini/skills/ai-training-data-class, .github/skills/ai-training-data-class and .opencode/skills/ai-training-data-class in your project.

What does AI Training Data Class need to run?

Going by SKILL.md and its folder, AI Training Data Class needs Python for the scripts in its folder. Our summary lists: Python 3.

Does AI Training Data Class access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Training Data Class safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does AI Training Data Class use?

AI Training Data Class is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Training Data Class use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to AI Training Data Class?

Skills that share tags, products or a category with AI Training Data Class: Compliance Testing (petrkindlmann/qa-skills, 165 stars), Compliance Os (alirezarezvani/claude-skills, 28k stars), Chief AI Officer Advisor (alirezarezvani/claude-skills, 28k stars) and Ra Qm Skills (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Training Data Class?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.