Compliance Testing
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Classifies sensitive data in AI/ML training datasets including bias detection for Art.
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-class --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/ai-training-data-class .claude/skills/ai-training-data-class && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-training-data-class" agent skill from https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-class into .claude/skills/ai-training-data-class/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-training-data-class", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-classType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-class --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/privacy/ai-training-data-class .agents/skills/ai-training-data-class && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-training-data-class" agent skill from https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-class into .agents/skills/ai-training-data-class/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-training-data-class", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-class --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/privacy/ai-training-data-class .cursor/skills/ai-training-data-class && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-training-data-class" agent skill from https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-class into .cursor/skills/ai-training-data-class/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-training-data-class", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mukul975/Privacy-Data-Protection-Skills.git --path skills/privacy/ai-training-data-class--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-class --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/privacy/ai-training-data-class .gemini/skills/ai-training-data-class && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-training-data-class" agent skill from https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-class into .gemini/skills/ai-training-data-class/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-training-data-class", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-classInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/privacy/ai-training-data-class .github/skills/ai-training-data-class && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-training-data-class" agent skill from https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-class into .github/skills/ai-training-data-class/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-training-data-class", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mukul975/Privacy-Data-Protection-Skills ai-training-data-class --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/privacy/ai-training-data-class .opencode/skills/ai-training-data-class && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-training-data-class" agent skill from https://github.com/mukul975/Privacy-Data-Protection-Skills/tree/main/skills/privacy/ai-training-data-class into .opencode/skills/ai-training-data-class/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-training-data-class", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-training-data-classClassifies sensitive data in AI/ML training datasets including bias detection for Art.
AI Training Data Class is an agent skill from mukul975/Privacy-Data-Protection-Skills. Classifies sensitive data in AI/ML training datasets including bias detection for Art. 9 categories, data card documentation, provenance tracking, and consent verification for model training. Keywords: AI training data, ML dataset, bias detection, data card, model training, Art 9, consent, GDPR AI.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).
It sits in Legal & Compliance, covering Privacy and GDPR, Fine-tuning and AI governance. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Training Data Class loads about 2.9k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 1,302 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 1,302 words, ~2,856 tokens.
.claude/skills/ai-training-data-class/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.AI and machine learning models trained on personal data raise distinct classification challenges. Training data may contain direct personal data, inferred special categories, proxy variables for protected characteristics, and data whose consent scope does not extend to model training. The EU AI Act (Regulation (EU) 2024/1689) imposes additional requirements for high-risk AI systems, including data governance obligations under Art. 10 that intersect with GDPR classification requirements. This skill provides a framework for classifying training data, detecting bias-relevant features, documenting data provenance, and verifying consent coverage.
| GDPR Article | Application to AI Training |
|---|---|
| Art. 5(1)(b) — Purpose limitation | Training a model is a distinct processing purpose; if data was collected for customer service, using it for ML training requires a compatible purpose assessment or new lawful basis |
| Art. 5(1)(c) — Data minimisation | Training datasets must not include more personal data than necessary for the model objective |
| Art. 6 — Lawful basis | Model training requires its own lawful basis; legitimate interests (Art. 6(1)(f)) is most common, but requires LIA documentation |
| Art. 9 — Special categories | If training data contains or enables inference of special category data, Art. 9(2) condition required |
| Art. 22 — Automated decision-making | If the trained model makes decisions with legal or significant effects, additional safeguards apply |
| Art. 25 — Data protection by design | Classification of training data is a by-design measure enabling appropriate technical protections |
| Art. 35 — DPIA | High-risk AI processing (profiling, automated decision-making) requires DPIA |
The AI Act Art. 10 requires that training, validation, and testing datasets for high-risk AI systems:
| Classification | Description | Example |
|---|---|---|
| TRAINING_PII_DIRECT | Dataset contains direct identifiers | Customer names, email addresses in NLP training corpus |
| TRAINING_PII_INDIRECT | Dataset contains indirect identifiers | Customer IDs, transaction patterns enabling singling out |
| TRAINING_SPECIAL_CAT | Dataset contains Art. 9 special category data | Health records for medical diagnosis model |
| TRAINING_CRIMINAL | Dataset contains Art. 10 criminal data | Fraud transaction labels derived from criminal investigations |
| TRAINING_PSEUDONYMISED | Personal data replaced with tokens but re-identification key exists | Pseudonymised customer data with mapping held by data team |
| TRAINING_ANONYMISED | Data verified as anonymised per WP29 criteria | Aggregated population statistics with k ≥ 10 |
| TRAINING_SYNTHETIC | Artificially generated data with no real personal data | GAN-generated synthetic transaction data |
| TRAINING_NON_PERSONAL | No personal data content | Market price data, weather data, product specifications |
Even when a dataset does not directly contain Art. 9 special category data, it may contain proxy variables that correlate with protected characteristics:
| Proxy Variable | Correlated Protected Characteristic | Detection Method |
|---|---|---|
| Postcode/ZIP code | Racial/ethnic origin, socioeconomic status | Geographic demographic analysis |
| First name | Gender, ethnic origin, age cohort | Name demographics database lookup |
| Language preference | Ethnic origin, nationality | Statistical correlation analysis |
| Shopping patterns | Religious belief (halal/kosher purchases), health status | Purchase category analysis |
| Web browsing history | Political opinions, sexual orientation, health status | Topic modelling on browsing categories |
| Employment gap patterns | Gender (maternity), disability, health | Statistical pattern analysis |
| Credit score | Racial/ethnic origin (documented correlation in US/UK studies) | Disparate impact analysis |
| Classification | Description | Compliance Requirement |
|---|---|---|
| CONSENT_COVERS_TRAINING | Original consent explicitly covers AI/ML training | Document consent text and verify specificity |
| CONSENT_DOES_NOT_COVER | Original consent did not anticipate ML training | New consent required or alternative lawful basis needed |
| LEGITIMATE_INTEREST | ML training relies on legitimate interests (Art. 6(1)(f)) | Documented LIA required |
| CONTRACT_PERFORMANCE | ML training is necessary for contract performance | Narrow scope — must be genuinely necessary |
| PUBLIC_DATA | Data sourced from publicly available sources | Still requires lawful basis; public availability is not a lawful basis |
| RESEARCH_EXEMPTION | Processing under Art. 89(1) research exemption | Appropriate safeguards including pseudonymisation required |
A data card is a structured document accompanying each training dataset, providing transparency about its contents, provenance, and limitations. Modelled on the "Datasheets for Datasets" framework (Gebru et al., 2021) and adapted for GDPR compliance.
| Section | Fields |
|---|---|
| 1. Dataset Identity | Name, version, creation date, owner, purpose |
| 2. Personal Data Classification | Tier 1 classification, data elements present, classification labels |
| 3. Data Subjects | Categories of data subjects, volume, geographic scope |
| 4. Provenance | Original collection purpose, source systems, processing chain from collection to training set |
| 5. Consent/Lawful Basis | Tier 3 classification, consent text reference or LIA reference, purpose compatibility assessment |
| 6. Special Category Assessment | Whether Art. 9 data is present (direct or inferred), Art. 9(2) condition if applicable |
| 7. Bias Assessment | Proxy variables identified, disparate impact analysis results, demographic representation statistics |
| 8. De-identification | Technique applied (pseudonymisation, anonymisation, synthetic generation), assessment reference |
| 9. Retention | Training data retention period, model retention period, deletion schedule |
| 10. Access Controls | Who can access the training data, who can access the model, audit logging |
| 11. DPIA Reference | DPIA document reference if applicable |
| 12. Limitations | Known biases, geographic limitations, temporal limitations, data quality issues |
For each training dataset, calculate representation statistics:
Scan all features for proxy correlation with Art. 9 protected characteristics:
For classification or scoring models:
If bias detection requires processing special category data:
© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/ai-training-data-class of mukul975/Privacy-Data-Protection-Skills.
Open the folder on GitHubat commit 9b2ef9e
AI Training Data Class next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Training Data Class this skillmukul975/Privacy-Data-Protection-Skills | 295 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Compliance Testingpetrkindlmann/qa-skills | 165 | — | ~4.6k | Automated safety check: Pass | MIT | |
| Compliance Osalirezarezvani/claude-skills | 28k | — | ~3.3k | Automated safety check: Pass | MIT | |
| Chief AI Officer Advisoralirezarezvani/claude-skills | 28k | — | ~3.5k | Automated safety check: Pass | MIT | |
| Ra Qm Skillsalirezarezvani/claude-skills | 28k | — | ~833 | Automated safety check: Pass | MIT | |
| Region Configindranilbanerjee/digital-marketing-pro | 855 | 1 repos | ~3.5k | Automated safety check: Pass | MIT |
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
alirezarezvani/claude-skills
Compliance OS — meta-orchestrator that lets compliance teams CONFIGURE which frameworks apply, COMPUTE cross-framework control overlap, SIMULATE internal audits, and CONSOLIDATE evidence across…
alirezarezvani/claude-skills
Chief AI Officer advisory for startups: model build-vs-buy decisions (API vs fine-tune vs in-house), AI risk classification under EU AI Act + US state patchwork, AI cost economics…
alirezarezvani/claude-skills
Router/index for the 15 regulatory & quality-management skills bundled in this plugin (ISO 13485 QMS, EU MDR 2017/745, FDA submissions under QMSR, ISO 14971 risk, CAPA, document control, ISO…
indranilbanerjee/digital-marketing-pro
Configure a brand's regional settings — timezone, languages, currency, compliance regulations (GDPR, CCPA, APPI, LGPD, EU AI Act Article 50, and more), local platforms, business hours, holiday…
lawve-ai/awesome-legal-skills
Analyzes how multiple regulations interact for a specific product, service, or business model.
mukul975/Privacy-Data-Protection-Skills
Implements age-gating mechanisms for online services to restrict access based on user age.
mukul975/Privacy-Data-Protection-Skills
Manages AI model retention and machine unlearning requirements.
mukul975/Privacy-Data-Protection-Skills
Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).
mukul975/Privacy-Data-Protection-Skills
Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.
mukul975/Privacy-Data-Protection-Skills
Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.
mukul975/Privacy-Data-Protection-Skills
Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.
Categories
Classifies sensitive data in AI/ML training datasets including bias detection for Art. AI Training Data Class is an agent skill from mukul975/Privacy-Data-Protection-Skills. Classifies sensitive data in AI/ML training datasets including bias detection for Art.
AI Training Data Class fits situations like: tasks that involve Privacy and GDPR; tasks that involve Fine-tuning; tasks that involve AI governance.
Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a claude-code`. Or copy the skill folder (skills/privacy/ai-training-data-class in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/ai-training-data-class in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a codex`. Or copy the skill folder (skills/privacy/ai-training-data-class in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/ai-training-data-class in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-data-class -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-training-data-class, .gemini/skills/ai-training-data-class, .github/skills/ai-training-data-class and .opencode/skills/ai-training-data-class in your project.
Going by SKILL.md and its folder, AI Training Data Class needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
AI Training Data Class is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with AI Training Data Class: Compliance Testing (petrkindlmann/qa-skills, 165 stars), Compliance Os (alirezarezvani/claude-skills, 28k stars), Chief AI Officer Advisor (alirezarezvani/claude-skills, 28k stars) and Ra Qm Skills (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.
Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.