Agent skill

Auto Data Discovery

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Implements automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie.

Apache-2.0Auto-check passedLegal & Compliance

Install Auto Data Discovery

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill auto-data-discovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills auto-data-discovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/auto-data-discovery .claude/skills/auto-data-discovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
auto-data-discovery
GitHub stars
295
Token cost
~3.3k tokens
SKILL.md length
1,272 words
Files
5 (incl. scripts, references, assets)
Skills in repo
278
Repo updated
First seen
Licence
Apache-2.0

At a glance

Implements automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie.

  • Tasks that involve Privacy and GDPR
  • SKILL.md covers Overview, Platform Comparison, Implementation Pattern —… and Implementation Pattern — BigID, plus 5 more sections
  • Runs Python scripts from its folder; calls aws
  • Tasks that involve Third-party API integration

What it does

Auto Data Discovery is an agent skill from mukul975/Privacy-Data-Protection-Skills. Implements automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie. Covers scanning schedules, accuracy tuning, false positive management, and integration patterns. Keywords: data discovery, PII scanning, Purview, BigID, Macie, OneTrust, automated classification, data cataloging.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Legal & Compliance, covering Privacy and GDPR and Third-party API integration. It works with Amazon Web Services. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Privacy and GDPR
  • Tasks that involve Third-party API integration

Example prompts

  • “Use the auto-data-discovery skill to implement automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust…”
  • “/auto-data-discovery”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Auto Data Discovery loads about 3.3k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 1,272 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 1,272 words, ~3,346 tokens.

Download SKILL.mdSave it as .claude/skills/auto-data-discovery/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
auto-data-discovery
description
Implements automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie. Covers scanning schedules, accuracy tuning, false positive management, and integration patterns. Keywords: data discovery, PII scanning, Purview, BigID, Macie, OneTrust, automated classification, data cataloging.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
data-classification
metadata.tags
data-discovery, pii-scanning, purview, bigid, macie, onetrust

Automated PII Discovery and Classification

Overview

Automated data discovery tools scan structured and unstructured data repositories to identify, classify, and catalogue personal data across the enterprise. Manual data inventories cannot keep pace with the volume, velocity, and variety of modern data processing. Automated discovery provides continuous visibility into where personal data resides, how it flows, and whether it is classified and protected according to policy. This skill covers implementation patterns for four leading platforms — Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie — with focus on scanning configuration, accuracy optimisation, and integration with privacy compliance workflows.

Platform Comparison

CapabilityMicrosoft PurviewBigIDOneTrust DataDiscoveryAWS Macie
Structured data scanningSQL Server, Azure SQL, Synapse, Cosmos DB, Oracle, PostgreSQL, MySQL, Teradata100+ connectors including all major RDBMS, NoSQL, data warehouses200+ connectors, pre-built integrations with SaaS applicationsS3, DynamoDB, RDS (via Lambda)
Unstructured data scanningSharePoint, OneDrive, Exchange, Azure Blob, Azure Files, AWS S3, GCP StorageFile shares, email, SharePoint, cloud storage, Slack, Teams, ConfluenceFile shares, email, cloud storage, collaboration platformsS3 buckets (primary focus)
Classification method300+ built-in sensitive information types (SITs), trainable classifiers, exact data match (EDM), custom regexML-based NER, correlation analysis, pattern matching, custom classifiersPattern matching, NER, contextual analysis, custom rulesML-based pattern matching, custom data identifiers, managed data identifiers
GDPR-specific classifiersEU national ID formats, EU passport numbers, EU debit/credit card numbers, EU tax ID numbers per Member StateGDPR personal data taxonomy, Art. 9 special category detection, cross-regulation mappingPre-built GDPR data subject types, purpose mapping, lawful basis taggingEU personal data identifiers (limited — primarily financial and identity patterns)
Accuracy tuningConfidence levels (low/medium/high), custom keyword dictionaries, EDM for exact matching, document fingerprintingML model retraining, feedback loop, confidence thresholds, correlation rulesConfidence scoring, validation rules, exception managementCustom data identifiers with regex and keyword proximity, severity scoring
Deployment modelSaaS (Microsoft 365/Azure), hybrid with Purview governanceSaaS, on-premises, hybridSaaS, on-premises agentAWS-native SaaS
Pricing modelPer information protection unit (Azure), per Microsoft 365 licence tier (E5 includes advanced)Per data source connector, per TB scannedPer data source module, per connectorPer S3 bucket evaluated, per GB scanned

Implementation Pattern — Microsoft Purview

Architecture
Data Sources                    Microsoft Purview
┌─────────────┐                ┌───────────────────────────┐
│ Azure SQL   │──────scanner──►│ Data Map (metadata store)  │
│ SharePoint  │──────scanner──►│ Data Catalog (search/tag)  │
│ AWS S3      │──────scanner──►│ Data Estate Insights       │
│ On-prem SQL │──self-hosted──►│ Information Protection     │
│ Power BI    │──────scanner──►│ Data Loss Prevention (DLP) │
└─────────────┘                └───────────────────────────┘
Configuration Steps for Vanguard Financial Services

Step 1: Register Data Sources

  • Register all Azure SQL databases, SharePoint sites, Azure Blob Storage accounts
  • Install self-hosted integration runtime for on-premises SQL Server and file shares
  • Register AWS S3 buckets via cross-cloud connector
  • Register Power BI workspaces for report-level sensitivity scanning

Step 2: Configure Scanning Rules

  • Enable built-in sensitive information types (SITs) for all EU jurisdictions:
    • UK National Insurance Number (pattern: [A-CEGHJ-PR-TW-Z]{2}\d{6}[A-D])
    • EU Passport Numbers (per Member State format)
    • International Banking Account Number (IBAN)
    • EU Tax Identification Numbers
    • EU Health Insurance Card Number
  • Create custom SITs for Vanguard-specific identifiers:
    • Customer account number format: VFS-\d{10}
    • Employee ID format: EMP-[A-Z]{2}\d{6}
    • Investment portfolio reference: PF-\d{8}-[A-Z]{2}
  • Configure confidence levels: HIGH for direct identifiers, MEDIUM for indirect, LOW for contextual matches
  • Enable exact data match (EDM) for customer name verification against CRM database

Step 3: Set Scanning Schedule

  • Full scan: Monthly (first Saturday of each month, 02:00 UTC)
  • Incremental scan: Weekly (every Saturday, 02:00 UTC)
  • On-demand scan: Triggered by new data source registration or incident response
  • Estimated scan duration: 8 hours for full scan across 15 TB of data

Step 4: Configure Sensitivity Labels

  • Map Purview sensitivity labels to Vanguard classification tiers:
    • Public → Purview label: Public
    • Internal → Purview label: General
    • Confidential → Purview label: Confidential
    • Restricted → Purview label: Highly Confidential (auto-applied to Art. 9/Art. 10 data)
  • Enable auto-labelling policies: Apply labels based on detected SITs
  • Configure label inheritance: Child items inherit parent container label unless higher

Step 5: DLP Policy Integration

  • Create DLP policies triggered by sensitivity labels:
    • Restricted data: Block external sharing, require encryption, notify DPO
    • Confidential data: Warn on external sharing, audit log, require justification
    • Internal data: Audit log only
  • Configure endpoint DLP for Windows devices (prevent copy to USB, print, screenshot for Restricted content)
Accuracy Tuning for Purview
IssueTuning Approach
False positive: UK phone numbers flagged as National Insurance numbersIncrease minimum confidence to HIGH for NINO SIT; add negative keyword list ("phone", "tel", "fax", "mobile")
False positive: Internal reference numbers flagged as account numbersCreate EDM schema for actual customer accounts; custom SIT with proximity to customer-related keywords
False negative: Health data in free-text email bodiesEnable trainable classifier for health content; train on sample of 50+ positive examples from occupational health correspondence
False negative: Genetic identifiers in research datasetsCreate custom SIT for rs-number pattern (rs\d{4,12}), ICD-10 codes, and HUGO gene names

Implementation Pattern — BigID

Architecture

BigID uses a distributed scanning architecture with correlation-based discovery:

Data Sources              BigID Platform
┌─────────────┐          ┌──────────────────────────┐
│ Databases   │──scan───►│ Discovery Engine          │
│ File Shares │──scan───►│ Correlation Engine (ML)   │
│ Cloud       │──scan───►│ Classification Engine     │
│ SaaS Apps   │──API────►│ Catalog & Inventory       │
│ Email       │──scan───►│ Privacy Rights Automation │
└─────────────┘          └──────────────────────────┘
Show full SKILL.md (509 more words)Show less
Key Differentiator: Correlation-Based Identity Resolution

BigID's ML-based correlation engine identifies personal data by correlating data elements across sources to build identity profiles. This approach detects personal data that pattern matching alone would miss — for example, a customer ID in one system linked to a name in another.

Configuration for Vanguard
  • Data source connections: Configure connectors for Salesforce CRM, Workday HR, ADP Payroll, SharePoint, Azure SQL, AWS S3, on-premises Oracle Data Warehouse
  • Correlation rules: Enable identity correlation across CRM customer ID, HR employee ID, and payroll employee number
  • Classification policies: Map BigID classifiers to GDPR categories (personal data, special category, criminal data, pseudonymised, anonymised)
  • Scanning schedule: Incremental scans daily at 01:00 UTC; full correlation rebuild weekly

Implementation Pattern — OneTrust DataDiscovery

Architecture

OneTrust integrates discovery with its broader privacy management platform:

Data Sources              OneTrust Platform
┌─────────────┐          ┌──────────────────────────┐
│ Cloud/SaaS  │──API────►│ DataDiscovery Module      │
│ Databases   │──agent──►│ Data Mapping (Art. 30)    │
│ File Shares │──agent──►│ Assessment Automation     │
│ Endpoints   │──agent──►│ Consent Management        │
└─────────────┘          │ DSAR Automation           │
                         └──────────────────────────┘
Key Differentiator: Privacy Program Integration

OneTrust's value proposition is tight integration between discovery results and privacy program management — discovered personal data feeds directly into Art. 30 records, DPIA assessments, and DSAR fulfilment workflows.

Configuration for Vanguard
  • Data source inventory: Use OneTrust's data mapping questionnaires to establish baseline, then deploy automated scanning to validate and enrich
  • Integration with RoPA: Map discovered data categories to Art. 30 processing activity records
  • DSAR integration: Configure discovery results to support automated data subject access request fulfilment (locate all data for a given identity across systems)
  • Vendor discovery: Scan for personal data shared with third-party processors; reconcile with vendor management records

Implementation Pattern — AWS Macie

Architecture

Macie is purpose-built for S3 data discovery within AWS:

AWS Environment
┌─────────────────────────────────────┐
│  S3 Buckets    ──scan──►  Macie     │
│                          │          │
│  EventBridge  ◄──alerts──┘          │
│  Security Hub ◄──findings──┘        │
│  CloudWatch   ◄──metrics──┘         │
└─────────────────────────────────────┘
Configuration for Vanguard
  • Bucket selection: Enable Macie for all S3 buckets in eu-west-1 (London) and eu-west-2 (Ireland) regions
  • Custom data identifiers: Create identifiers for Vanguard-specific patterns (customer account numbers, employee IDs, portfolio references)
  • Managed identifiers: Enable all EU-region managed data identifiers (UK NINO, IBAN, EU passport)
  • Job scheduling: Scheduled classification jobs weekly; event-driven scan on new object upload for high-sensitivity buckets
  • Findings integration: Route findings to AWS Security Hub for centralised monitoring; trigger Lambda function for automated label application

Scanning Schedule Best Practices

Scan TypeFrequencyDuration WindowTrigger
Full discovery scanMonthlyWeekend maintenance window (8-12 hours)Scheduled
Incremental scanWeeklyOff-peak hours (2-4 hours)Scheduled
New source onboarding scanOn registrationWithin 48 hours of source registrationEvent-driven
Post-incident scanAs neededImmediate (targeted scope)Incident response
Pre-DPIA scanBefore DPIA commencement1-2 weeks before DPIA startProject-triggered

Accuracy Metrics and Tuning

Key Performance Indicators
MetricTargetMeasurement Method
Precision (true positive rate)> 90%Sample 100 classified items monthly; verify classification accuracy
Recall (detection rate)> 85%Plant known PII test data in scan scope; measure detection rate
False positive rate< 10%Count items classified as personal data that are not
False negative rate< 15%Count personal data items missed by the scanner
Classification consistency> 95%Same data element classified consistently across repeat scans
Continuous Improvement Cycle
Month 1: Baseline scan → establish initial accuracy metrics
Month 2: Review false positives/negatives → tune rules and thresholds
Month 3: Re-scan → measure improvement
Month 4: Expand scope (new data sources) → re-baseline
Month 5: Review edge cases → create custom classifiers
Month 6: Accuracy audit by DPO → formal accuracy report
[Repeat cycle]

Integration Points

  • data-inventory-mapping: Automated discovery feeds Art. 30 data inventory
  • pii-in-unstructured: Specialized techniques for unstructured data augment platform capabilities
  • data-labeling-system: Discovery results drive automated label application
  • classification-policy: Discovery validates compliance with classification policy tiers
  • data-lineage-tracking: Discovery combined with lineage provides complete data visibility

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/auto-data-discovery of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Auto Data Discovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Auto Data Discovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Auto Data Discovery this skillmukul975/Privacy-Data-Protection-Skills295—~3.3kAutomated safety check: PassApache-2.0
Eks Securityaws-samples/appmod-blueprints113—~4.7kAutomated safety check: PassMIT-0
Eu Data Act Oliver Schmidt Prietzlawve-ai/awesome-legal-skills836—~3.9kAutomated safety check: PassAGPL-3.0
Healthcare Phi Complianceaffaan-m/ECC275k1 repos~1.4kAutomated safety check: PassMIT
Storage S3 Resiliency Expertiseaws/tools-for-devops-agent100—~2.8kAutomated safety check: PassApache-2.0
Amazon Location Serviceawslabs/agent-plugins915—~4kAutomated safety check: PassApache-2.0

Similar skills

  • Eks Security

    aws-samples/appmod-blueprints

    Official

    A skill your agent uses whenever someone needs security or compliance guidance for Amazon EKS — phrased as "CIS Benchmark for EKS", "HIPAA / PCI-DSS / FedRAMP / SOC 2 / GDPR on EKS", "harden my EKS…

    113 GitHub stars~4.7k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Eu Data Act Oliver Schmidt Prietz

    lawve-ai/awesome-legal-skills

    Practitioner skill for advising on EU Regulation 2023/2854 (Data Act).

    836 GitHub stars~3.9k tokensUpdated 5 days ago
    Legal & ComplianceAuto-check passed
  • Protected Health Information (PHI) and PII compliance patterns for healthcare applications: data classification, row-level access control, tamper-proof audit trails, schema tagging, and common leak…

    275k GitHub starsUsed in 1 repo~1.4k tokens
    Legal & ComplianceAuto-check passed
  • Storage S3 Resiliency Expertise

    aws/tools-for-devops-agent

    Official

    S3 resiliency, security, and data protection review. An agent skill from aws/tools-for-devops-agent.

    100 GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Amazon Location Service

    awslabs/agent-plugins

    Official

    Integrates Amazon Location Service APIs for AWS applications.

    915 GitHub stars~4k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Apollo Data Handling

    jeremylongshore/tons-of-skills-marketplace

    Apollo.io data management and compliance. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2.3k tokensUpdated today
    Legal & ComplianceAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 278 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    295 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    295 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    295 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed
  • Retention Schedule

    mukul975/Privacy-Data-Protection-Skills

    Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.

    295 GitHub stars~3.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Auto Data Discovery

What does Auto Data Discovery do?

Implements automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie. Auto Data Discovery is an agent skill from mukul975/Privacy-Data-Protection-Skills. Implements automated PII discovery and classification using tools like Microsoft Purview, BigID, OneTrust DataDiscovery, and AWS Macie.

When should I use Auto Data Discovery?

Auto Data Discovery fits situations like: tasks that involve Privacy and GDPR; tasks that involve Third-party API integration.

How do I install Auto Data Discovery in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill auto-data-discovery -a claude-code`. Or copy the skill folder (skills/privacy/auto-data-discovery in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/auto-data-discovery in your project. Claude Code loads it when a task matches its description.

How do I install Auto Data Discovery in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill auto-data-discovery -a codex`. Or copy the skill folder (skills/privacy/auto-data-discovery in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/auto-data-discovery in your project. Codex loads it when a task matches its description.

Can I use Auto Data Discovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill auto-data-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/auto-data-discovery, .gemini/skills/auto-data-discovery, .github/skills/auto-data-discovery and .opencode/skills/auto-data-discovery in your project.

What does Auto Data Discovery need to run?

Going by SKILL.md and its folder, Auto Data Discovery needs Python for the scripts in its folder and the command-line tools its instructions call (aws). Our summary lists: Python 3.

Does Auto Data Discovery access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Auto Data Discovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Auto Data Discovery use?

Auto Data Discovery is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Auto Data Discovery use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to Auto Data Discovery?

Skills that share tags, products or a category with Auto Data Discovery: Eks Security (aws-samples/appmod-blueprints, 113 stars), Eu Data Act Oliver Schmidt Prietz (lawve-ai/awesome-legal-skills, 836 stars), Healthcare Phi Compliance (affaan-m/ECC, 275k stars) and Storage S3 Resiliency Expertise (aws/tools-for-devops-agent, 100 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Auto Data Discovery?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.