Agent skill

Privacy Data Sharing

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Build privacy-preserving data sharing platforms using synthetic data generation with the SDV library, data clean rooms, secure enclaves, and utility measurement.

Apache-2.0Auto-check passedLegal & Compliance

Install Privacy Data Sharing

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill privacy-data-sharing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills privacy-data-sharing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/privacy-data-sharing .claude/skills/privacy-data-sharing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
privacy-data-sharing
GitHub stars
295
Token cost
~3.4k tokens
SKILL.md length
426 words
Files
5 (incl. scripts, references, assets)
Skills in repo
278
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build privacy-preserving data sharing platforms using synthetic data generation with the SDV library, data clean rooms, secure enclaves, and utility measurement.

  • Tasks that involve Test data and fixtures
  • SKILL.md covers Overview, Approach Selection Framework, Synthetic Data Generation with… and Data Clean Room Architecture, plus 3 more sections
  • Runs Python scripts from its folder
  • Tasks that involve Privacy and GDPR

What it does

Privacy Data Sharing is an agent skill from mukul975/Privacy-Data-Protection-Skills. Build privacy-preserving data sharing platforms using synthetic data generation with the SDV library, data clean rooms, secure enclaves, and utility measurement. Covers end-to-end architecture for sharing analytical datasets while preserving individual privacy guarantees.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Legal & Compliance, covering Test data and fixtures and Privacy and GDPR. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Test data and fixtures
  • Tasks that involve Privacy and GDPR

Example prompts

  • “/privacy-data-sharing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Privacy Data Sharing loads about 3.4k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 426 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 426 words, ~3,385 tokens.

Download SKILL.mdSave it as .claude/skills/privacy-data-sharing/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
privacy-data-sharing
description
Build privacy-preserving data sharing platforms using synthetic data generation with the SDV library, data clean rooms, secure enclaves, and utility measurement. Covers end-to-end architecture for sharing analytical datasets while preserving individual privacy guarantees.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
privacy-engineering
metadata.tags
data-sharing, synthetic-data, data-clean-rooms, secure-enclaves, sdv-library

Privacy-Preserving Data Sharing Platform

Overview

Privacy-preserving data sharing enables organizations to derive analytical value from combined datasets without exposing raw personal data. This skill covers four primary approaches: synthetic data generation, data clean rooms, secure enclaves, and federated analytics, along with utility measurement frameworks to ensure shared data remains useful.

Approach Selection Framework

ApproachPrivacy GuaranteeData UtilityComputational CostTrust Model
Synthetic DataStatistical (configurable)High for distributions, lower for edge casesMedium (training)No trust required
Data Clean RoomsContractual + technicalHigh (real data, restricted queries)Low-MediumTrusted operator
Secure Enclaves (TEE)Hardware-backed isolationVery high (real data)MediumTrust hardware vendor
Federated AnalyticsCryptographic/DPMedium-HighHigh (communication)Minimal trust
Homomorphic EncryptionCryptographicHighVery HighNo trust required
Secure Multi-Party ComputationCryptographicHighHighHonest majority

Synthetic Data Generation with SDV

Architecture
Source Data --> Statistical Profiling --> Model Training --> Synthetic Generation
                    |                        |                      |
                    v                        v                      v
            Metadata Analysis        Model Selection         Quality Assessment
            - Column types           - GaussianCopula        - Statistical tests
            - Distributions          - CTGAN                 - Privacy metrics
            - Correlations           - CopulaGAN             - Utility metrics
            - Constraints            - TVAE                  - Visual comparison
SDV Implementation
python
"""
Synthetic data generation using the Synthetic Data Vault (SDV) library.
Generates privacy-preserving synthetic datasets that maintain statistical
properties of the original data.
"""

import pandas as pd
import numpy as np
from sdv.metadata import SingleTableMetadata
from sdv.single_table import GaussianCopulaSynthesizer, CTGANSynthesizer, TVAESynthesizer
from sdv.evaluation.single_table import run_diagnostic, evaluate_quality
from sdmetrics.reports.single_table import QualityReport


def create_metadata(df: pd.DataFrame) -> SingleTableMetadata:
    """Auto-detect and create metadata for a DataFrame."""
    metadata = SingleTableMetadata()
    metadata.detect_from_dataframe(df)
    return metadata


def train_gaussian_copula(
    df: pd.DataFrame,
    metadata: SingleTableMetadata,
    enforce_min_max: bool = True
) -> GaussianCopulaSynthesizer:
    """
    Train a Gaussian Copula model for synthetic data generation.
    Best for: Datasets with mostly numerical data and linear correlations.
    """
    synthesizer = GaussianCopulaSynthesizer(
        metadata,
        enforce_min_max_values=enforce_min_max,
        numerical_distributions={
            "norm": "beta",  # Fit beta distributions for bounded numerical data
        }
    )
    synthesizer.fit(df)
    return synthesizer


def train_ctgan(
    df: pd.DataFrame,
    metadata: SingleTableMetadata,
    epochs: int = 300,
    batch_size: int = 500
) -> CTGANSynthesizer:
    """
    Train a CTGAN model for synthetic data generation.
    Best for: Complex distributions, mixed data types, mode-specific patterns.
    """
    synthesizer = CTGANSynthesizer(
        metadata,
        epochs=epochs,
        batch_size=batch_size,
        verbose=True
    )
    synthesizer.fit(df)
    return synthesizer


def train_tvae(
    df: pd.DataFrame,
    metadata: SingleTableMetadata,
    epochs: int = 300
) -> TVAESynthesizer:
    """
    Train a TVAE model for synthetic data generation.
    Best for: Datasets where CTGAN struggles, faster training than CTGAN.
    """
    synthesizer = TVAESynthesizer(
        metadata,
        epochs=epochs
    )
    synthesizer.fit(df)
    return synthesizer


def generate_synthetic_data(
    synthesizer,
    num_rows: int
) -> pd.DataFrame:
    """Generate synthetic data from a trained synthesizer."""
    return synthesizer.sample(num_rows=num_rows)


def evaluate_synthetic_quality(
    real_data: pd.DataFrame,
    synthetic_data: pd.DataFrame,
    metadata: SingleTableMetadata
) -> dict:
    """
    Evaluate the quality of synthetic data against real data.

    Returns diagnostic and quality scores.
    """
    # Run diagnostic checks
    diagnostic = run_diagnostic(
        real_data=real_data,
        synthetic_data=synthetic_data,
        metadata=metadata
    )

    # Run quality evaluation
    quality = evaluate_quality(
        real_data=real_data,
        synthetic_data=synthetic_data,
        metadata=metadata
    )

    return {
        "diagnostic_score": diagnostic.get_score(),
        "quality_score": quality.get_score(),
    }


def measure_privacy_risk(
    real_data: pd.DataFrame,
    synthetic_data: pd.DataFrame,
    metadata: SingleTableMetadata,
    key_fields: list[str]
) -> dict:
    """
    Measure re-identification risk in synthetic data.

    Checks for exact matches and nearest-neighbor distances
    between real and synthetic records.
    """
    # Check for exact record matches
    merged = real_data.merge(synthetic_data, how="inner")
    exact_match_rate = len(merged) / len(real_data)

    # Check key field matches
    if key_fields:
        key_merged = real_data[key_fields].merge(
            synthetic_data[key_fields], how="inner"
        )
        key_match_rate = len(key_merged) / len(real_data)
    else:
        key_match_rate = 0.0

    return {
        "exact_match_rate": exact_match_rate,
        "key_match_rate": key_match_rate,
        "privacy_safe": exact_match_rate < 0.01 and key_match_rate < 0.05,
    }
Model Selection Guide
FactorGaussianCopulaCTGANTVAE
Training speedFast (minutes)Slow (hours)Medium (30-60 min)
Small datasets (<1K rows)GoodPoorFair
Large datasets (>100K rows)GoodGoodGood
Numerical dataExcellentGoodGood
Categorical data (high cardinality)FairGoodGood
Complex correlationsFairGoodGood
Constraint handlingGoodFairFair
ReproducibilityExcellentFair (seed-dependent)Fair

Data Clean Room Architecture

Components
Organization A                    Clean Room                    Organization B
+-------------+    encrypted    +------------------+    encrypted    +-------------+
| Source Data  |  ----------->  | Ingestion Layer  |  <-----------  | Source Data  |
+-------------+                 +------------------+                 +-------------+
                                        |
                                        v
                                +------------------+
                                | Data Preparation |
                                | - Schema mapping |
                                | - Normalization  |
                                | - Deduplication  |
                                +------------------+
                                        |
                                        v
                                +------------------+
                                | Approved Queries |
                                | - Pre-approved   |
                                |   query templates|
                                | - Aggregate only |
                                | - Min group size |
                                +------------------+
                                        |
                                        v
                                +------------------+
                                | Output Validation|
                                | - k-anonymity    |
                                | - DP noise       |
                                | - Disclosure risk|
                                +------------------+
                                        |
                            +-----------+-----------+
                            |                       |
                            v                       v
                    Results for Org A        Results for Org B
Clean Room Policy Engine
python
"""
Policy engine for data clean room query validation.
Enforces privacy rules on all queries before execution.
"""

from dataclasses import dataclass


@dataclass
class CleanRoomPolicy:
    min_group_size: int = 50
    allowed_operations: list[str] = None
    blocked_columns: list[str] = None
    max_output_rows: int = 1000
    require_aggregation: bool = True
    dp_epsilon: float = 1.0

    def __post_init__(self):
        if self.allowed_operations is None:
            self.allowed_operations = ["COUNT", "SUM", "AVG", "MEDIAN", "PERCENTILE"]
        if self.blocked_columns is None:
            self.blocked_columns = ["ssn", "email", "phone", "full_name", "address"]


class QueryValidator:
    """Validate clean room queries against privacy policies."""

    def __init__(self, policy: CleanRoomPolicy):
        self.policy = policy

    def validate(self, query_ast: dict) -> tuple[bool, list[str]]:
        """
        Validate a parsed query against the policy.

        Returns (is_valid, list_of_violations).
        """
        violations = []

        # Check for blocked columns
        referenced_columns = query_ast.get("columns", [])
        for col in referenced_columns:
            if col.lower() in self.policy.blocked_columns:
                violations.append(f"Column '{col}' is blocked by policy")

        # Check aggregation requirement
        if self.policy.require_aggregation:
            if not query_ast.get("has_aggregation", False):
                violations.append("Query must include aggregation (no raw record output)")

        # Check operations
        operations = query_ast.get("operations", [])
        for op in operations:
            if op.upper() not in self.policy.allowed_operations:
                violations.append(f"Operation '{op}' is not in allowed operations list")

        # Check output size
        if query_ast.get("limit", float("inf")) > self.policy.max_output_rows:
            violations.append(
                f"Output exceeds max rows ({self.policy.max_output_rows})"
            )

        return (len(violations) == 0, violations)

Secure Enclave Integration

Intel SGX / Azure Confidential Computing
Data Owner A           Confidential Computing          Data Owner B
                       +----------------------+
Data (encrypted) ----> | Enclave Environment  | <---- Data (encrypted)
                       | - Decryption in TEE  |
                       | - Join/Analysis      |
                       | - Re-encrypt results |
                       +----------------------+
                               |
                       Encrypted Results
                       (only to authorized parties)
Key Properties
  • Confidentiality: Data is encrypted outside the enclave; only decrypted within the TEE
  • Integrity: Enclave code is measured and attested; tampering is detectable
  • Attestation: Remote parties can verify the enclave is running approved code
  • Isolation: Even the cloud provider cannot access data inside the enclave
Show full SKILL.md (169 more words)Show less

Utility Measurement Framework

Statistical Utility Metrics
MetricDescriptionTarget
Column ShapesDistribution similarity per column (KS test)> 0.85
Column Pair TrendsCorrelation preservation between column pairs> 0.80
Boundary AdherenceValues within real data min/max ranges> 0.95
Category CoverageAll categories in real data appear in synthetic> 0.90
Range CoverageNumeric ranges adequately represented> 0.85
Privacy Metrics
MetricDescriptionTarget
Exact Match Rate% of synthetic records identical to real records< 1%
Nearest Neighbor DistanceMinimum distance from synthetic to nearest real record> threshold
Membership Inference AUCAbility of attack model to determine membership< 0.55
Attribute Inference AccuracyAbility to infer sensitive attributes< random + 5%
k-Anonymity of outputMinimum equivalence class sizek >= 5

References

  • Patki, N., Wedge, R., and Veeramachaneni, K. "The Synthetic Data Vault." IEEE DSAA, 2016.
  • SDV Documentation: docs.sdv.dev
  • Xu, L. et al. "Modeling Tabular Data Using Conditional GAN." NeurIPS, 2019.
  • Google BigQuery Clean Rooms Documentation
  • AWS Clean Rooms Service Documentation
  • Microsoft Azure Confidential Computing Documentation
  • Stadler, T. et al. "Synthetic Data — Anonymisation Groundhog Day." USENIX Security, 2022.

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/privacy-data-sharing of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Privacy Data Sharing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Privacy Data Sharing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Privacy Data Sharing this skillmukul975/Privacy-Data-Protection-Skills295—~3.4kAutomated safety check: PassApache-2.0
Test Data Managementproffesor-for-testing/agentic-qe494—~1.2kAutomated safety check: PassMIT
Qe Test Data Managementproffesor-for-testing/agentic-qe494—~1.7kAutomated safety check: PassMIT
C15tc15t/c15t1.9k1 repos~1.6kAutomated safety check: PassApache-2.0
HIPAA Safe Harbor Coverage Auditmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
Korean Privacy Termskimlawtech/korean-privacy-terms586—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Test Data Management

    proffesor-for-testing/agentic-qe

    Strategic test data generation, management, and privacy compliance.

    494 GitHub stars~1.2k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Qe Test Data Management

    proffesor-for-testing/agentic-qe

    Strategic test data generation, management, and privacy compliance.

    494 GitHub stars~1.7k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Korean Privacy Terms

    kimlawtech/korean-privacy-terms

    처리방침·이용약관 자동 생성 스킬 패키지 (v4.0). An agent skill from kimlawtech/korean-privacy-terms.

    586 GitHub stars~2.9k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Gdpr Compliance

    Sushegaad/Claude-Skills-Governance-Risk-and-Compliance

    Expert GDPR compliance assistant covering all four core workflows: (1) auditing code and systems for GDPR violations, (2) drafting GDPR-compliant documents such as privacy policies, Data Processing…

    942 GitHub starsUsed in 1 repo~3.9k tokens
    Legal & ComplianceAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 278 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    295 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    295 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    295 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed
  • Retention Schedule

    mukul975/Privacy-Data-Protection-Skills

    Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.

    295 GitHub stars~3.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Privacy Data Sharing

What does Privacy Data Sharing do?

Build privacy-preserving data sharing platforms using synthetic data generation with the SDV library, data clean rooms, secure enclaves, and utility measurement. Privacy Data Sharing is an agent skill from mukul975/Privacy-Data-Protection-Skills. Build privacy-preserving data sharing platforms using synthetic data generation with the SDV library, data clean rooms, secure enclaves, and utility measurement.

When should I use Privacy Data Sharing?

Privacy Data Sharing fits situations like: tasks that involve Test data and fixtures; tasks that involve Privacy and GDPR.

How do I install Privacy Data Sharing in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill privacy-data-sharing -a claude-code`. Or copy the skill folder (skills/privacy/privacy-data-sharing in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/privacy-data-sharing in your project. Claude Code loads it when a task matches its description.

How do I install Privacy Data Sharing in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill privacy-data-sharing -a codex`. Or copy the skill folder (skills/privacy/privacy-data-sharing in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/privacy-data-sharing in your project. Codex loads it when a task matches its description.

Can I use Privacy Data Sharing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill privacy-data-sharing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/privacy-data-sharing, .gemini/skills/privacy-data-sharing, .github/skills/privacy-data-sharing and .opencode/skills/privacy-data-sharing in your project.

What does Privacy Data Sharing need to run?

Going by SKILL.md and its folder, Privacy Data Sharing needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Privacy Data Sharing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Privacy Data Sharing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Privacy Data Sharing use?

Privacy Data Sharing is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Privacy Data Sharing use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 643 tokens, read only when the agent opens those files.

What are the alternatives to Privacy Data Sharing?

Skills that share tags, products or a category with Privacy Data Sharing: Test Data Management (proffesor-for-testing/agentic-qe, 494 stars), Qe Test Data Management (proffesor-for-testing/agentic-qe, 494 stars), C15t (c15t/c15t, 1.9k stars) and HIPAA Safe Harbor Coverage Audit (maziyarpanahi/openmed, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Privacy Data Sharing?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.