Agent skill

Designing Privacy Preserving Analytics

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Design privacy-preserving analytics systems using differential privacy, k-anonymity, l-diversity, and t-closeness.

Apache-2.0Auto-check passedLegal & Compliance

Install Designing Privacy Preserving Analytics

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill designing-privacy-preserving-analytics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills designing-privacy-preserving-analytics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/designing-privacy-preserving-analytics .claude/skills/designing-privacy-preserving-analytics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
designing-privacy-preserving-analytics
GitHub stars
301
Token cost
~2.5k tokens
SKILL.md length
937 words
Files
5 (incl. scripts, references, assets)
Skills in repo
280
Repo updated
First seen
Licence
Apache-2.0

At a glance

Design privacy-preserving analytics systems using differential privacy, k-anonymity, l-diversity, and t-closeness.

  • Works in 4 steps: Identify quasi-identifiers in the dataset → Apply generalization hierarchies (e.g.,… → Apply suppression for records that… → …
  • Tasks that involve Privacy and GDPR
  • SKILL.md covers Overview, Statistical Disclosure Control…, Privacy Budget Management and Architecture Design, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Designing Privacy Preserving Analytics is an agent skill from mukul975/Privacy-Data-Protection-Skills. Design privacy-preserving analytics systems using differential privacy, k-anonymity, l-diversity, and t-closeness. Covers privacy budget allocation with epsilon tracking, references Google DP library, OpenDP, and Apple PPML. Includes Python differential privacy implementation for GDPR-compliant statistical analysis.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Legal & Compliance, covering Privacy and GDPR and Statistics. It works with Python. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Privacy and GDPR
  • Tasks that involve Statistics

Example prompts

  • “/designing-privacy-preserving-analytics”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify quasi-identifiers in the dataset
  2. Apply generalization hierarchies (e.g., age 27 → age range 25-30, postal code 8001 → 80**)
  3. Apply suppression for records that cannot achieve k-anonymity through generalization alone
  4. Verify k-value across all equivalence classes

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Designing Privacy Preserving Analytics loads about 2.5k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 937 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 937 words, ~2,541 tokens.

Download SKILL.mdSave it as .claude/skills/designing-privacy-preserving-analytics/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
designing-privacy-preserving-analytics
description
Design privacy-preserving analytics systems using differential privacy, k-anonymity, l-diversity, and t-closeness. Covers privacy budget allocation with epsilon tracking, references Google DP library, OpenDP, and Apple PPML. Includes Python differential privacy implementation for GDPR-compliant statistical analysis.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
privacy-by-design
metadata.tags
differential-privacy, k-anonymity, privacy-budget, opendp, privacy-preserving-analytics

Designing Privacy-Preserving Analytics

Overview

Privacy-preserving analytics enables organizations to extract statistical insights from personal data without exposing individual-level information. This directly supports GDPR Article 5(1)(c) (data minimization) and Recital 26 (which exempts truly anonymous data from the regulation). The Article 29 Working Party Opinion 05/2014 on Anonymisation Techniques (WP216) established that effective anonymization must resist singling out, linkability, and inference attacks.

Four primary statistical disclosure control techniques form the foundation of privacy-preserving analytics: differential privacy, k-anonymity, l-diversity, and t-closeness. Each offers different trade-offs between privacy guarantees and data utility.

Statistical Disclosure Control Techniques

Differential Privacy

Differential privacy (Dwork et al., 2006) provides a mathematical guarantee that the output of an analysis is approximately the same whether or not any single individual's data is included. Formally, a randomized mechanism M satisfies (epsilon, delta)-differential privacy if for all datasets D1 and D2 differing in at most one record, and for all sets of outputs S:

P[M(D1) ∈ S] ≤ e^ε × P[M(D2) ∈ S] + δ

Where epsilon (ε) is the privacy loss parameter and delta (δ) bounds the probability of privacy breach.

Privacy Budget Allocation:

Epsilon RangePrivacy LevelTypical Use Cases
0.01 — 0.1Very strongMedical research, genetic data, highly sensitive analytics
0.1 — 1.0StrongGeneral-purpose analytics, demographic analysis
1.0 — 5.0ModerateAggregate business metrics, trend analysis
5.0 — 10.0WeakLow-sensitivity counts, already-public statistics

Key Libraries:

LibraryMaintainerLanguageMechanism Types
Google DP LibraryGoogleC++/Java/GoLaplace, Gaussian, partition selection
OpenDPHarvard IQSS & MicrosoftRust/PythonComposable framework, Laplace, Gaussian, exponential
Apple PPMLAppleSwiftLocal DP, count-mean-sketch, Hadamard response
IBM diffprivlibIBM ResearchPythonScikit-learn compatible, ML with DP
PyDPOpenMinedPython (C++ backend)Python wrapper around Google DP library
k-Anonymity

A dataset satisfies k-anonymity (Sweeney, 2002) if every record is indistinguishable from at least k-1 other records with respect to quasi-identifier attributes. Quasi-identifiers are attributes that could be combined with external data to re-identify individuals (e.g., age, postal code, gender).

Implementation approach:

  1. Identify quasi-identifiers in the dataset
  2. Apply generalization hierarchies (e.g., age 27 → age range 25-30, postal code 8001 → 80**)
  3. Apply suppression for records that cannot achieve k-anonymity through generalization alone
  4. Verify k-value across all equivalence classes

Limitations: k-anonymity does not protect against attribute disclosure when sensitive values within an equivalence class are homogeneous.

l-Diversity

l-Diversity (Machanavajjhala et al., 2007) extends k-anonymity by requiring that each equivalence class contains at least l "well-represented" values of the sensitive attribute. This prevents attribute disclosure attacks.

Variants:

  • Distinct l-diversity: Each equivalence class has at least l distinct sensitive values
  • Entropy l-diversity: Entropy of sensitive values in each class ≥ log(l)
  • Recursive (c,l)-diversity: The most frequent sensitive value appears less than c times the frequency of the least frequent value
t-Closeness

t-Closeness (Li et al., 2007) requires that the distribution of a sensitive attribute in any equivalence class is within distance t of the distribution of the attribute in the entire dataset, measured using Earth Mover's Distance (EMD).

This prevents skewness attacks where an adversary can infer sensitive attributes from the distribution within an equivalence class, even when l-diversity is satisfied.

Show full SKILL.md (437 more words)Show less

Privacy Budget Management

Privacy budgets track cumulative privacy loss across multiple queries. Under sequential composition, the total privacy loss is the sum of individual epsilons. Under parallel composition (disjoint subsets), the total privacy loss is the maximum individual epsilon.

Budget Allocation Framework for Prism Data Systems AG:

Analytics FunctionEpsilon AllocationRefresh CadenceJustification
Daily active user counts0.1 per dayDailyLow sensitivity, high frequency
Revenue by region0.5 per quarterQuarterlyMedium sensitivity, aggregate metric
Feature usage patterns0.3 per monthMonthlyUsed for product development under Art. 6(1)(f)
Customer churn analysis0.2 per quarterQuarterlyInvolves behavioral profiling
A/B test results0.1 per experimentPer experimentBinary outcome, low disclosure risk
Total annual budget≤ 8.0Sum across all functions with composition

Budget Exhaustion Protocol:

  1. When 80% of annual budget is consumed, alert the Data Protection Officer
  2. When 95% is consumed, require DPO approval for each additional query
  3. When 100% is consumed, block all further differentially private queries until the next budget period
  4. Emergency queries require joint approval from DPO and Chief Information Security Officer

Architecture Design

┌──────────────────────────────────────────────────────────┐
│                    Analyst Interface                      │
│              (SQL-like query submission)                  │
└──────────────────────┬───────────────────────────────────┘
                       │
┌──────────────────────▼───────────────────────────────────┐
│                 Privacy Gateway                           │
│  ┌─────────────┐  ┌──────────────┐  ┌────────────────┐  │
│  │ Query Parser │  │ Budget Check │  │ Sensitivity    │  │
│  │ & Validator  │  │ (ε tracker)  │  │ Calibration    │  │
│  └──────┬──────┘  └──────┬───────┘  └───────┬────────┘  │
│         └────────────────┼──────────────────┘            │
└──────────────────────────┼───────────────────────────────┘
                           │
┌──────────────────────────▼───────────────────────────────┐
│              Noise Injection Layer                        │
│  ┌──────────────┐  ┌──────────────┐  ┌───────────────┐  │
│  │ Laplace      │  │ Gaussian     │  │ Exponential   │  │
│  │ Mechanism    │  │ Mechanism    │  │ Mechanism     │  │
│  └──────────────┘  └──────────────┘  └───────────────┘  │
└──────────────────────────┬───────────────────────────────┘
                           │
┌──────────────────────────▼───────────────────────────────┐
│              Data Processing Layer                        │
│  ┌─────────────────┐  ┌──────────────────────────────┐   │
│  │ k-Anonymization │  │ Aggregation Engine           │   │
│  │ Pre-processing  │  │ (min group size: 11)         │   │
│  └─────────────────┘  └──────────────────────────────┘   │
└──────────────────────────┬───────────────────────────────┘
                           │
┌──────────────────────────▼───────────────────────────────┐
│              Encrypted Data Store                         │
│         (Field-level AES-256-GCM encrypted)              │
└──────────────────────────────────────────────────────────┘

Implementation Workflow

  1. Classify Analytics Queries — Categorize each analytics use case by sensitivity (direct identifiers accessed, quasi-identifiers combined, sensitive attributes involved) and assign an epsilon budget allocation.

  2. Select Mechanism — Choose the appropriate noise mechanism based on query type: Laplace for counting queries, Gaussian for mean/variance queries requiring (ε,δ)-DP, exponential mechanism for selection queries.

  3. Calibrate Sensitivity — Determine the global sensitivity of each query function (maximum change in output when one record is added or removed). Use bounded sensitivity where possible by clipping input values.

  4. Implement Budget Tracking — Deploy a centralized privacy budget ledger that records every query's epsilon consumption, enforces composition bounds, and blocks queries when the budget is exhausted.

  5. Apply Pre-processing — Where differential privacy alone provides insufficient utility, apply k-anonymity as a pre-processing step to reduce the sensitivity of downstream DP queries.

  6. Validate Output — Verify that released statistics do not violate minimum group sizes (11 records per Prism Data Systems AG policy) and that confidence intervals are reported alongside noised results.

  7. Audit Trail — Log every privacy-preserving query with: timestamp, analyst identity, query hash, epsilon consumed, mechanism used, and cumulative budget remaining.

Key Regulatory References

  • GDPR Article 5(1)(c) — Data minimization principle
  • GDPR Article 5(1)(e) — Storage limitation
  • GDPR Article 25 — Data protection by design and by default
  • GDPR Article 89 — Safeguards for processing for scientific/historical research or statistical purposes
  • GDPR Recital 26 — Scope of anonymous information
  • GDPR Recital 162 — Statistical purposes
  • Article 29 Working Party Opinion 05/2014 on Anonymisation Techniques (WP216)
  • ENISA Report: Pseudonymisation techniques and best practices (November 2019)

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/designing-privacy-preserving-analytics of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Designing Privacy Preserving Analytics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Designing Privacy Preserving Analytics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Designing Privacy Preserving Analytics this skillmukul975/Privacy-Data-Protection-Skills301—~2.5kAutomated safety check: PassApache-2.0
HIPAA Safe Harbor Coverage Auditmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
De-identification Leakage Auditmaziyarpanahi/openmed5.5k—~1.9kAutomated safety check: PassApache-2.0
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
C15tc15t/c15t1.9k1 repos~1.6kAutomated safety check: PassApache-2.0
Hand Drawnthreerocks/hand-drawn-styles2.2k—~398Automated safety check: PassMIT

Similar skills

  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • De-identification Leakage Audit

    maziyarpanahi/openmed

    Scans text that has already been de-identified for leftover identifiers such as SSNs, card numbers, emails and dates, and blocks release if anything turns up.

    5.5k GitHub stars~1.9k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed
  • Hand Drawn

    threerocks/hand-drawn-styles

    A skill your agent uses when users ask for a hand-drawn or illustrated image prompt, name one of the repository's 20 numbered styles or the 3.1 stable variant, or mention triggers such as…

    2.2k GitHub stars~398 tokensUpdated 2 days ago
    Legal & ComplianceAuto-check passed
  • Tuya Smart Control

    tuya/tuya-openclaw-skills

    Control Tuya smart home devices via natural language. An agent skill from tuya/tuya-openclaw-skills.

    480 GitHub stars~5.2k tokensUpdated yesterday
    Backend & APIsAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 280 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    301 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    301 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Dpia

    mukul975/Privacy-Data-Protection-Skills

    Conducts Data Protection Impact Assessments for AI and ML systems per EDPB Guidelines 04/2025 on AI processing.

    301 GitHub stars~3.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    301 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    301 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    301 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Designing Privacy Preserving Analytics

What does Designing Privacy Preserving Analytics do?

Design privacy-preserving analytics systems using differential privacy, k-anonymity, l-diversity, and t-closeness. Designing Privacy Preserving Analytics is an agent skill from mukul975/Privacy-Data-Protection-Skills. Design privacy-preserving analytics systems using differential privacy, k-anonymity, l-diversity, and t-closeness.

When should I use Designing Privacy Preserving Analytics?

Designing Privacy Preserving Analytics fits situations like: tasks that involve Privacy and GDPR; tasks that involve Statistics.

How do I install Designing Privacy Preserving Analytics in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill designing-privacy-preserving-analytics -a claude-code`. Or copy the skill folder (skills/privacy/designing-privacy-preserving-analytics in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/designing-privacy-preserving-analytics in your project. Claude Code loads it when a task matches its description.

How do I install Designing Privacy Preserving Analytics in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill designing-privacy-preserving-analytics -a codex`. Or copy the skill folder (skills/privacy/designing-privacy-preserving-analytics in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/designing-privacy-preserving-analytics in your project. Codex loads it when a task matches its description.

Can I use Designing Privacy Preserving Analytics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill designing-privacy-preserving-analytics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/designing-privacy-preserving-analytics, .gemini/skills/designing-privacy-preserving-analytics, .github/skills/designing-privacy-preserving-analytics and .opencode/skills/designing-privacy-preserving-analytics in your project.

What does Designing Privacy Preserving Analytics need to run?

Going by SKILL.md and its folder, Designing Privacy Preserving Analytics needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Designing Privacy Preserving Analytics access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Designing Privacy Preserving Analytics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Designing Privacy Preserving Analytics use?

Designing Privacy Preserving Analytics is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Designing Privacy Preserving Analytics use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Designing Privacy Preserving Analytics?

Skills that share tags, products or a category with Designing Privacy Preserving Analytics: HIPAA Safe Harbor Coverage Audit (maziyarpanahi/openmed, 5.5k stars), De-identification Leakage Audit (maziyarpanahi/openmed, 5.5k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars) and C15t (c15t/c15t, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Designing Privacy Preserving Analytics?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 301 GitHub stars. The repository holds 280 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.