Agent skill

Pseudonymization Risk

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Assessment of pseudonymization techniques and re-identification risk.

Apache-2.0Auto-check passedLegal & Compliance

Install Pseudonymization Risk

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill pseudonymization-risk -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills pseudonymization-risk --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/pseudonymization-risk .claude/skills/pseudonymization-risk && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pseudonymization-risk
GitHub stars
295
Token cost
~3.2k tokens
SKILL.md length
1,418 words
Files
5 (incl. scripts, references, assets)
Skills in repo
278
Repo updated
First seen
Licence
Apache-2.0

At a glance

Assessment of pseudonymization techniques and re-identification risk.

  • Works in 3 steps: Technique Assessment (0-100) → Data Environment Assessment (0-100) → Residual Risk Score
  • Tasks that involve Cryptography
  • SKILL.md covers Overview, Pseudonymization Techniques…, Technique Comparison Matrix and Re-Identification Risk…, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Pseudonymization Risk is an agent skill from mukul975/Privacy-Data-Protection-Skills. Assessment of pseudonymization techniques and re-identification risk. Covers tokenization, hashing, encryption-based pseudonymization, and hybrid approaches. Includes re-identification risk scoring using the motivated intruder test, quantitative metrics (marketer, journalist, prosecutor models), and linkage attack resilience evaluation. References ENISA 2019 pseudonymization report.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Legal & Compliance, covering Cryptography, Natural language processing and Privacy and GDPR. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Cryptography
  • Tasks that involve Natural language processing
  • Tasks that involve Privacy and GDPR

Example prompts

  • “/pseudonymization-risk”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Technique Assessment (0-100)
  2. Data Environment Assessment (0-100)
  3. Residual Risk Score

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pseudonymization Risk loads about 3.2k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 102 tokens; SKILL.md has 1,418 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 1,418 words, ~3,240 tokens.

Download SKILL.mdSave it as .claude/skills/pseudonymization-risk/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
pseudonymization-risk
description
Assessment of pseudonymization techniques and re-identification risk. Covers tokenization, hashing, encryption-based pseudonymization, and hybrid approaches. Includes re-identification risk scoring using the motivated intruder test, quantitative metrics (marketer, journalist, prosecutor models), and linkage attack resilience evaluation. References ENISA 2019 pseudonymization report.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
privacy-by-design
metadata.tags
pseudonymization, re-identification, tokenization, hashing, enisa, motivated-intruder, linkage-attack

Assessing Pseudonymization and Re-Identification Risk

Overview

Pseudonymization (GDPR Article 4(5)) means processing personal data in such a manner that the data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure non-attribution. Unlike anonymization (Recital 26), pseudonymized data remains personal data — but it significantly reduces risk and is explicitly recognized as an appropriate safeguard under Articles 25, 32, and 89.

The ENISA report "Pseudonymisation Techniques and Best Practices" (November 2019) provides a comprehensive taxonomy of pseudonymization techniques and their properties. This skill implements a structured assessment framework for selecting appropriate techniques and quantifying residual re-identification risk.

Pseudonymization Techniques Taxonomy

Technique 1: Counter-Based Pseudonyms
PropertyValue
MechanismSequential or random counter assigned to each identifier
ReversibilityReversible via lookup table
Collision resistanceGuaranteed (bijective mapping)
PerformanceO(1) lookup
Storage overheadMapping table size proportional to identifier count
Key managementProtect the mapping table as additional information

Prism Data Systems AG Implementation: Customer IDs are replaced with UUIDv4 pseudonyms at the analytics ingestion boundary. The mapping table is stored in a separate HSM-backed database accessible only to the DPO and the identity resolution service (which requires dual-approval access).

Technique 2: Cryptographic Hashing (HMAC)
PropertyValue
MechanismHMAC-SHA256 with secret key applied to identifier
ReversibilityNot directly reversible; vulnerable to brute-force on low-entropy inputs
Collision resistanceNegligible collision probability (256-bit output)
PerformanceO(1) computation
Storage overheadFixed 32-byte output per identifier
Key managementHMAC key must be protected as additional information

Risk: Plain SHA-256 hashing (without key) is vulnerable to rainbow table attacks on low-entropy identifiers like email addresses. HMAC with a secret key mitigates this but the key becomes the single point of re-identification.

Prism Data Systems AG Implementation: HMAC-SHA256 with a 256-bit key rotated quarterly is used for pseudonymizing user identifiers in the analytics pipeline. The HMAC key is stored in AWS KMS with access restricted to the pseudonymization service account. Key rotation triggers re-pseudonymization of the active analytics dataset.

Technique 3: Encryption-Based Pseudonymization
PropertyValue
MechanismAES-256-GCM or format-preserving encryption (FF1/FF3-1)
ReversibilityFully reversible with decryption key
Collision resistanceGuaranteed (bijective for same key)
PerformanceO(1) computation; FPE slower than standard AES
Storage overheadCiphertext size ≈ plaintext + 16-byte tag (GCM)
Key managementEncryption key is the additional information per Art. 4(5)

Format-Preserving Encryption preserves the format of the original identifier (e.g., encrypting a 10-digit phone number produces another 10-digit number), enabling pseudonymization without schema changes.

Prism Data Systems AG Implementation: FPE (FF3-1 mode) is used for pseudonymizing phone numbers and IBAN fields in the support ticket system. This allows Tier 1 support agents to see structurally valid but pseudonymized identifiers, reducing accidental exposure while maintaining workflow compatibility.

Technique 4: Tokenization
PropertyValue
MechanismRandom token generated and stored in a secure vault; no mathematical relationship to original
ReversibilityReversible only via vault lookup
Collision resistanceGuaranteed by vault uniqueness constraint
PerformanceO(1) vault lookup; network latency for centralized vault
Storage overheadVault stores original-to-token mapping
Key managementVault access controls replace key management

Prism Data Systems AG Implementation: Payment card data is tokenized using a PCI DSS-compliant token vault. The vault is operated by the payment processor and is logically separated from Prism's infrastructure. Detokenization requires a PCI-scoped service account with transaction-level authorization.

Technique 5: Synthetic Data Replacement
PropertyValue
MechanismReplace real identifiers with structurally similar but fictional values
ReversibilityNot reversible (one-way replacement)
Collision resistanceDepends on generation method; risk of accidental collision with real data
PerformanceO(1) generation
Storage overheadNo mapping table needed (irreversible)
Key managementNot applicable

Use case: Test environments where referential integrity must be maintained but no re-identification path is acceptable.

Technique Comparison Matrix

CriterionCounterHMACEncryptionTokenizationSynthetic
ReversibilityLookup tableBrute-force onlyDecryption keyVault lookupIrreversible
DeterministicConfigurableYes (same key)Yes (same key+IV)ConfigurableNo
Format preservingNoNoFPE onlyConfigurableYes
Linkability (same dataset)YesYesYesConfigurableNo
Linkability (cross-dataset)NoSame key: YesSame key: YesNoNo
Brute-force resistanceN/AKey-dependentKey-dependentN/AN/A
Suited for analyticsYesYesLimitedYesLimited
Suited for testingNoNoNoYesYes
ENISA recommendation levelBasicIntermediateAdvancedAdvancedSpecialized

Re-Identification Risk Assessment Framework

The Motivated Intruder Test

The UK Information Commissioner's Office (ICO) Anonymisation Code of Practice defines the motivated intruder test: would a reasonably competent, motivated person with access to resources such as the internet, public libraries, and public records be able to identify an individual from the pseudonymized data?

Assessment Factors:

FactorLow Risk (1)Medium Risk (3)High Risk (5)
Population uniquenessCommon attributes, large equivalence classesModerate uniqueness, some rare combinationsHighly unique attribute combinations
Auxiliary information availabilityNo public datasets linkableSome public records overlapRich public profiles (social media, registers)
Identifier entropyHigh entropy (e.g., UUID)Medium entropy (e.g., hashed email)Low entropy (e.g., hashed phone number)
Dataset richnessFew quasi-identifiersModerate quasi-identifiers (5-10)Many quasi-identifiers (>10)
Temporal precisionYearly or coarserMonthlyDaily or finer
Geographic precisionCountry levelRegion/city levelPostcode or finer
Domain sensitivityLow sensitivity domainModerate sensitivityHealth, financial, political data
Show full SKILL.md (556 more words)Show less
Quantitative Risk Models
Prosecutor Model

The attacker knows a specific individual is in the dataset and attempts to find their record. Risk = 1 / k, where k is the size of the equivalence class containing the target record.

Journalist Model

The attacker does not know if the target is in the dataset but attempts to re-identify any individual for a story. Risk = max(1/k) across all equivalence classes (the smallest group is most vulnerable).

Marketer Model

The attacker attempts to re-identify as many individuals as possible. Risk = (1/n) * Σ(1/k_i) — the average re-identification probability across all records.

Linkage Attack Categories
Attack TypeDescriptionMitigation
Record linkageMatching pseudonymized records with identified records in external datasets using quasi-identifiersGeneralize quasi-identifiers, enforce k-anonymity
Attribute linkageInferring sensitive attributes from group characteristics even without identifying the individualEnforce l-diversity on sensitive attributes
Table linkageDetermining that an individual is present in a sensitive dataset (membership inference)Apply differential privacy, enforce t-closeness
Composition attackCombining multiple independent pseudonymized releases to narrow equivalence classesTrack cumulative privacy budget, limit releases
Temporal linkageLinking records across time periods using behavioral patterns or trajectory dataApply temporal generalization, add noise to timestamps

Risk Scoring Methodology

Step 1: Technique Assessment (0-100)

Score the pseudonymization technique on five properties:

PropertyWeightScore Range
Key/mapping security30%0-100 based on key management maturity
Brute-force resistance25%0-100 based on input entropy and computational cost
Cross-dataset unlinkability20%0-100 based on determinism and key reuse
Implementation maturity15%0-100 based on library validation and audit history
Operational resilience10%0-100 based on key rotation, backup, disaster recovery
Step 2: Data Environment Assessment (0-100)

Score the data environment on re-identification risk factors:

FactorWeightScore Range
Population uniqueness25%0-100 (higher = more unique = higher risk)
Auxiliary data availability25%0-100 (higher = more auxiliary data = higher risk)
Quasi-identifier count20%0-100 based on number and granularity
Dataset size15%0-100 (smaller datasets = higher risk per record)
Release frequency15%0-100 (more frequent releases = higher composition risk)
Step 3: Residual Risk Score
Residual Risk = Data Environment Risk × (1 - Technique Score / 100)
Residual Risk RangeClassificationAction Required
0-15LowAcceptable. Document and monitor
16-35ModerateAdditional controls recommended
36-60HighTechnique upgrade or data reduction required
61-100Very HighProcessing should not proceed without anonymization

Prism Data Systems AG — Assessment Summary

Data FlowTechniqueTechnique ScoreEnvironment RiskResidual RiskClassification
Analytics pipelineHMAC-SHA256 (keyed)78429.2Low
Support dashboardFPE (FF3-1)82356.3Low
Payment processingTokenization (vault)91555.0Low
Test environmentsSynthetic replacement95201.0Low
Research exportsCounter + k-anonymity725816.2Moderate
ML training dataHMAC + differential privacy85487.2Low

The research exports flow received a "Moderate" classification due to the combination of rich quasi-identifiers in the exported dataset and the counter-based technique's susceptibility to record linkage. Remediation: apply l-diversity (l≥3) on sensitive attributes and generalize quasi-identifiers to achieve k≥11 before export.

Key Regulatory References

  • GDPR Article 4(5) — Definition of pseudonymization
  • GDPR Recital 26 — Anonymization vs. pseudonymization distinction
  • GDPR Article 25(1) — Pseudonymization as a data protection by design measure
  • GDPR Article 32(1)(a) — Pseudonymization as a security measure
  • GDPR Article 89(1) — Pseudonymization for research processing safeguards
  • ENISA "Pseudonymisation Techniques and Best Practices" (November 2019)
  • ICO Anonymisation Code of Practice — Motivated intruder test
  • Article 29 Working Party Opinion 05/2014 (WP216) on Anonymisation Techniques
  • EDPB Guidelines 4/2019 on Article 25 — Pseudonymization as by-design measure

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/pseudonymization-risk of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Pseudonymization Risk next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pseudonymization Risk compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pseudonymization Risk this skillmukul975/Privacy-Data-Protection-Skills295—~3.2kAutomated safety check: PassApache-2.0
Generating Synthetic Surrogatesmaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Data Protection And Encryptioncbrock84/headcount2k—~1.3kAutomated safety check: PassMIT
Clade Security Basicsjeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: NotesMIT
Stellar DevVelaPayments/vela-payments131—~1.8kAutomated safety check: PassMIT
Security Compliancesangrokjung/claude-forge8502 repos~7.2kAutomated safety check: PassMIT

Similar skills

  • Generating Synthetic Surrogates

    maziyarpanahi/openmed

    Replace detected PHI with realistic, type-matched fake values in OpenMed so clinical notes stay readable and parseable instead of full of [REDACTED] markers.

    5.5k GitHub stars~1.8k tokensUpdated yesterday
    Legal & ComplianceAuto-check passed
  • Protects data itself rather than the systems around it — classifying what you hold, encrypting in transit and at rest and understanding what each actually defends against, managing keys and their…

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Legal & ComplianceAuto-check passed
  • Clade Security Basics

    jeremylongshore/tons-of-skills-marketplace

    Secure your Anthropic integration — API key management, input validation, Use when working with security-basics patterns.

    2.8k GitHub stars~1.1k tokensUpdated today
    SecurityAuto-check: notes
  • Stellar Dev

    VelaPayments/vela-payments

    End-to-end Stellar development playbook. An agent skill from VelaPayments/vela-payments.

    131 GitHub stars~1.8k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Security Compliance

    sangrokjung/claude-forge

    Guides security professionals in implementing defense-in-depth security architectures, achieving compliance with industry frameworks (SOC2, ISO27001, GDPR, HIPAA), conducting threat modeling and…

    850 GitHub starsUsed in 2 repos~7.2k tokens
    Legal & ComplianceAuto-check passed
  • Incident Reporting Navigator

    davila7/claude-code-templates

    A skill your agent uses when a security incident, data breach, or actively exploited vulnerability raises the question "who must we notify, where, and by when?" Screens one incident across the EU…

    32k GitHub starsUsed in 1 repo~4.7k tokens
    Legal & ComplianceAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 278 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    295 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    295 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    295 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed
  • Retention Schedule

    mukul975/Privacy-Data-Protection-Skills

    Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.

    295 GitHub stars~3.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Pseudonymization Risk

What does Pseudonymization Risk do?

Assessment of pseudonymization techniques and re-identification risk. Pseudonymization Risk is an agent skill from mukul975/Privacy-Data-Protection-Skills. Assessment of pseudonymization techniques and re-identification risk.

When should I use Pseudonymization Risk?

Pseudonymization Risk fits situations like: tasks that involve Cryptography; tasks that involve Natural language processing; tasks that involve Privacy and GDPR.

How do I install Pseudonymization Risk in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill pseudonymization-risk -a claude-code`. Or copy the skill folder (skills/privacy/pseudonymization-risk in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/pseudonymization-risk in your project. Claude Code loads it when a task matches its description.

How do I install Pseudonymization Risk in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill pseudonymization-risk -a codex`. Or copy the skill folder (skills/privacy/pseudonymization-risk in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/pseudonymization-risk in your project. Codex loads it when a task matches its description.

Can I use Pseudonymization Risk in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill pseudonymization-risk -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pseudonymization-risk, .gemini/skills/pseudonymization-risk, .github/skills/pseudonymization-risk and .opencode/skills/pseudonymization-risk in your project.

What does Pseudonymization Risk need to run?

Going by SKILL.md and its folder, Pseudonymization Risk needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Pseudonymization Risk access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pseudonymization Risk safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pseudonymization Risk use?

Pseudonymization Risk is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pseudonymization Risk use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.1k tokens, read only when the agent opens those files.

What are the alternatives to Pseudonymization Risk?

Skills that share tags, products or a category with Pseudonymization Risk: Generating Synthetic Surrogates (maziyarpanahi/openmed, 5.5k stars), Data Protection And Encryption (cbrock84/headcount, 2k stars), Clade Security Basics (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Stellar Dev (VelaPayments/vela-payments, 131 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pseudonymization Risk?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.