Agent skill

Data Lineage Tracking

by mukul975 in mukul975/Privacy-Data-Protection-Skills

Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs.

Apache-2.0Auto-check passedLegal & Compliance

Install Data Lineage Tracking

skills CLI
$ npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-lineage-tracking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Privacy-Data-Protection-Skills data-lineage-tracking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/privacy/data-lineage-tracking .claude/skills/data-lineage-tracking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-lineage-tracking
GitHub stars
295
Token cost
~1.8k tokens
SKILL.md length
833 words
Files
5 (incl. scripts, references, assets)
Skills in repo
278
Repo updated
First seen
Licence
Apache-2.0

At a glance

Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs.

  • Works in 8 steps: Define Lineage Scope → Identify Data Sources (Collection Points) → Map Data Transformations → …
  • Tasks that involve Data governance
  • SKILL.md covers Purpose, Prerequisites, Workflow and Verification
  • Runs Python scripts from its folder

What it does

Data Lineage Tracking is an agent skill from mukul975/Privacy-Data-Protection-Skills. Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs. Covers source-to-sink mapping, GDPR Art. 30 RoPA integration, automated lineage discovery, and breach impact scoping. Keywords: data lineage, data provenance, data flow mapping, transformation logging, deletion verification.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts, reference files and assets (for example `assets/template.md`, `references/standards.md` and `references/workflows.md`).

It sits in Legal & Compliance, covering Data governance, Privacy and GDPR and Reproducible research. The repository describes itself as: 282+ structured privacy & data protection skills for AI agents. GDPR, CCPA, EU AI Act, HIPAA, LGPD, PIPL, DPDP Act. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Data governance
  • Tasks that involve Privacy and GDPR
  • Tasks that involve Reproducible research

Example prompts

  • “Use the data-lineage-tracking skill to implement data lineage tracking for privacy compliance including origin tracking, transformation logging…”
  • “/data-lineage-tracking”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Define Lineage Scope
  2. Identify Data Sources (Collection Points)
  3. Map Data Transformations
  4. Document Data Flows and Transfers
  5. Map Data Storage and Retention
  6. Implement Automated Lineage Discovery
  7. Integrate with RoPA and Compliance
  8. Maintain and Validate

What it can do on your machine

Read from SKILL.md and the folder at commit 9b2ef9e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Lineage Tracking loads about 1.8k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 833 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Privacy-Data-Protection-Skills at commit 9b2ef9e, republished under its Apache-2.0 licence (© mukul975). 833 words, ~1,815 tokens.

Download SKILL.mdSave it as .claude/skills/data-lineage-tracking/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
data-lineage-tracking
description
Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs. Covers source-to-sink mapping, GDPR Art. 30 RoPA integration, automated lineage discovery, and breach impact scoping. Keywords: data lineage, data provenance, data flow mapping, transformation logging, deletion verification.
license
Apache-2.0
metadata.author
mukul975
metadata.version
1.0
metadata.domain
privacy
metadata.subdomain
data-classification
metadata.tags
data-lineage, data-provenance, data-flow-mapping, transformation-logging, deletion-verification

Data Lineage Tracking

Purpose

Map and document the complete lifecycle of personal data through organizational systems, from initial collection point through all transformations, transfers, and storage locations to final deletion or anonymization.

Prerequisites

  • Data inventory or asset register identifying systems processing personal data
  • Network architecture documentation showing system interconnections
  • Access to system metadata, ETL configurations, and API documentation
  • Records of Processing Activities (RoPA) under GDPR Article 30

Workflow

Step 1: Define Lineage Scope

Determine scope boundaries for the lineage mapping exercise:

  1. Data categories: Identify which personal data categories to trace (identity data, contact data, financial data, behavioral data, special category data under Art. 9)
  2. System boundaries: Define which systems are in-scope (production databases, data warehouses, analytics platforms, third-party SaaS, backup systems)
  3. Temporal scope: Determine whether to map current-state lineage only or include historical data flows
  4. Legal basis mapping: Link each processing activity in the lineage to its GDPR Article 6 lawful basis
Step 2: Identify Data Sources (Collection Points)

Document every point where personal data enters the organization:

  1. Direct collection: Web forms, mobile apps, point-of-sale terminals, customer service interactions, paper forms digitized via scanning
  2. Indirect collection: Third-party data providers, publicly available sources, data brokers, partner organizations
  3. Derived data: Data generated through processing (risk scores, customer segments, behavioral profiles)
  4. Inferred data: Data inferred from other data points (creditworthiness, health predictions, preferences)

For each source, record:

  • Source identifier and type
  • Data categories collected (referencing Art. 30(1)(c) categories)
  • Legal basis under Art. 6(1) and, if applicable, Art. 9(2)
  • Information provided to data subjects per Art. 13 or Art. 14
  • Volume and frequency of collection
Step 3: Map Data Transformations

Document every transformation applied to personal data:

  1. ETL processes: Extract-Transform-Load pipelines moving data between systems
  2. Aggregation: Grouping individual records into summary statistics
  3. Pseudonymization: Replacing identifiers with tokens per Art. 4(5) and Recital 26
  4. Anonymization: Irreversible de-identification per WP29 Opinion 05/2014
  5. Enrichment: Combining data from multiple sources to create enriched profiles
  6. Format conversion: Changing data formats (CSV to JSON, database migration)

For each transformation, record:

  • Input data categories and source system
  • Transformation logic description
  • Output data categories and destination system
  • Whether transformation changes the identifiability of data subjects
  • Retention period at destination per Art. 5(1)(e) storage limitation
Step 4: Document Data Flows and Transfers

Map all movements of personal data between systems and parties:

  1. Internal flows: Between departments, systems, databases within the organization
  2. Processor transfers: To data processors under Art. 28 agreements
  3. Third-country transfers: Cross-border transfers requiring Art. 44-49 safeguards
  4. Third-party disclosures: To independent controllers (regulators, partners, law enforcement)

For each flow, record:

  • Source and destination system/entity
  • Transfer mechanism (API, file transfer, database replication, manual export)
  • Legal safeguard for international transfers (SCCs, BCRs, adequacy decision)
  • Encryption in transit and at rest
  • Frequency and volume
Show full SKILL.md (366 more words)Show less
Step 5: Map Data Storage and Retention

Document where personal data resides at each stage:

  1. Primary storage: Production databases, CRM systems, HRIS
  2. Secondary storage: Data warehouses, analytics databases, reporting systems
  3. Archival storage: Long-term archives, cold storage, compliance archives
  4. Backup storage: Disaster recovery systems, backup tapes, cloud backup
  5. Temporary storage: Caches, message queues, log files, session storage

For each storage location, record:

  • Storage technology and location (on-premises, cloud region)
  • Retention period and legal basis for retention
  • Access controls and encryption
  • Deletion or anonymization mechanism at end of retention period
Step 6: Implement Automated Lineage Discovery

Deploy tooling to automate lineage tracking:

  1. Database-level lineage: Query log analysis, column-level lineage from SQL parsing
  2. Application-level lineage: API call tracing, service mesh observability
  3. Pipeline-level lineage: ETL tool metadata (Apache Airflow lineage backend, dbt documentation)
  4. Infrastructure-level lineage: Network flow logs, data lake audit trails

Use the scripts/process.py helper to parse system metadata and generate lineage graphs.

Step 7: Integrate with RoPA and Compliance

Link lineage data to GDPR compliance documentation:

  1. Art. 30 RoPA: Each lineage path should map to a processing activity in the RoPA
  2. DPIA triggers: Flag lineage paths involving Art. 35(3) processing (large-scale profiling, systematic monitoring, special categories)
  3. Data subject rights: Use lineage to locate all data for DSAR responses (Art. 15 access, Art. 17 erasure, Art. 20 portability)
  4. Breach impact scoping: Use lineage to determine affected data subjects and categories during incident response per Art. 33(3)
Step 8: Maintain and Validate

Establish ongoing lineage maintenance:

  1. Change management: Update lineage when new systems, data flows, or processing activities are introduced
  2. Periodic validation: Quarterly review to verify lineage accuracy against actual system behavior
  3. Stakeholder review: Annual sign-off from data owners, system architects, and DPO
  4. Completeness check: Cross-reference lineage against data inventory and RoPA to identify gaps

Verification

  • All personal data categories have documented source-to-sink lineage paths
  • Each lineage node includes legal basis, retention period, and access controls
  • Transformation steps document identifiability impact
  • Cross-border transfers identified with Art. 44-49 safeguards documented
  • Lineage integrates with Art. 30 RoPA entries
  • Automated discovery covers at least database and pipeline layers
  • Quarterly validation schedule established with responsible owners assigned

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references, assets) in skills/privacy/data-lineage-tracking of mukul975/Privacy-Data-Protection-Skills.

  • SKILL.md
  • assets/template.md
  • references/standards.md
  • references/workflows.md
  • scripts/process.py

Open the folder on GitHubat commit 9b2ef9e

Compare with similar skills

Data Lineage Tracking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Lineage Tracking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Lineage Tracking this skillmukul975/Privacy-Data-Protection-Skills295—~1.8kAutomated safety check: PassApache-2.0
PII Annotation and Agreement Scorerglebis/claude-skills389—~1.3kAutomated safety check: PassMIT
Implementing Cloud Dlp For Data Protectionmukul975/Anthropic-Cybersecurity-Skills34k—~4.2kAutomated safety check: PassApache-2.0
Cursor Compliance Auditjeremylongshore/tons-of-skills-marketplace2.8k—~2.3kAutomated safety check: NotesMIT
Posthog Data Handlingjeremylongshore/tons-of-skills-marketplace2.8k—~2.7kAutomated safety check: PassMIT
Cometchat Compliancecometchat/cometchat-skills130—~1.7kAutomated safety check: PassMIT

Similar skills

  • Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

    389 GitHub stars~1.3k tokensUpdated 12 days ago
    Legal & ComplianceAuto-check passed
  • Implementing Cloud Dlp For Data Protection

    mukul975/Anthropic-Cybersecurity-Skills

    Implement cloud DLP using Amazon Macie, Google Cloud DLP API, Microsoft Purview, Azure Information Protection, and Nightfall AI to discover, classify, label, de-identify, and protect sensitive data…

    34k GitHub stars~4.2k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Cursor Compliance Audit

    jeremylongshore/tons-of-skills-marketplace

    Compliance and security auditing for Cursor IDE usage: SOC 2, GDPR, HIPAA assessment, evidence collection, and remediation.

    2.8k GitHub stars~2.3k tokensUpdated today
    Legal & ComplianceAuto-check: notes
  • Posthog Data Handling

    jeremylongshore/tons-of-skills-marketplace

    Implement consent-aware PostHog collection, PII minimization, masking, retention, and deletion workflows.

    2.8k GitHub stars~2.7k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Cometchat Compliance

    cometchat/cometchat-skills

    Data governance & compliance for CometChat — pick the data-residency region, satisfy GDPR/CCPA (right-to-erasure and data export), plan message retention & purge, and produce audit / eDiscovery…

    130 GitHub stars~1.7k tokensUpdated 3 days ago
    Legal & ComplianceAuto-check passed
  • Data Policy

    ericrisco/rsc-harness

    A skill your agent uses when building internal data-governance machinery: a retention schedule (period, lawful basis, expiry action, system where deletion runs), an Art.

    167 GitHub stars~3.1k tokensUpdated yesterday
    Legal & ComplianceAuto-check passed

More from mukul975/Privacy-Data-Protection-Skills

All 278 skills in this repo
  • Age Gating Services

    mukul975/Privacy-Data-Protection-Skills

    Implements age-gating mechanisms for online services to restrict access based on user age.

    295 GitHub stars~3.7k tokensUpdated 6 mo ago
    Auto-check passed
  • AI Data Retention

    mukul975/Privacy-Data-Protection-Skills

    Manages AI model retention and machine unlearning requirements.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Dpia Mitigation Plan

    mukul975/Privacy-Data-Protection-Skills

    Structures risk mitigation planning and residual risk tracking for Data Protection Impact Assessments under GDPR Article 35(7)(d).

    295 GitHub stars~846 tokensUpdated 6 mo ago
    Auto-check passed
  • Gdpr Accountability

    mukul975/Privacy-Data-Protection-Skills

    Guides implementation of the GDPR accountability principle under Articles 5(2) and 24, including documentation requirements for policies, DPIAs, RoPA, training records, and breach logs.

    295 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Pia Threshold Screening

    mukul975/Privacy-Data-Protection-Skills

    Conducts pre-DPIA threshold screening to determine whether a full Data Protection Impact Assessment is required under GDPR Article 35.

    295 GitHub stars~880 tokensUpdated 6 mo ago
    Auto-check passed
  • Retention Schedule

    mukul975/Privacy-Data-Protection-Skills

    Designs and implements data retention schedules compliant with GDPR Article 5(1)(e) storage limitation principle.

    295 GitHub stars~3.3k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Data Lineage Tracking

What does Data Lineage Tracking do?

Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs. Data Lineage Tracking is an agent skill from mukul975/Privacy-Data-Protection-Skills. Implements data lineage tracking for privacy compliance including origin tracking, transformation logging, access auditing, deletion verification, and cross-system lineage graphs.

When should I use Data Lineage Tracking?

Data Lineage Tracking fits situations like: tasks that involve Data governance; tasks that involve Privacy and GDPR; tasks that involve Reproducible research.

How do I install Data Lineage Tracking in Claude Code?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-lineage-tracking -a claude-code`. Or copy the skill folder (skills/privacy/data-lineage-tracking in mukul975/Privacy-Data-Protection-Skills) into .claude/skills/data-lineage-tracking in your project. Claude Code loads it when a task matches its description.

How do I install Data Lineage Tracking in Codex?

Run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-lineage-tracking -a codex`. Or copy the skill folder (skills/privacy/data-lineage-tracking in mukul975/Privacy-Data-Protection-Skills) into .agents/skills/data-lineage-tracking in your project. Codex loads it when a task matches its description.

Can I use Data Lineage Tracking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Privacy-Data-Protection-Skills --skill data-lineage-tracking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-lineage-tracking, .gemini/skills/data-lineage-tracking, .github/skills/data-lineage-tracking and .opencode/skills/data-lineage-tracking in your project.

What does Data Lineage Tracking need to run?

Going by SKILL.md and its folder, Data Lineage Tracking needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Data Lineage Tracking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Lineage Tracking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Lineage Tracking use?

Data Lineage Tracking is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Lineage Tracking use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Data Lineage Tracking?

Skills that share tags, products or a category with Data Lineage Tracking: PII Annotation and Agreement Scorer (glebis/claude-skills, 389 stars), Implementing Cloud Dlp For Data Protection (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Cursor Compliance Audit (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Posthog Data Handling (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Lineage Tracking?

mukul975 (a GitHub user) maintains it in mukul975/Privacy-Data-Protection-Skills, which has 295 GitHub stars. The repository holds 278 skills in this directory. The repository was last updated on March 16, 2026.

Source: mukul975/Privacy-Data-Protection-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.