Agent skill

Data Pipeline Spec

by mohitagw15856 in mohitagw15856/pm-claude-skills

Design an ETL/ELT data pipeline specification. An agent skill from mohitagw15856/pm-claude-skills.

MITAuto-check passedData & Analytics

Install Data Pipeline Spec

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill data-pipeline-spec -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills data-pipeline-spec --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-pipeline-spec .claude/skills/data-pipeline-spec && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-pipeline-spec
GitHub stars
1.4k
Token cost
~2.5k tokens
SKILL.md length
1,123 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Design an ETL/ELT data pipeline specification. An agent skill from mohitagw15856/pm-claude-skills.

  • Works in 11 steps: Overview → Sources → Ingestion Layer → …
  • Asked to design a data pipeline
  • SKILL.md covers Required Inputs, Output Structure, 1. Overview and 2. Sources, plus 12 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Pipeline Spec is an agent skill from mohitagw15856/pm-claude-skills. Design an ETL/ELT data pipeline specification. Use when asked to design a data pipeline, spec an ETL or ELT process, document a data ingestion workflow, or plan a data integration. Produces a complete pipeline spec with sources, transforms, destinations, SLAs, error handling, and data quality rules.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to design a data pipeline
  • Document a data ingestion workflow
  • Plan a data integration

Example prompts

  • “/data-pipeline-spec”

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Overview
  2. Sources
  3. Ingestion Layer
  4. Transformation Logic
  5. Destination
  6. Scheduling & SLAs
  7. Data Quality Rules
  8. Error Handling & Recovery
  9. Monitoring & Observability
  10. Dependencies & Sequencing
  11. Security & Compliance

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Pipeline Spec loads about 2.5k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 1,123 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 1,123 words, ~2,469 tokens.

Download SKILL.mdSave it as .claude/skills/data-pipeline-spec/SKILL.md (or your agent's skills folder).
name
data-pipeline-spec
description
Design an ETL/ELT data pipeline specification. Use when asked to design a data pipeline, spec an ETL or ELT process, document a data ingestion workflow, or plan a data integration. Produces a complete pipeline spec with sources, transforms, destinations, SLAs, error handling, and data quality rules.

Data Pipeline Spec Skill

This skill produces a complete data pipeline specification covering sources, transformations, destinations, scheduling, SLAs, error handling, data quality checks, and monitoring requirements. Output is ready for engineering handoff or architecture review.

Required Inputs

Ask the user for these if not provided:

  • Pipeline purpose — what business question or workflow does this pipeline serve?
  • Source systems — where does data come from? (databases, APIs, files, event streams)
  • Destination — where does data land? (data warehouse, data lake, downstream DB, reporting tool)
  • Transformation type — ETL (transform before loading) or ELT (load raw, transform in warehouse)?
  • Frequency / SLA — how often must data be fresh? (real-time / hourly / daily / weekly)
  • Volume estimate — approximate rows/events per run
  • Data quality requirements — completeness, deduplication, freshness, schema enforcement
  • Team or stack — any specific tools in use? (Airflow, dbt, Fivetran, Spark, Kafka, etc.)

Output Structure


Data Pipeline Spec: [Pipeline Name]

Purpose: [One sentence — what decision or workflow does this pipeline enable?] Type: [ETL / ELT / Streaming / Batch] Owner: [Team or individual] Version: [1.0] Date: [Date] Status: [Draft / Under Review / Approved]


1. Overview

[2–3 sentences describing the pipeline end-to-end: what data moves, from where to where, at what cadence, and why.]

Architecture diagram (text):

[Source A] ──┐
[Source B] ──┤──► [Ingestion Layer] ──► [Transform Layer] ──► [Destination] ──► [Consumers]
[Source C] ──┘

2. Sources

SourceSystemConnection typeData formatUpdate patternVolume
[Source 1][PostgreSQL / Salesforce / S3 / Kafka][JDBC / REST API / SDK / Webhook][JSON / CSV / Parquet / CDC][Append / Full refresh / Incremental][X rows/day]
[Source 2][...][...][...][...][...]

Incremental key (if applicable): [The column used to identify new or changed records — e.g. updated_at, event_id]

Authentication: [API key / OAuth / IAM role / connection string — note where credentials are stored]


3. Ingestion Layer

Tool: [Fivetran / Airbyte / Kafka Connect / custom script / dbt source]

Ingestion method:

  • Full extract (full table refresh each run)
  • Incremental extract (only new/changed rows since last run)
  • CDC (change data capture from database transaction log)
  • Event streaming (continuous ingestion from Kafka/Kinesis)

Raw landing zone: [Where raw data lands before transformation — e.g. raw.salesforce_opportunities in Snowflake, S3 bucket s3://data-raw/crm/]

Schema handling: [Strict schema enforcement / Schema evolution allowed / Union schema]


4. Transformation Logic

List each transformation in execution order. For ELT pipelines, this is the dbt model or SQL layer.

StepNameDescriptionInputOutputTool
1[Deduplicate events][Remove duplicate event rows based on event_id]raw.eventsstaging.events_deduped[dbt / SQL / Spark]
2[Join user profile][Enrich events with user attributes from CRM]staging.events_deduped, raw.usersstaging.events_enriched[...]
3[Aggregate to daily][Roll up to user×day grain]staging.events_enrichedmart.user_daily_activity[...]

Business logic rules:

  • [e.g. Revenue is recognised on payment_confirmed_at, not payment_initiated_at]
  • [e.g. Users in the internal@company.com domain are excluded from all metrics]
  • [e.g. Currency conversion uses the ECB rate from the first business day of each month]

Slowly Changing Dimensions (SCD) — if applicable:

  • [e.g. users.plan_tier is SCD Type 2 — keep history of plan changes with valid_from / valid_to]

5. Destination

DestinationSystemSchema / TableWrite modeConsumers
[Primary][Snowflake / BigQuery / Redshift / PostgreSQL][analytics.mart_user_activity][Append / Upsert / Full replace][Looker / Metabase / downstream pipeline]
[Secondary][...][...][...][...]

Partitioning / Clustering: [e.g. Partitioned by event_date, clustered by user_id — reduces query cost for time-range scans]

Retention policy: [e.g. Raw data retained for 90 days; mart tables retained indefinitely]


6. Scheduling & SLAs

SLATargetBreach action
Data freshness[Data must be ≤ X hours old by HH:MM UTC][Page on-call / alert Slack channel]
Pipeline completion[Must complete within X minutes of trigger][Alert and auto-retry]
Availability[Pipeline must run successfully X% of days per month][Incident review]

Schedule: [Cron expression and human description — e.g. 0 6 * * * — daily at 06:00 UTC]

Trigger type:

  • Time-based (cron)
  • Event-based (triggered by upstream pipeline success / file arrival / Kafka lag)
  • Manual (ad hoc runs only)

Backfill strategy: [How to reprocess historical data if the pipeline fails or logic changes — e.g. parameterised date range, full drop-and-reload]


7. Data Quality Rules

CheckTableRuleFailure action
Completenessstaging.eventsevent_id IS NOT NULL — 100% of rowsBlock load / Alert
Uniquenessmart.user_daily_activity(user_id, date) must be uniqueBlock load
Freshnessmart.user_daily_activitymax(event_date) >= CURRENT_DATE - 1Alert
Volumestaging.eventsRow count within ±20% of 7-day averageAlert
Referential integritystaging.eventsAll user_id values exist in users tableAlert

DQ tool: [dbt tests / Great Expectations / Monte Carlo / custom SQL assertions]


Show full SKILL.md (456 more words)Show less

8. Error Handling & Recovery

Retry policy: [e.g. 3 retries with exponential back-off: 5 min, 20 min, 60 min]

Failure modes and responses:

FailureDetectionResponseOwner
Source unavailableHTTP 5xx / connection timeoutRetry 3×, then alert and skip runData engineering
Schema change in sourceColumn missing or type mismatchBlock load, alert schema ownerData owner + engineering
DQ check failsdbt test failure / assertion errorBlock load for P1 checks; alert for P2Data engineering
Partial loadRow count < expected thresholdAlert; do not publish to consumers until resolvedData engineering

Dead-letter queue: [Where failed records are routed for manual inspection — e.g. raw.dlq_events]


9. Monitoring & Observability

Metrics to track:

  • Pipeline run duration (p50, p95)
  • Rows processed per run
  • DQ check pass rate
  • Source freshness lag
  • Error rate per source

Alerting:

  • [Slack channel: #data-alerts]
  • [PagerDuty: data-on-call escalation for P1 SLA breaches]
  • [Dashboard: [link to monitoring dashboard]]

Logging: [What gets logged and where — e.g. Airflow task logs to CloudWatch, structured JSON to data lake]


10. Dependencies & Sequencing

Upstream dependencies: [Which pipelines or data sources must succeed before this pipeline runs?]

Downstream dependents: [Which dashboards, pipelines, or models depend on this pipeline's output?]

[upstream pipeline A] ──► THIS PIPELINE ──► [downstream dashboard B]
                                          └──► [downstream pipeline C]

Coordination mechanism: [Airflow DAG dependency / dbt ref() / event trigger / manual gate]


11. Security & Compliance

  • PII fields: [List columns containing PII — e.g. email, ip_address, name]
  • Masking / Pseudonymisation: [e.g. email hashed with SHA-256 before landing in mart layer]
  • Access control: [Who can query the destination tables? — e.g. Role-based access in Snowflake]
  • Data residency: [Which regions is data permitted to transit and rest in?]
  • Audit trail: [Is pipeline execution auditable for compliance purposes? Where are logs retained?]

Quality Checks

  • Every source has an incremental key or full-refresh justification
  • Business logic rules are documented, not just the SQL
  • SLAs are agreed with consumers, not set unilaterally by engineering
  • DQ checks cover completeness, uniqueness, freshness, and volume
  • Failure modes include a documented recovery owner
  • PII fields are identified and a treatment plan is specified

Anti-Patterns

  • Do not spec a pipeline without defining SLAs — "as fast as possible" is not an acceptable freshness target
  • Do not omit error handling and dead-letter queue strategy — every pipeline must specify what happens to failed records
  • Do not design idempotent loads without documenting the deduplication key — assume reruns will happen
  • Do not leave data quality rules implicit — schema validation, null checks, and referential integrity must be explicit
  • Do not ignore schema evolution — specify how upstream schema changes are detected and handled

Example Trigger Phrases

  • "Design a data pipeline for our Salesforce to Snowflake sync"
  • "Write a pipeline spec for ingesting Stripe events into our data warehouse"
  • "Build an ETL spec for our user activity data"
  • "Document our dbt pipeline from raw events to the analytics mart"
  • "Spec out the pipeline that feeds the executive dashboard"

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/data-pipeline-spec of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Data Pipeline Spec next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Pipeline Spec compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Pipeline Spec this skillmohitagw15856/pm-claude-skills1.4k—~2.5kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Glue 09 10 Migrationaws-samples/aws-glue-samples1.5k—~2.4kAutomated safety check: PassMIT-0
Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-0
Dbt Databricks PR Readydatabricks/dbt-databricks380—~2.8kAutomated safety check: PassApache-2.0
Apache Spark EngineerJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Glue 09 10 Migration

    aws-samples/aws-glue-samples

    Official

    Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.

    1.5k GitHub stars~2.4k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Official

    Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

    1.5k GitHub stars~3.6k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Dbt Databricks PR Ready

    databricks/dbt-databricks

    Official

    A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.

    380 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Mz Dbt Release

    MaterializeInc/materialize

    Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.

    6.4k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Data Pipeline Spec

What does Data Pipeline Spec do?

Design an ETL/ELT data pipeline specification. An agent skill from mohitagw15856/pm-claude-skills. Data Pipeline Spec is an agent skill from mohitagw15856/pm-claude-skills. Design an ETL/ELT data pipeline specification.

When should I use Data Pipeline Spec?

Data Pipeline Spec fits situations like: asked to design a data pipeline; document a data ingestion workflow; plan a data integration.

How do I install Data Pipeline Spec in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill data-pipeline-spec -a claude-code`. Or copy the skill folder (skills/data-pipeline-spec in mohitagw15856/pm-claude-skills) into .claude/skills/data-pipeline-spec in your project. Claude Code loads it when a task matches its description.

How do I install Data Pipeline Spec in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill data-pipeline-spec -a codex`. Or copy the skill folder (skills/data-pipeline-spec in mohitagw15856/pm-claude-skills) into .agents/skills/data-pipeline-spec in your project. Codex loads it when a task matches its description.

Can I use Data Pipeline Spec in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill data-pipeline-spec -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-pipeline-spec, .gemini/skills/data-pipeline-spec, .github/skills/data-pipeline-spec and .opencode/skills/data-pipeline-spec in your project.

What does Data Pipeline Spec need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Pipeline Spec is instructions for the agent only.

Does Data Pipeline Spec access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Pipeline Spec safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Pipeline Spec use?

Data Pipeline Spec is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Pipeline Spec use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Pipeline Spec?

Skills that share tags, products or a category with Data Pipeline Spec: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Pipeline Spec?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.