Agent skill

Windmill Data Pipeline Author

by windmill-labs in windmill-labs/windmill

Builds Windmill data pipelines as independent, pipeline-annotated scripts forming a DAG over shared storage, defaulting to DuckDB nodes that materialize into DuckLake tables.

Custom licenceAuto-check passedData & Analytics

Install Windmill Data Pipeline Author

skills CLI
$ npx skills add windmill-labs/windmill --skill write-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install windmill-labs/windmill write-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/system_prompts/auto-generated/skills/write-pipeline .claude/skills/write-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
write-pipeline
GitHub stars
18k
Token cost
~4k tokens
SKILL.md length
2,083 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
Custom licence

At a glance

Builds Windmill data pipelines as independent, pipeline-annotated scripts forming a DAG over shared storage, defaulting to DuckDB nodes that materialize into DuckLake tables.

  • Works in 5 steps: Put every node in the same folder: f//.… → Write each node as its own script.… → Start each body with // pipeline, then… → …
  • Building a multi-step data pipeline out of independent scripts
  • SKILL.md covers Default to DuckDB + DuckLake, Ingesting from Python /…, Storage prerequisites and What makes a script a pipeline…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill draws a firm line between a flow, which is one runnable orchestrating its own steps, and a data pipeline, which is a set of separately deployed scripts wired together by on and materialize annotations and visualized at /pipeline/<folder>. When a request calls for ingesting, transforming or materializing data across steps, it builds pipeline-annotated scripts rather than a flow.

Its default is a duckdb node per table-producing step, writing a bare SELECT with a materialize-ducklake comment so the runtime handles the write into DuckLake, the lakehouse store the pipeline editor assumes. PostgreSQL is reserved for row-level, transactional mutations against an existing table, and bun or python3 nodes are reserved for non-tabular glue work such as calling an external API, landing any resulting tabular output back in DuckLake rather than a separate store. A DuckLake pipeline needs object storage and a configured DuckLake catalog before it can materialize or read anything, checked with a list command.

When your agent uses it

  • Building a multi-step data pipeline out of independent scripts
  • Deciding whether a step belongs in DuckDB, Postgres or a glue script
  • Diagnosing a pipeline that fails because DuckLake storage is not configured

Example prompts

  • “Build a pipeline that ingests the orders CSV and materializes it as a DuckLake table.”
  • “Add a step to this pipeline that calls the weather API and lands the result in DuckLake.”
  • “Check which DuckLake catalogs this workspace has before I build a new pipeline.”

Requirements

  • A Windmill workspace with object storage and a DuckLake catalog configured

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Put every node in the same folder: f//. The folder is the pipeline.
  2. Write each node as its own script. Default to duckdb materializing into DuckLake (see "Default to DuckDB + DuckLake" above); pick…
  3. Start each body with // pipeline, then the // on input declarations, then the transform that writes the output.
  4. Chain nodes by asset URI: read an upstream node's output asset, then // on in the downstream node so the edge forms. Reuse exact asset…
  5. Don't deploy nodes unless the user asks to. A pipeline only "runs" once its scripts are deployed and their triggers exist.

What it can do on your machine

Read from SKILL.md and the folder at commit fb22e5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql, python and typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Windmill Data Pipeline Author loads about 4k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 2,083 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 2,083 words (~3,996 tokens).

“A data pipeline is NOT a flow. A flow is one runnable that orchestrates steps internally. A data pipeline is a set of independent scripts, each deployed on its own, that form a DAG by reading and writing shared storage…”

— opening of SKILL.md by windmill-labs, Custom licence
name
write-pipeline

Read the full SKILL.md on GitHub

Files

Just SKILL.md in system_prompts/auto-generated/skills/write-pipeline of windmill-labs/windmill.

Open the folder on GitHubat commit fb22e5c

Compare with similar skills

Windmill Data Pipeline Author next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Windmill Data Pipeline Author compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Windmill Data Pipeline Author this skillwindmill-labs/windmill18k—~4kAutomated safety check: PassCustom licence
Mz Dbt ReleaseMaterializeInc/materialize6.4k—~1.2kAutomated safety check: PassCustom licence
Dinobase Business Data Querieskappa90/dinobase263—~1.5kAutomated safety check: PassCustom licence
Create Examplegodatadriven/whirl205—~1.1kAutomated safety check: PassApache-2.0
Rocky Pocrocky-data/rocky304—~1.7kAutomated safety check: PassApache-2.0
Altimate Data Warehouse DelegateAltimateAI/data-engineering-skills128—~1.4kAutomated safety check: PassMIT

Similar skills

  • Mz Dbt Release

    MaterializeInc/materialize

    Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.

    6.4k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.

    263 GitHub stars~1.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Create Example

    godatadriven/whirl

    Create a new Whirl example project in the examples/ directory.

    205 GitHub stars~1.1k tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Rocky Poc

    rocky-data/rocky

    Authoring a new POC under examples/playground/pocs/. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.7k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Altimate Data Warehouse Delegate

    AltimateAI/data-engineering-skills

    Delegates dbt and warehouse tasks such as lineage, migrations and cost attribution to the altimate-code CLI agent and relays its answer back.

    128 GitHub stars~1.4k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Cocoindex

    davila7/claude-code-templates

    Comprehensive toolkit for developing with the CocoIndex library.

    32k GitHub starsUsed in 2 repos~6.4k tokens
    AI & LLM EngineeringAuto-check: notes

More from windmill-labs/windmill

All 42 skills in this repo
  • Windmill Trigger Type Checklist

    windmill-labs/windmill

    Checklist of every backend, frontend, CLI and capture change needed to add a new TriggerCrud-based trigger type, such as Azure, GCP or Kafka, to Windmill.

    18k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated today
    Auto-check: notes
  • Domain Modeling and Glossary

    windmill-labs/windmill

    Actively challenges vague or conflicting terminology as you design, and keeps a living domain glossary file up to date in real time.

    18k GitHub stars~622 tokensUpdated today
    Auto-check passed
  • Local PR Review

    windmill-labs/windmill

    Runs the same code review locally that GitHub's auto-review actions run on a PR, delegating to a fresh-context subagent so the review isn't biased by the main session's own reasoning.

    18k GitHub stars~995 tokensUpdated today
    Auto-check passed
  • Draft PR with CI Rounds

    windmill-labs/windmill

    Opens a draft GitHub pull request with a conventional title and explicit body, then drives CI review rounds before marking it ready.

    18k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Windmill Dev Page Preview

    windmill-labs/windmill

    Opens the Windmill dev page to preview a flow, script or app, choosing between proxy and direct mode and deciding whether the agent or the runtime starts the wmill dev server.

    18k GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Windmill Data Pipeline Author

What does Windmill Data Pipeline Author do?

Builds Windmill data pipelines as independent, pipeline-annotated scripts forming a DAG over shared storage, defaulting to DuckDB nodes that materialize into DuckLake tables. The skill draws a firm line between a flow, which is one runnable orchestrating its own steps, and a data pipeline, which is a set of separately deployed scripts wired together by on and materialize annotations and visualized at /pipeline/<folder>. When a request calls for ingesting, transforming or materializing data across steps, it builds pipeline-annotated scripts rather than a flow.

When should I use Windmill Data Pipeline Author?

Windmill Data Pipeline Author fits situations like: building a multi-step data pipeline out of independent scripts; deciding whether a step belongs in DuckDB, Postgres or a glue script; diagnosing a pipeline that fails because DuckLake storage is not configured.

How do I install Windmill Data Pipeline Author in Claude Code?

Run `npx skills add windmill-labs/windmill --skill write-pipeline -a claude-code`. Or copy the skill folder (system_prompts/auto-generated/skills/write-pipeline in windmill-labs/windmill) into .claude/skills/write-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Windmill Data Pipeline Author in Codex?

Run `npx skills add windmill-labs/windmill --skill write-pipeline -a codex`. Or copy the skill folder (system_prompts/auto-generated/skills/write-pipeline in windmill-labs/windmill) into .agents/skills/write-pipeline in your project. Codex loads it when a task matches its description.

Can I use Windmill Data Pipeline Author in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add windmill-labs/windmill --skill write-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/write-pipeline, .gemini/skills/write-pipeline, .github/skills/write-pipeline and .opencode/skills/write-pipeline in your project.

What does Windmill Data Pipeline Author need to run?

SKILL.md names no scripts, command-line tools or credentials: Windmill Data Pipeline Author is instructions for the agent only. Our summary lists: A Windmill workspace with object storage and a DuckLake catalog configured.

Does Windmill Data Pipeline Author access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Windmill Data Pipeline Author safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Windmill Data Pipeline Author use?

Windmill Data Pipeline Author has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Windmill Data Pipeline Author use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Windmill Data Pipeline Author?

Skills that share tags, products or a category with Windmill Data Pipeline Author: Mz Dbt Release (MaterializeInc/materialize, 6.4k stars), Dinobase Business Data Queries (kappa90/dinobase, 263 stars), Create Example (godatadriven/whirl, 205 stars) and Rocky Poc (rocky-data/rocky, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Windmill Data Pipeline Author?

windmill-labs (a GitHub organization) maintains it in windmill-labs/windmill, which has 18,146 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 9, 2026.

Source: windmill-labs/windmill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.