Agent skill

Datachain Knowledge

by datachain-ai in datachain-ai/datachain

A skill your agent uses whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure…

Apache-2.0Auto-check passedKnowledge Management

Install Datachain Knowledge

skills CLI
$ npx skills add datachain-ai/datachain --skill datachain-knowledge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install datachain-ai/datachain datachain-knowledge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/datachain-ai/datachain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/datachain/skill/knowledge .claude/skills/datachain-knowledge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datachain-knowledge
GitHub stars
2.8k
Token cost
~3k tokens
SKILL.md length
1,356 words
Files
21 (incl. scripts)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure…

  • Works in 7 steps: Bucket Enlistment → Sync → Save Data → …
  • Cloud storage buckets
  • SKILL.md covers Critical Rules, Common gotchas in UDF scripts, Running the work and Workflow Mode Detection, plus 7 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Datachain Knowledge is an agent skill from datachain-ai/datachain. Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts (for example `__init__.py`, `collect.py` and `prompts/enrich.md`).

It sits in Knowledge Management, covering Knowledge bases, File uploads and storage and Data pipelines and ETL. It works with Microsoft Azure and Pydantic. The repository describes itself as: The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure. The licence is Apache-2.0.

When your agent uses it

  • Cloud storage buckets
  • Data pipelines are mentioned — creating
  • Processing data in S3
  • Running any script that may create datasets as a side effect

Example prompts

  • “/datachain-knowledge”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Bucket Enlistment
  2. Sync
  3. Save Data
  4. Enrich
  5. Build Index
  6. Cleanup
  7. Report

What it can do on your machine

Read from SKILL.md and the folder at commit cbb008d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 14 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datachain Knowledge loads about 3k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,356 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from datachain-ai/datachain at commit cbb008d, republished under its Apache-2.0 licence (© datachain-ai). 1,356 words, ~3,032 tokens.

Download SKILL.mdSave it as .claude/skills/datachain-knowledge/SKILL.md (or your agent's skills folder). This skill also uses 20 other files; get the full folder from GitHub.
name
datachain-knowledge
description
Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.
triggers
what datasets exist, show me the schema, list datasets, datachain knowledge, update the knowledge base, refresh dataset docs, what's in this bucket, explore…

Maintain a knowledge base at dc-knowledge/. .md files are the persistent output. .json files are intermediate (generated in Step 3, consumed in Step 4, then deleted).

{core_skill_dir}/SDK.md owns how the pipeline code is written — dataset shape, row grain, provenance, naming. This file owns the knowledge base and how the work is run: what already exists, what a run will cost, and where the results land.

Critical Rules

  1. Path is dc-knowledge/ — NOT .datachain/. The .datachain/ directory is the internal database; the knowledge base lives at dc-knowledge/.
  2. Never pass update=True to dc.read_storage() in query or exploration code unless the user explicitly asks to refresh the listing. Build scripts that read storage are the exception — they pass update=True, delta=True.
  3. Prefer DataChain operations over plain Python for all metadata analysis.
  4. Bounded output — JSON and markdown files stay small regardless of data size.
  5. Stop on auth/connection errors — bucket_scan.py runs a fast access check. If it exits with an error JSON on stderr, stop immediately and show the error to the user. Do not retry with different regions, profiles, or endpoints — ask for the missing credentials.
  6. Follow the enrichment prompt template literally in Step 4. Downstream tooling (render_index.py) parses the exact frontmatter the prompt prescribes.

Common gotchas in UDF scripts

  • parallel=N vs workers=N. parallel=N is local multiprocessing (works anywhere). workers=N is Studio-only and MUST be guarded: chain = chain.settings(parallel=N); if dc.is_studio(): chain = chain.settings(workers=N).
  • No from __future__ import annotations in UDF modules. It stringifies type hints and DataChain's signal-schema resolution rejects the string-vs-class mismatch.
  • Type the UDF return precisely. Iterator[object] / Iterator[Any] / bare dict fail schema resolution. Return a specific Iterator[T], a Pydantic BaseModel, or a primitive.
  • Generators aren't subscriptable. Iterators returned by file APIs do not support [:N]. Use enumerate + break, or list(...) only when the result is genuinely small.
  • Use datachain.__version__ to get the package version (e.g. dc.__version__).

Running the work

Reuse before building

Read dc-knowledge/index.md first. When an existing dataset covers the task — even partially — read it with dc.read_dataset(...) and filter / merge / extend from there instead of going back to raw storage. Re-running a pass that already ran is the most expensive mistake available here. Say which dataset was reused and what it saved.

Save what was expensive

A UDF that ran a model, decoded file bodies, or called a paid API produces rows worth keeping: save that operation's full output, unfiltered, under a descriptive name with a description=. Chains that only list, filter, or select are cheap to recompute and need no dataset. Cost is the only criterion — there is no hierarchy of datasets that has to be built.

Estimate before a long run

Quote a number before starting anything that may run for minutes:

wall ≈ files × per-row × 1.5 / parallel
Op classPer-row
header / metadata parse (bounded-prefix reads)~1 ms
file-body decodesize / 10-50 MB/s
small CPU model (text, light CV)5-50 ms
mid CPU model (detection, segmentation)50-500 ms
streaming CPU model (ASR, audio)0.1-0.5× realtime
local GPU10-100× faster than the CPU row
paid API (LLM / VLM)$0.001-0.01 per row + 0.5-2 s, rate-limited

Measure instead of estimating when the implementation is untested, the model or library has no row in the table, or files are large enough that decode dominates: run 3-5 items with .persist() (never .save()), budget 60 s, and extrapolate wall_full = (wall_sample / N) × total_files × 1.5. Kill at 60 s and fall back to the estimate.

Label which is which — estimated ~X or measured on N=5: ~X. Never present an estimate as a measurement, and write not measured literally when nothing was.

Watch the first minutes

For any run estimated over 5 minutes, read the throughput line DataChain prints (Processed: N rows [elapsed, rate]) over the first 60-90 s:

  • at or above ~0.66× the expected rate → carry on;
  • below ~0.5× → kill it, report the gap and the revised estimate;
  • no throughput line within 2 minutes → kill it and investigate (model download, auth retry, startup cost).
Results land in datasets

Aggregations and final answers are DataChain chains — .filter(), .group_by(), .mutate(), .distinct() — ending in .save(). Two bypasses are forbidden: writing results to .json / .csv / .parquet through open(), json.dump or pandas.to_csv, and pulling rows out with .to_iter() / .to_list() to walk them in Python loops and print the answer. Both leave no dataset, no lineage and no KB record, so the next session recomputes everything. .show() on a saved dataset is fine. If answering needs a Python loop over nested lists, the row grain is wrong — see "Row shape" in SDK.md.

One script per stage

A pipeline that produces several datasets is several scripts, each named after the dataset it produces, each with exactly one .save(). Never batch or shard by hand: DataChain checkpoints UDF progress, so re-running a killed script resumes where it stopped.


Show full SKILL.md (590 more words)Show less

Workflow Mode Detection

Mode A — Discovery/Exploration (e.g., "what datasets exist", "show schema", "explore bucket"): → If the user references a specific bucket URI, run Step 1 (Bucket Enlistment) for its root first. → Then run Steps 2–7.

Mode B — Dataset Creation/Pipeline (e.g., "create dataset X from ...", "process files and save"):

Precondition (do this FIRST — before ANY tool call):

$ cat dc-knowledge/index.md

If index.md exists and the task can be solved by reading an existing dataset, do not write a pipeline — read it directly with dc.read_dataset("name") and filter/merge/extend from there. This avoids recomputing expensive operations.

Never parse files under dc-knowledge/datasets/*.json or dc-knowledge/buckets/**/*.json directly — those are pre-render intermediates that get deleted. The information you need is in index.md.

If dc-knowledge/index.md does not exist, proceed with Steps 1–7 to build it.

→ If the pipeline reads from a bucket, run Step 1 (Bucket Enlistment) for the bucket root first. → Run the access check (if not already done in Step 1): datachain bucket status <uri>. If not found / denied, stop and ask for credentials. → Read {core_skill_dir}/SDK.md for DataChain SDK rules. → Work through "Running the work" above — reuse, estimate, then write the script. → While the pipeline is running, enrich any Step 1 bucket JSON that does not yet have a .md (parallel work). → After the pipeline completes, run Steps 2–7 to update the knowledge base. → Report both: pipeline result AND knowledge base update status.

Mode C — Script Execution (e.g., user runs an existing .py file that touches data): → If the script references bucket URIs, run Step 1 for each bucket root first. → Scripts can create datasets as side effects. → While the script is running, enrich Step 1 bucket JSON in parallel. → After ANY data-related script finishes, run Steps 2–7 to detect and record new/changed datasets.

Mode D — Knowledge Base Maintenance (e.g., "update the knowledge base", "refresh dataset docs"): → Run Steps 2–7. Existing session context in .md files is preserved automatically during re-enrichment.


Step 1 — Bucket Enlistment

When any storage URI is encountered, enlist the whole bucket first.

  1. Extract bucket root. From any URI, derive {scheme}://{bucket}/.
  2. Check if already enlisted. Look for dc-knowledge/buckets/{scheme}/{bucket_slug}.md or .json. If either exists, skip.
  3. Access check. Run datachain bucket status {root_uri}. If denied / not found, stop and ask.
  4. Scan with timeout. Default 60s; user can override:
    bash
    python3 {skill_dir}/scripts/bucket_scan.py {root_uri} \
      --output dc-knowledge/buckets/{scheme}/{bucket_slug}.json --timeout 60
  5. Handle timeout (exit code 124). Run the hierarchical fallback:
    bash
    python3 {skill_dir}/scripts/bucket_overview.py {root_uri} \
      --bucket-json dc-knowledge/buckets/{scheme}/{bucket_slug}.json
  6. Report. "Enlisted bucket {bucket} — {N} files, total size {size}, primarily {top 2-3 extensions}." Do not enrich here; Step 4 batches it.

Step 1 runs once per bucket root per session.


Step 2 — Sync

bash
python3 {skill_dir}/scripts/plan.py [--studio] --output dc-knowledge/.plan.json

Buckets are auto-discovered from catalog listings. Do not add --studio unless requested. If "up_to_date": true, print "Knowledge base is up to date." and stop. Entries with status of "new" or "stale" need processing in Step 3.


Step 3 — Save Data

For each dataset where status != "ok":

bash
python3 {skill_dir}/scripts/dataset_all.py <name> \
  --plan dc-knowledge/.plan.json --output dc-knowledge/<file_path>.json

For each bucket where status != "ok" (and not enlisted in Step 1):

bash
python3 {skill_dir}/scripts/bucket_scan.py <uri> --output dc-knowledge/<file_path>.json

Run independent calls concurrently.


Step 4 — Enrich

Generate .md from .json for each entry processed in Step 3 (and any Step 1 bucket JSON that lacks a .md).

  • Datasets: read {skill_dir}/prompts/enrich.md, then write dc-knowledge/<file_path>.md per the template.
  • Buckets: read {skill_dir}/prompts/enrich_bucket.md, then write dc-knowledge/<file_path>.md.

The prompt template is authoritative — downstream tooling parses the exact frontmatter it prescribes. Skip this step only if the user requests raw output only.


Step 5 — Build Index

bash
python3 {skill_dir}/scripts/render_index.py --plan dc-knowledge/.plan.json --output dc-knowledge/index.md

Step 6 — Cleanup

bash
python3 {skill_dir}/scripts/cleanup_json.py --plan dc-knowledge/.plan.json

Keeps .plan.json for Step 7. Skip if the user asks to retain JSON for debugging.


Step 7 — Report

Knowledge base updated: <N> datasets (<M> updated, <K> unchanged), <B> buckets (<X> scanned, <Y> unchanged).

If any buckets have listing_expired: true, add:

Warning: Listing for <bucket> is expired (last scanned: <date>). Run dc.read_storage("<uri>", update=True) to refresh.

© datachain-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 20 other files (scripts) in src/datachain/skill/knowledge of datachain-ai/datachain.

  • SKILL.md
  • __init__.py
  • collect.py
  • prompts/enrich.md
  • prompts/enrich_bucket.md
  • scripts/__init__.py
  • scripts/bucket_overview.py
  • scripts/bucket_scan.py
  • scripts/changes.py
  • scripts/cleanup_json.py
  • scripts/dataset.py
  • scripts/dataset_all.py
  • scripts/db_mtime.py
  • scripts/list_datasets.py
  • scripts/plan.py
  • scripts/render_index.py
  • scripts/schema.py
  • scripts/summary.py
  • scripts/utils.py
  • … and 2 more

Open the folder on GitHubat commit cbb008d

Compare with similar skills

Datachain Knowledge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datachain Knowledge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datachain Knowledge this skilldatachain-ai/datachain2.8k—~3kAutomated safety check: PassApache-2.0
Wiki Viewerrohitg00/pro-workflow2.9k—~1.1kAutomated safety check: PassNone
Kb SearchAzure/azure-sdk-tools134—~1.6kAutomated safety check: PassMIT
Update GuidelinesAzure/azure-sdk-tools134—~1.9kAutomated safety check: PassMIT
Foundry Iqmicrosoft/GitHub-Copilot-for-Azure255—~872Automated safety check: PassMIT
Create Environmentgodatadriven/whirl205—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Wiki Viewer

    rohitg00/pro-workflow

    Render a self-contained HTML viewer for a pro-workflow wiki.

    2.9k GitHub stars~1.1k tokensUpdated 8 days ago
    Knowledge ManagementAuto-check passed
  • Kb Search

    Azure/azure-sdk-tools

    Official

    Query the APIView Copilot knowledge base for guidelines, examples, and memories.

    134 GitHub stars~1.6k tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Update Guidelines

    Azure/azure-sdk-tools

    Official

    Ingest guideline changes from the azure-sdk repo into the knowledge base.

    134 GitHub stars~1.9k tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Foundry Iq

    microsoft/GitHub-Copilot-for-Azure

    Official

    Foundry IQ knowledge bases. An agent skill from microsoft/GitHub-Copilot-for-Azure.

    255 GitHub stars~872 tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Create Environment

    godatadriven/whirl

    Create a new Whirl environment in the envs/ directory. An agent skill from godatadriven/whirl.

    205 GitHub stars~1.9k tokensUpdated 6 days ago
    Backend & APIsAuto-check passed
  • Azure Eventhub Dotnet

    microsoft/skills

    Official

    Azure Event Hubs SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 6 repos~2.7k tokens
    Backend & APIsAuto-check passed

More from datachain-ai/datachain

  • Datachain Core

    datachain-ai/datachain

    Use ONLY for abstract DataChain SDK questions — API usage, method signatures, or code patterns — when no specific dataset or bucket is referenced.

    2.8k GitHub stars~423 tokensUpdated today
    Auto-check passed

Questions about Datachain Knowledge

What does Datachain Knowledge do?

A skill your agent uses whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure…. Datachain Knowledge is an agent skill from datachain-ai/datachain. Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage.

When should I use Datachain Knowledge?

Datachain Knowledge fits situations like: cloud storage buckets; data pipelines are mentioned — creating; processing data in S3; running any script that may create datasets as a side effect.

How do I install Datachain Knowledge in Claude Code?

Run `npx skills add datachain-ai/datachain --skill datachain-knowledge -a claude-code`. Or copy the skill folder (src/datachain/skill/knowledge in datachain-ai/datachain) into .claude/skills/datachain-knowledge in your project. Claude Code loads it when a task matches its description.

How do I install Datachain Knowledge in Codex?

Run `npx skills add datachain-ai/datachain --skill datachain-knowledge -a codex`. Or copy the skill folder (src/datachain/skill/knowledge in datachain-ai/datachain) into .agents/skills/datachain-knowledge in your project. Codex loads it when a task matches its description.

Can I use Datachain Knowledge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datachain-ai/datachain --skill datachain-knowledge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datachain-knowledge, .gemini/skills/datachain-knowledge, .github/skills/datachain-knowledge and .opencode/skills/datachain-knowledge in your project.

What does Datachain Knowledge need to run?

Going by SKILL.md and its folder, Datachain Knowledge needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Datachain Knowledge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Datachain Knowledge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Datachain Knowledge use?

Datachain Knowledge is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datachain Knowledge use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datachain Knowledge?

Skills that share tags, products or a category with Datachain Knowledge: Wiki Viewer (rohitg00/pro-workflow, 2.9k stars), Kb Search (Azure/azure-sdk-tools, 134 stars), Update Guidelines (Azure/azure-sdk-tools, 134 stars) and Foundry Iq (microsoft/GitHub-Copilot-for-Azure, 255 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datachain Knowledge?

datachain-ai (a GitHub organization) maintains it in datachain-ai/datachain, which has 2,823 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.

Source: datachain-ai/datachain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.