A skill your agent uses whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure…
Install the "datachain-knowledge" agent skill from https://github.com/datachain-ai/datachain/tree/main/src/datachain/skill/knowledge into .claude/skills/datachain-knowledge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datachain-knowledge", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add datachain-ai/datachain --skill datachain-knowledge -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "datachain-knowledge" agent skill from https://github.com/datachain-ai/datachain/tree/main/src/datachain/skill/knowledge into .agents/skills/datachain-knowledge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datachain-knowledge", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add datachain-ai/datachain --skill datachain-knowledge -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "datachain-knowledge" agent skill from https://github.com/datachain-ai/datachain/tree/main/src/datachain/skill/knowledge into .cursor/skills/datachain-knowledge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datachain-knowledge", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add datachain-ai/datachain --skill datachain-knowledge -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "datachain-knowledge" agent skill from https://github.com/datachain-ai/datachain/tree/main/src/datachain/skill/knowledge into .gemini/skills/datachain-knowledge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datachain-knowledge", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add datachain-ai/datachain --skill datachain-knowledge -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "datachain-knowledge" agent skill from https://github.com/datachain-ai/datachain/tree/main/src/datachain/skill/knowledge into .github/skills/datachain-knowledge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datachain-knowledge", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add datachain-ai/datachain --skill datachain-knowledge -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "datachain-knowledge" agent skill from https://github.com/datachain-ai/datachain/tree/main/src/datachain/skill/knowledge into .opencode/skills/datachain-knowledge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datachain-knowledge", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
datachain-knowledge
GitHub stars
2.8k
Token cost
~3k tokens
SKILL.md length
1,356 words
Files
21 (incl. scripts)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0
At a glance
A skill your agent uses whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure…
Works in 7 steps: Bucket Enlistment → Sync → Save Data → …
Cloud storage buckets
SKILL.md covers Critical Rules, Common gotchas in UDF scripts, Running the work and Workflow Mode Detection, plus 7 more sections
Runs Python scripts from its folder; calls python3
What it does
Datachain Knowledge is an agent skill from datachain-ai/datachain. Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts (for example `__init__.py`, `collect.py` and `prompts/enrich.md`).
It sits in Knowledge Management, covering Knowledge bases, File uploads and storage and Data pipelines and ETL. It works with Microsoft Azure and Pydantic. The repository describes itself as: The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure. The licence is Apache-2.0.
When your agent uses it
Cloud storage buckets
Data pipelines are mentioned — creating
Processing data in S3
Running any script that may create datasets as a side effect
Example prompts
“/datachain-knowledge”
Requirements
Python 3
Workflow steps
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cbb008d. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Ships 14 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python3
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Datachain Knowledge loads about 3k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,356 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~122
When it runs· the whole SKILL.md, loaded when a task matches
~3k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Download SKILL.mdSave it as .claude/skills/datachain-knowledge/SKILL.md (or your agent's skills folder). This skill also uses 20 other files; get the full folder from GitHub.
name
datachain-knowledge
description
Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.
triggers
what datasets exist, show me the schema, list datasets, datachain knowledge, update the knowledge base, refresh dataset docs, what's in this bucket, explore…
Maintain a knowledge base at dc-knowledge/. .md files are the persistent
output. .json files are intermediate (generated in Step 3, consumed in
Step 4, then deleted).
{core_skill_dir}/SDK.md owns how the pipeline code is written — dataset
shape, row grain, provenance, naming. This file owns the knowledge base
and how the work is run: what already exists, what a run will cost, and
where the results land.
Critical Rules
Path is dc-knowledge/ — NOT .datachain/. The .datachain/ directory is the internal database; the knowledge base lives at dc-knowledge/.
Never pass update=True to dc.read_storage() in query or exploration code unless the user explicitly asks to refresh the listing. Build scripts that read storage are the exception — they pass update=True, delta=True.
Prefer DataChain operations over plain Python for all metadata analysis.
Bounded output — JSON and markdown files stay small regardless of data size.
Stop on auth/connection errors — bucket_scan.py runs a fast access check. If it exits with an error JSON on stderr, stop immediately and show the error to the user. Do not retry with different regions, profiles, or endpoints — ask for the missing credentials.
Follow the enrichment prompt template literally in Step 4. Downstream tooling (render_index.py) parses the exact frontmatter the prompt prescribes.
Common gotchas in UDF scripts
parallel=N vs workers=N.parallel=N is local multiprocessing (works anywhere). workers=N is Studio-only and MUST be guarded: chain = chain.settings(parallel=N); if dc.is_studio(): chain = chain.settings(workers=N).
No from __future__ import annotations in UDF modules. It stringifies type hints and DataChain's signal-schema resolution rejects the string-vs-class mismatch.
Type the UDF return precisely.Iterator[object] / Iterator[Any] / bare dict fail schema resolution. Return a specific Iterator[T], a Pydantic BaseModel, or a primitive.
Generators aren't subscriptable. Iterators returned by file APIs do not support [:N]. Use enumerate + break, or list(...) only when the result is genuinely small.
Use datachain.__version__ to get the package version (e.g. dc.__version__).
Running the work
Reuse before building
Read dc-knowledge/index.md first. When an existing dataset covers the task —
even partially — read it with dc.read_dataset(...) and filter / merge / extend
from there instead of going back to raw storage. Re-running a pass that already
ran is the most expensive mistake available here. Say which dataset was reused
and what it saved.
Save what was expensive
A UDF that ran a model, decoded file bodies, or called a paid API produces rows
worth keeping: save that operation's full output, unfiltered, under a
descriptive name with a description=. Chains that only list, filter, or select
are cheap to recompute and need no dataset. Cost is the only criterion — there is
no hierarchy of datasets that has to be built.
Estimate before a long run
Quote a number before starting anything that may run for minutes:
wall ≈ files × per-row × 1.5 / parallel
Op class
Per-row
header / metadata parse (bounded-prefix reads)
~1 ms
file-body decode
size / 10-50 MB/s
small CPU model (text, light CV)
5-50 ms
mid CPU model (detection, segmentation)
50-500 ms
streaming CPU model (ASR, audio)
0.1-0.5× realtime
local GPU
10-100× faster than the CPU row
paid API (LLM / VLM)
$0.001-0.01 per row + 0.5-2 s, rate-limited
Measure instead of estimating when the implementation is untested, the model or
library has no row in the table, or files are large enough that decode dominates:
run 3-5 items with .persist() (never .save()), budget 60 s, and extrapolate
wall_full = (wall_sample / N) × total_files × 1.5. Kill at 60 s and fall back to
the estimate.
Label which is which — estimated ~X or measured on N=5: ~X. Never present an
estimate as a measurement, and write not measured literally when nothing was.
Watch the first minutes
For any run estimated over 5 minutes, read the throughput line DataChain prints
(Processed: N rows [elapsed, rate]) over the first 60-90 s:
at or above ~0.66× the expected rate → carry on;
below ~0.5× → kill it, report the gap and the revised estimate;
no throughput line within 2 minutes → kill it and investigate (model download,
auth retry, startup cost).
Results land in datasets
Aggregations and final answers are DataChain chains — .filter(), .group_by(),
.mutate(), .distinct() — ending in .save(). Two bypasses are forbidden:
writing results to .json / .csv / .parquet through open(), json.dump or
pandas.to_csv, and pulling rows out with .to_iter() / .to_list() to walk them
in Python loops and print the answer. Both leave no dataset, no lineage and no KB
record, so the next session recomputes everything. .show() on a saved dataset is
fine. If answering needs a Python loop over nested lists, the row grain is wrong —
see "Row shape" in SDK.md.
One script per stage
A pipeline that produces several datasets is several scripts, each named after the
dataset it produces, each with exactly one .save(). Never batch or shard by hand:
DataChain checkpoints UDF progress, so re-running a killed script resumes where it
stopped.
Show full SKILL.md (590 more words)Show less
Workflow Mode Detection
Mode A — Discovery/Exploration (e.g., "what datasets exist", "show schema", "explore bucket"):
→ If the user references a specific bucket URI, run Step 1 (Bucket Enlistment) for its root first.
→ Then run Steps 2–7.
Mode B — Dataset Creation/Pipeline (e.g., "create dataset X from ...", "process files and save"):
Precondition (do this FIRST — before ANY tool call):
$ cat dc-knowledge/index.md
If index.md exists and the task can be solved by reading an existing
dataset, do not write a pipeline — read it directly with
dc.read_dataset("name") and filter/merge/extend from there. This avoids
recomputing expensive operations.
Never parse files under dc-knowledge/datasets/*.json or
dc-knowledge/buckets/**/*.json directly — those are pre-render
intermediates that get deleted. The information you need is in index.md.
If dc-knowledge/index.md does not exist, proceed with Steps 1–7 to build it.
→ If the pipeline reads from a bucket, run Step 1 (Bucket Enlistment) for the bucket root first.
→ Run the access check (if not already done in Step 1): datachain bucket status <uri>. If not found / denied, stop and ask for credentials.
→ Read {core_skill_dir}/SDK.md for DataChain SDK rules.
→ Work through "Running the work" above — reuse, estimate, then write the script.
→ While the pipeline is running, enrich any Step 1 bucket JSON that does not yet have a .md (parallel work).
→ After the pipeline completes, run Steps 2–7 to update the knowledge base.
→ Report both: pipeline result AND knowledge base update status.
Mode C — Script Execution (e.g., user runs an existing .py file that touches data):
→ If the script references bucket URIs, run Step 1 for each bucket root first.
→ Scripts can create datasets as side effects.
→ While the script is running, enrich Step 1 bucket JSON in parallel.
→ After ANY data-related script finishes, run Steps 2–7 to detect and record new/changed datasets.
Mode D — Knowledge Base Maintenance (e.g., "update the knowledge base", "refresh dataset docs"):
→ Run Steps 2–7. Existing session context in .md files is preserved automatically during re-enrichment.
Step 1 — Bucket Enlistment
When any storage URI is encountered, enlist the whole bucket first.
Extract bucket root. From any URI, derive {scheme}://{bucket}/.
Check if already enlisted. Look for dc-knowledge/buckets/{scheme}/{bucket_slug}.md or .json. If either exists, skip.
Access check. Run datachain bucket status {root_uri}. If denied / not found, stop and ask.
Scan with timeout. Default 60s; user can override:bash
Buckets are auto-discovered from catalog listings. Do not add --studio unless requested. If "up_to_date": true, print "Knowledge base is up to date." and stop. Entries with status of "new" or "stale" need processing in Step 3.
Generate .md from .json for each entry processed in Step 3 (and any Step 1 bucket JSON that lacks a .md).
Datasets: read {skill_dir}/prompts/enrich.md, then write dc-knowledge/<file_path>.md per the template.
Buckets: read {skill_dir}/prompts/enrich_bucket.md, then write dc-knowledge/<file_path>.md.
The prompt template is authoritative — downstream tooling parses the exact frontmatter it prescribes. Skip this step only if the user requests raw output only.
Datachain Knowledge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Datachain Knowledge compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Datachain Knowledge this skilldatachain-ai/datachain
A skill your agent uses whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure…. Datachain Knowledge is an agent skill from datachain-ai/datachain. Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage.
When should I use Datachain Knowledge?
Datachain Knowledge fits situations like: cloud storage buckets; data pipelines are mentioned — creating; processing data in S3; running any script that may create datasets as a side effect.
How do I install Datachain Knowledge in Claude Code?
Run `npx skills add datachain-ai/datachain --skill datachain-knowledge -a claude-code`. Or copy the skill folder (src/datachain/skill/knowledge in datachain-ai/datachain) into .claude/skills/datachain-knowledge in your project. Claude Code loads it when a task matches its description.
How do I install Datachain Knowledge in Codex?
Run `npx skills add datachain-ai/datachain --skill datachain-knowledge -a codex`. Or copy the skill folder (src/datachain/skill/knowledge in datachain-ai/datachain) into .agents/skills/datachain-knowledge in your project. Codex loads it when a task matches its description.
Can I use Datachain Knowledge in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add datachain-ai/datachain --skill datachain-knowledge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datachain-knowledge, .gemini/skills/datachain-knowledge, .github/skills/datachain-knowledge and .opencode/skills/datachain-knowledge in your project.
What does Datachain Knowledge need to run?
Going by SKILL.md and its folder, Datachain Knowledge needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
Does Datachain Knowledge access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is Datachain Knowledge safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
What licence does Datachain Knowledge use?
Datachain Knowledge is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Datachain Knowledge use?
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Datachain Knowledge?
Skills that share tags, products or a category with Datachain Knowledge: Wiki Viewer (rohitg00/pro-workflow, 2.9k stars), Kb Search (Azure/azure-sdk-tools, 134 stars), Update Guidelines (Azure/azure-sdk-tools, 134 stars) and Foundry Iq (microsoft/GitHub-Copilot-for-Azure, 255 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Datachain Knowledge?
datachain-ai (a GitHub organization) maintains it in datachain-ai/datachain, which has 2,823 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.
Source: datachain-ai/datachain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.