Agent skill

Add Dataset

by marin-community in marin-community/marin

Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Dataset

skills CLI
$ npx skills add marin-community/marin --skill add-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install marin-community/marin add-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-dataset .claude/skills/add-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-dataset
GitHub stars
3.9k
Token cost
~458 tokens
SKILL.md length
193 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module.

  • Tasks that involve Model hubs and datasets
  • Calls uv

What it does

Add Dataset is an agent skill from marin-community/marin. Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module.

Its SKILL.md is about 460 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with Hugging Face. The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Model hubs and datasets

Example prompts

  • “/add-dataset”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 61bb85c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Dataset loads about 458 tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 193 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~458

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from marin-community/marin at commit 61bb85c, republished under its Apache-2.0 licence (© marin-community). 193 words, ~458 tokens.

Download SKILL.mdSave it as .claude/skills/add-dataset/SKILL.md (or your agent's skills folder).
name
add-dataset
description
Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module.

Register a Hugging Face dataset

Inspect the schema without downloading the full dataset:

sh
uv run lib/marin/tools/get_hf_dataset_schema.py <dataset_name> [options]

For programmatic inspection:

python
from marin.tools.get_hf_dataset_schema import get_schema

schema = get_schema(dataset_name="wikitext", config_name="wikitext-103-v1")

Use repo-managed dependencies. For a one-off inspection without a provisioned environment, add --with datasets --with pyyaml to uv run.

If the result says a config is required, select one of available_configs and retry with --config_name. Add --trust_remote_code only after inspecting the dataset repository and accepting its code-execution boundary. The tool streams; do not replace it with a full dataset download.

If the dataset cannot be found, stop and report the identifier, path, or access failure instead of guessing a replacement.

Choose the text field from the reported schema. Prefer an exact text field, then a field containing text, then another string field. Inspect sample_row to verify the content; it may be empty for some datasets. The result also reports splits, text_field_candidates, and features.

Add a leaf module under experiments/datasets/ using the lazy builders in marin.experiment.data:

  • expose <name>_dataset() for one corpus;
  • expose <name>_datasets() -> dict[str, ...] for a keyed family;
  • for Hugging Face subsets, follow experiments/datasets/nemotron.py and return one keyed handle per subset.

Validate the selected config, splits, text mapping, and one sample before adding tokenization or downstream experiment configuration.

© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/add-dataset of marin-community/marin.

Open the folder on GitHubat commit 61bb85c

Compare with similar skills

Add Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Dataset this skillmarin-community/marin3.9k—~458Automated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Upload Post Imagehuggingface/blog3.5k—~1.1kAutomated safety check: PassNone
Add Archon Modelareal-project/AReaL5.8k—~4.9kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Upload Post Image

    huggingface/blog

    Official

    A skill your agent uses when adding or migrating non-thumbnail images for a Hugging Face Blog post.

    3.5k GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Add Archon Model

    areal-project/AReaL

    Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed

More from marin-community/marin

All 41 skills in this repo
  • Noslop

    marin-community/marin

    Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Use Iris

    marin-community/marin

    Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.

    3.9k GitHub stars~745 tokensUpdated today
    Auto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Auto-check passed
  • Marina Applet

    marin-community/marin

    Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.

    3.9k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Query Finelog

    marin-community/marin

    Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Trace Pulumi Diff

    marin-community/marin

    Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…

    3.9k GitHub stars~663 tokensUpdated today
    Auto-check passed

Works with

Questions about Add Dataset

What does Add Dataset do?

Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module. Add Dataset is an agent skill from marin-community/marin. Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module.

When should I use Add Dataset?

Add Dataset fits situations like: tasks that involve Model hubs and datasets.

How do I install Add Dataset in Claude Code?

Run `npx skills add marin-community/marin --skill add-dataset -a claude-code`. Or copy the skill folder (.agents/skills/add-dataset in marin-community/marin) into .claude/skills/add-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Add Dataset in Codex?

Run `npx skills add marin-community/marin --skill add-dataset -a codex`. Or copy the skill folder (.agents/skills/add-dataset in marin-community/marin) into .agents/skills/add-dataset in your project. Codex loads it when a task matches its description.

Can I use Add Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill add-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-dataset, .gemini/skills/add-dataset, .github/skills/add-dataset and .opencode/skills/add-dataset in your project.

What does Add Dataset need to run?

Going by SKILL.md and its folder, Add Dataset needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Add Dataset access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Dataset use?

Add Dataset is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Dataset use?

About 458 tokens (SKILL.md is roughly 1.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Dataset?

Skills that share tags, products or a category with Add Dataset: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Upload Post Image (huggingface/blog, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Dataset?

marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,920 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 9, 2026.

Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.