Official agent skill

Hugging Face Dataset Viewer

by huggingface in huggingface/skills

Explores Hugging Face datasets through the read-only Dataset Viewer API: list splits, preview and page through rows, search, filter, and fetch parquet links and statistics.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Hugging Face Dataset Viewer

skills CLI
$ npx skills add huggingface/skills --skill huggingface-datasets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huggingface/skills huggingface-datasets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface-datasets .claude/skills/huggingface-datasets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huggingface-datasets
GitHub stars
11k
Used in
3 other repos
Token cost
~1.1k tokens
SKILL.md length
394 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Explores Hugging Face datasets through the read-only Dataset Viewer API: list splits, preview and page through rows, search, filter, and fetch parquet links and statistics.

  • Works in 6 steps: Optionally validate dataset availability… → Resolve config + split with /splits. → Preview with /first-rows. → …
  • Previewing a dataset's splits and first rows before downloading anything
  • SKILL.md covers Core workflow, Defaults, Dataset Viewer and Creating and Uploading Datasets, plus 1 more section
  • Calls hf, curl and npx; reaches datasets-server.huggingface.co and huggingface.co; needs HF_TOKEN

What it does

The skill sends read-only GET calls to the Dataset Viewer API at datasets-server.huggingface.co. A typical run checks that the dataset is valid, resolves its config and split, previews the first rows, and then pages through the content with offset and length, where length tops out at 100 for row-style endpoints.

Beyond browsing, it covers text search over string columns, filtering with a where predicate and optional orderby, parquet shard links, size totals, per-column statistics and Croissant metadata when the dataset has it. For partial pages it follows response fields such as num_rows_total and partial. Gated or private datasets need a bearer token taken from HF_TOKEN.

When your agent uses it

  • Previewing a dataset's splits and first rows before downloading anything
  • Searching text columns or filtering rows in a hosted dataset
  • Getting parquet URLs and size statistics for a Hugging Face dataset

Example prompts

  • “List the configs and splits for stanfordnlp/imdb and show me the first rows.”
  • “Search the train split of that dataset for rows mentioning refunds.”
  • “Give me the parquet links and column statistics for this dataset.”

Requirements

  • Network access to datasets-server.huggingface.co
  • An HF_TOKEN for gated or private datasets

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Optionally validate dataset availability with /is-valid.
  2. Resolve config + split with /splits.
  3. Preview with /first-rows.
  4. Paginate content with /rows using offset and length (max 100).
  5. Use /search for text matching and /filter for row predicates.
  6. Retrieve parquet links via /parquet and totals/metadata via /size and /statistics.

What it can do on your machine

Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • hf
    • curl
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • datasets-server.huggingface.co
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hugging Face Dataset Viewer loads about 1.1k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 394 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 394 words, ~1,141 tokens.

Download SKILL.mdSave it as .claude/skills/huggingface-datasets/SKILL.md (or your agent's skills folder).
name
huggingface-datasets
description
Use this skill for Hugging Face Dataset Viewer API workflows that fetch subset/split metadata, paginate rows, search text, apply filters, download parquet URLs, and read size or statistics.

Hugging Face Dataset Viewer

Use this skill to execute read-only Dataset Viewer API calls for dataset exploration and extraction.

Core workflow

  1. Optionally validate dataset availability with /is-valid.
  2. Resolve config + split with /splits.
  3. Preview with /first-rows.
  4. Paginate content with /rows using offset and length (max 100).
  5. Use /search for text matching and /filter for row predicates.
  6. Retrieve parquet links via /parquet and totals/metadata via /size and /statistics.

Defaults

  • Base URL: https://datasets-server.huggingface.co
  • Default API method: GET
  • Query params should be URL-encoded.
  • offset is 0-based.
  • length max is usually 100 for row-like endpoints.
  • Gated/private datasets require Authorization: Bearer <HF_TOKEN>.

Dataset Viewer

  • Validate dataset: /is-valid?dataset=<namespace/repo>
  • List subsets and splits: /splits?dataset=<namespace/repo>
  • Preview first rows: /first-rows?dataset=<namespace/repo>&config=<config>&split=<split>
  • Paginate rows: /rows?dataset=<namespace/repo>&config=<config>&split=<split>&offset=<int>&length=<int>
  • Search text: /search?dataset=<namespace/repo>&config=<config>&split=<split>&query=<text>&offset=<int>&length=<int>
  • Filter with predicates: /filter?dataset=<namespace/repo>&config=<config>&split=<split>&where=<predicate>&orderby=<sort>&offset=<int>&length=<int>
  • List parquet shards: /parquet?dataset=<namespace/repo>
  • Get size totals: /size?dataset=<namespace/repo>
  • Get column statistics: /statistics?dataset=<namespace/repo>&config=<config>&split=<split>
  • Get Croissant metadata (if available): /croissant?dataset=<namespace/repo>

Pagination pattern:

bash
curl "https://datasets-server.huggingface.co/rows?dataset=stanfordnlp/imdb&config=plain_text&split=train&offset=0&length=100"
curl "https://datasets-server.huggingface.co/rows?dataset=stanfordnlp/imdb&config=plain_text&split=train&offset=100&length=100"

When pagination is partial, use response fields such as num_rows_total, num_rows_per_page, and partial to drive continuation logic.

Search/filter notes:

  • /search matches string columns (full-text style behavior is internal to the API).
  • /filter requires predicate syntax in where and optional sort in orderby.
  • Keep filtering and searches read-only and side-effect free.

For CLI-based parquet URL discovery or SQL, use the hf-cli skill with hf datasets parquet and hf datasets sql.

Show full SKILL.md (179 more words)Show less

Creating and Uploading Datasets

Use one of these flows depending on dependency constraints.

Zero local dependencies (Hub UI):

  • Create dataset repo in browser: https://huggingface.co/new-dataset
  • Upload parquet files in the repo "Files and versions" page.
  • Verify shards appear in Dataset Viewer:
bash
curl -s "https://datasets-server.huggingface.co/parquet?dataset=<namespace>/<repo>"

Low dependency CLI flow (npx @huggingface/hub / hfjs):

  • Set auth token:
bash
export HF_TOKEN=<your_hf_token>
  • Upload parquet folder to a dataset repo (auto-creates repo if missing):
bash
npx -y @huggingface/hub upload datasets/<namespace>/<repo> ./local/parquet-folder data
  • Upload as private repo on creation:
bash
npx -y @huggingface/hub upload datasets/<namespace>/<repo> ./local/parquet-folder data --private

After upload, call /parquet to discover <config>/<split>/<shard> values for querying with @~parquet.

Agent Traces

The Hub supports raw agent session traces from Claude Code, Codex, and Pi Agent. Upload them to Hugging Face Datasets as original JSONL files and the Hub can auto-detect the trace format, tag the dataset as Traces, and enable the trace viewer for browsing sessions, turns, tool calls, and model responses. Common local session directories:

  • Claude Code: ~/.claude/projects
  • Codex: ~/.codex/sessions
  • Pi: ~/.pi/agent/sessions

Default to private dataset repos because traces can contain prompts, file paths, tool outputs, secrets, or PII. Preserve the raw .jsonl files and nest them by project/cwd instead of uploading every session at the dataset root.

bash
hf repos create <namespace>/<repo> --type dataset --private --exist-ok
hf upload <namespace>/<repo> ~/.codex/sessions codex/<project-or-cwd> --type dataset

© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/huggingface-datasets of huggingface/skills.

Open the folder on GitHubat commit ca0325b

Used in 3 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hugging Face Dataset Viewer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hugging Face Dataset Viewer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hugging Face Dataset Viewer this skillhuggingface/skills11k3 repos~1.1kAutomated safety check: PassApache-2.0
Python Data AnalysisA-EVO-Lab/a-evolve805—~476Automated safety check: PassNone
Veomni New ModelByteDance-Seed/VeOmni2.2k—~2kAutomated safety check: PassApache-2.0
Dataset FinderLeoYeAI/openclaw-master-skills2.2k—~5.4kAutomated safety check: PassProprietary
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Upload Post Imagehuggingface/blog3.5k—~1.1kAutomated safety check: PassNone

Similar skills

  • Python Data Analysis

    A-EVO-Lab/a-evolve

    Best practices for multi-step Python tasks including data analysis, HuggingFace datasets, token counting, and any task requiring state across multiple python() calls.

    805 GitHub stars~476 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Veomni New Model

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding support for a new model to VeOmni.

    2.2k GitHub stars~2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Finder

    LeoYeAI/openclaw-master-skills

    A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.

    2.2k GitHub stars~5.4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Upload Post Image

    huggingface/blog

    Official

    A skill your agent uses when adding or migrating non-thumbnail images for a Hugging Face Blog post.

    3.5k GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed

More from huggingface/skills

All 25 skills in this repo
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    Auto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Works with

Questions about Hugging Face Dataset Viewer

What does Hugging Face Dataset Viewer do?

Explores Hugging Face datasets through the read-only Dataset Viewer API: list splits, preview and page through rows, search, filter, and fetch parquet links and statistics. co. A typical run checks that the dataset is valid, resolves its config and split, previews the first rows, and then pages through the content with offset and length, where length tops out at 100 for row-style endpoints.

When should I use Hugging Face Dataset Viewer?

Hugging Face Dataset Viewer fits situations like: previewing a dataset's splits and first rows before downloading anything; searching text columns or filtering rows in a hosted dataset; getting parquet URLs and size statistics for a Hugging Face dataset.

How do I install Hugging Face Dataset Viewer in Claude Code?

Run `npx skills add huggingface/skills --skill huggingface-datasets -a claude-code`. Or copy the skill folder (skills/huggingface-datasets in huggingface/skills) into .claude/skills/huggingface-datasets in your project. Claude Code loads it when a task matches its description.

How do I install Hugging Face Dataset Viewer in Codex?

Run `npx skills add huggingface/skills --skill huggingface-datasets -a codex`. Or copy the skill folder (skills/huggingface-datasets in huggingface/skills) into .agents/skills/huggingface-datasets in your project. Codex loads it when a task matches its description.

Can I use Hugging Face Dataset Viewer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill huggingface-datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-datasets, .gemini/skills/huggingface-datasets, .github/skills/huggingface-datasets and .opencode/skills/huggingface-datasets in your project.

What does Hugging Face Dataset Viewer need to run?

Going by SKILL.md and its folder, Hugging Face Dataset Viewer needs the command-line tools its instructions call (hf, curl and npx) and credentials named HF_TOKEN. Our summary lists: Network access to datasets-server.huggingface.co; An HF_TOKEN for gated or private datasets.

Does Hugging Face Dataset Viewer access the network?

SKILL.md names 2 domains. In commands or code: datasets-server.huggingface.co and huggingface.co; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Hugging Face Dataset Viewer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hugging Face Dataset Viewer use?

Hugging Face Dataset Viewer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hugging Face Dataset Viewer use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hugging Face Dataset Viewer?

Skills that share tags, products or a category with Hugging Face Dataset Viewer: Python Data Analysis (A-EVO-Lab/a-evolve, 805 stars), Veomni New Model (ByteDance-Seed/VeOmni, 2.2k stars), Dataset Finder (LeoYeAI/openclaw-master-skills, 2.2k stars) and LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hugging Face Dataset Viewer?

huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,142 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.

Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.