Agent skill

ML Dataset Discovery

by OpenLAIR in OpenLAIR/dr-claw

Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.

Custom licenceAuto-check passedAI & LLM Engineering

Install ML Dataset Discovery

skills CLI
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenLAIR/dr-claw dataset-discovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-discovery .claude/skills/dataset-discovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-discovery
GitHub stars
1.2k
Token cost
~741 tokens
SKILL.md length
221 words
Files
3 (incl. scripts)
Skills in repo
36
Repo updated
First seen
Licence
Custom licence

At a glance

Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.

  • Works in 5 steps: SCOPE → SEARCH → PRESENT → …
  • Finding training or benchmark datasets for a new research task
  • SKILL.md covers Overview, Agent Workflow, Workspace Layout and Dependencies
  • Runs Python scripts from its folder; calls python3 and huggingface-cli

What it does

The agent first clarifies the research task, the data modality (image, text, tabular, audio or any), a size preference and a licence preference. It then runs `scripts/search_ml_datasets.py search` with your query across four sources, huggingface, openml, github and papers, with options for sources, a maximum result count (30 by default), modality and workspace folder. It can add Hugging Face's `hub_repo_search` MCP tool for semantic search on top.

Results are shown as a markdown table of name, source, downloads, size, licence, tags and URL, sorted by relevance. For a dataset you pick, a `detail` command writes `metadata.json` and a README, and a `pull` command saves sample rows to `sample.jsonl`; a full download needs your confirmation first. Output goes to `./datasets/discovery/` by default, together with a dated search log. The script needs Python 3.8 or newer and `requests`, plus the `gh` CLI for the GitHub source.

When your agent uses it

  • Finding training or benchmark datasets for a new research task
  • Comparing candidate datasets by size, licence and download counts
  • Previewing sample rows before committing to a full download

Example prompts

  • “Find English text datasets for sentiment analysis with a permissive licence.”
  • “What datasets exist for medical image segmentation? Check Hugging Face and OpenML.”
  • “Show details for huggingface:stanfordnlp/imdb and pull 20 sample rows.”

Requirements

  • Python 3.8 or newer with `requests`
  • The `gh` CLI for the GitHub source
  • Network access to the dataset sources

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. SCOPE
  2. SEARCH
  3. PRESENT
  4. DETAIL
  5. PULL

What it can do on your machine

Read from SKILL.md and the folder at commit d51b64e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Dataset Discovery loads about 741 tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 221 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~741

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 221 words (~741 tokens).

“Search multiple ML dataset sources (HuggingFace Hub, OpenML, GitHub, Semantic Scholar) and return a ranked, deduplicated list of relevant datasets.”

— opening of SKILL.md by OpenLAIR, Custom licence
name
dataset-discovery

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (scripts) in skills/dataset-discovery of OpenLAIR/dr-claw.

  • SKILL.md
  • scripts/search_ml_datasets.py
  • tests/prompts.txt

Open the folder on GitHubat commit d51b64e

Compare with similar skills

ML Dataset Discovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Dataset Discovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Dataset Discovery this skillOpenLAIR/dr-claw1.2k—~741Automated safety check: PassCustom licence
Ideer Daily PaperAI45Lab/iDeer416—~2.3kAutomated safety check: NotesAGPL-3.0
Hugging Face Paper Pageshuggingface/skills11k3 repos~2.3kAutomated safety check: PassApache-2.0
Academic AioAperivue/medsci-skills333—~4.8kAutomated safety check: PassMIT
News Aggregator Skillcclank/news-aggregator-skill1.3k—~2.1kAutomated safety check: PassNone
Morning AIdavepoon/buildwithclaude3.6k—~405Automated safety check: PassMIT

Similar skills

  • Ideer Daily Paper

    AI45Lab/iDeer

    Daily paper/repo digest where YOU are the reader. An agent skill from AI45Lab/iDeer.

    416 GitHub stars~2.3k tokensUpdated 2 mo ago
    Research & ScienceAuto-check: notes
  • Hugging Face Paper Pages

    huggingface/skills

    Official

    Fetches Hugging Face paper pages as markdown and reads paper metadata through the papers API when you share a paper URL, an arXiv link or an arXiv ID.

    11k GitHub starsUsed in 3 repos~2.3k tokens
    Research & ScienceAuto-check passed
  • Academic Aio

    Aperivue/medsci-skills

    A skill your agent uses when a medical AI paper should be found and cited by AI search engines and RAG tools.

    333 GitHub stars~4.8k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • News Aggregator Skill

    cclank/news-aggregator-skill

    Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…

    1.3k GitHub stars~2.1k tokensUpdated 4 mo ago
    Writing & ContentAuto-check passed
  • Morning AI

    davepoon/buildwithclaude

    AI news tracking skill that monitors 80+ entities across 6 free sources (Reddit, HN, GitHub, HuggingFace, arXiv, X/Twitter).

    3.6k GitHub stars~405 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed

More from OpenLAIR/dr-claw

All 36 skills in this repo
  • Analyzes reviewer comments and drafts venue-specific rebuttals for AI and computer science conferences, with an issue board, task list and paper edit plan.

    1.2k GitHub stars~4.9k tokensUpdated 23 days ago
    Auto-check: notes
  • Turns a research paper into a slide deck and, optionally, a narrated demo video, through script, slide generation, text-to-speech and video assembly stages you control.

    1.2k GitHub stars~1.5k tokensUpdated 23 days ago
    Auto-check passed
  • Clusters the latest news-feed results by topic and writes a briefing of research idea seeds with citations, plus a structured seeds file, without crawling new sources.

    1.2k GitHub stars~1.3k tokensUpdated 23 days ago
    Auto-check: notes
  • Gemini Deep Research

    OpenLAIR/dr-claw

    Runs multi-source web research through Google's Gemini Deep Research Agent with a bundled Python script and saves a structured, cited report as files.

    1.2k GitHub stars~996 tokensUpdated 23 days ago
    Auto-check passed
  • Six-phase workflow for writing, revising and adapting grant proposals for NSF, NIH, DOE, DARPA, NASA and China's NSFC, from profiling through simulated peer review.

    1.2k GitHub stars~9.2k tokensUpdated 23 days ago
    Auto-check passed
  • Inno Rclone To Overleaf

    OpenLAIR/dr-claw

    Access Overleaf projects via CLI. An agent skill from OpenLAIR/dr-claw.

    1.2k GitHub stars~1.2k tokensUpdated 23 days ago
    Auto-check passed

Questions about ML Dataset Discovery

What does ML Dataset Discovery do?

Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table. The agent first clarifies the research task, the data modality (image, text, tabular, audio or any), a size preference and a licence preference.py search` with your query across four sources, huggingface, openml, github and papers, with options for sources, a maximum result count (30 by default), modality and workspace folder.

When should I use ML Dataset Discovery?

ML Dataset Discovery fits situations like: finding training or benchmark datasets for a new research task; comparing candidate datasets by size, licence and download counts; previewing sample rows before committing to a full download.

How do I install ML Dataset Discovery in Claude Code?

Run `npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a claude-code`. Or copy the skill folder (skills/dataset-discovery in OpenLAIR/dr-claw) into .claude/skills/dataset-discovery in your project. Claude Code loads it when a task matches its description.

How do I install ML Dataset Discovery in Codex?

Run `npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a codex`. Or copy the skill folder (skills/dataset-discovery in OpenLAIR/dr-claw) into .agents/skills/dataset-discovery in your project. Codex loads it when a task matches its description.

Can I use ML Dataset Discovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-discovery, .gemini/skills/dataset-discovery, .github/skills/dataset-discovery and .opencode/skills/dataset-discovery in your project.

What does ML Dataset Discovery need to run?

Going by SKILL.md and its folder, ML Dataset Discovery needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and huggingface-cli). Our summary lists: Python 3.8 or newer with `requests`; The `gh` CLI for the GitHub source; Network access to the dataset sources.

Does ML Dataset Discovery access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Dataset Discovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Dataset Discovery use?

ML Dataset Discovery has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does ML Dataset Discovery use?

About 741 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Dataset Discovery?

Skills that share tags, products or a category with ML Dataset Discovery: Ideer Daily Paper (AI45Lab/iDeer, 416 stars), Hugging Face Paper Pages (huggingface/skills, 11k stars), Academic Aio (Aperivue/medsci-skills, 333 stars) and News Aggregator Skill (cclank/news-aggregator-skill, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Dataset Discovery?

OpenLAIR (a GitHub organization) maintains it in OpenLAIR/dr-claw, which has 1,155 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 17, 2026.

Source: OpenLAIR/dr-claw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.