Agent skill

Dataset Discovery

by LigphiDonk in LigphiDonk/Oh-my--paper

“Multi-source ML dataset discovery.”

— description from SKILL.md by LigphiDonk
MITAuto-check passedResearch & Science

Install Dataset Discovery

skills CLI
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LigphiDonk/Oh-my--paper dataset-discovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-discovery .claude/skills/dataset-discovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-discovery
GitHub stars
739
Token cost
~1.3k tokens
SKILL.md length
390 words
Files
3 (incl. scripts)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

  • Works in 5 steps: SCOPE → SEARCH → PRESENT → …
  • SKILL.md covers Canonical Summary, Trigger Rules, Resource Use Rules and Execution Contract, plus 5 more sections
  • Runs Python scripts from its folder; calls python3 and huggingface-cli

About this skill

Dataset Discovery is a skill in LigphiDonk/Oh-my--paper (739 stars). Its SKILL.md is about 1.3k tokens, with 2 other files in the folder (scripts). Licence: MIT.

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. SCOPE
  2. SEARCH
  3. PRESENT
  4. DETAIL
  5. PULL

What it can do on your machine

Read from SKILL.md and the folder at commit 6baece9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Discovery loads about 1.3k tokens when it runs. Until then it costs about 13 tokens; SKILL.md has 390 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~13
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LigphiDonk/Oh-my--paper at commit 6baece9, republished under its MIT licence (© LigphiDonk). 390 words, ~1,252 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-discovery/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dataset-discovery
description
Multi-source ML dataset discovery.
id
dataset-discovery
version
1.0.0
stages
survey, ideation, experiment
tools
read_file, search_project, write_file, run_terminal
summary
Multi-source ML dataset discovery. Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets relevant to a research task. Use when asked…
primaryIntent
data
intents
data, research
capabilities
search-retrieval, data-processing
domains
data-engineering
keywords
dataset-discovery, resource prep, search-retrieval, data-processing, data-engineering, dataset, discovery, multi, source, ml, search, huggingface

dataset-discovery

Canonical Summary

Multi-source ML dataset discovery. Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets relevant to a research task. Use when asked to "find datasets for", "search ML datasets", "what datasets exist for", or "dis...

Trigger Rules

Use this skill when the user request matches its research workflow scope. Prefer the bundled resources instead of recreating templates or reference material. Keep outputs traceable to project files, citations, scripts, or upstream evidence.

Resource Use Rules

  • Treat scripts/ as optional helpers. Run them only when their dependencies are available, keep outputs in the project workspace, and explain a manual fallback if execution is blocked.

Execution Contract

  • Resolve every relative path from this skill directory first.
  • Prefer inspection before mutation when invoking bundled scripts.
  • If a required runtime, CLI, credential, or API is unavailable, explain the blocker and continue with the best manual fallback instead of silently skipping the step.
  • Do not write generated artifacts back into the skill directory; save them inside the active project workspace.

Upstream Instructions

Dataset Discovery Skill

Overview

Search multiple ML dataset sources (HuggingFace Hub, OpenML, GitHub, Semantic Scholar) and return a ranked, deduplicated list of relevant datasets.

Agent Workflow

Phase 1: SCOPE

Clarify the user's needs before searching:

  • Research task: What problem or domain? (e.g., "sentiment analysis", "medical image segmentation")
  • Modality: image / text / tabular / audio / any
  • Size preference: small (< 10K rows), medium (10K–1M), large (> 1M), any
  • License preference: permissive (MIT/Apache/CC-BY), any, or specific
Show full SKILL.md (151 more words)Show less

Run the search script with the user's query:

bash
python3 scripts/search_ml_datasets.py search --query "<query>" --sources huggingface,openml,github,papers --max 30

Options:

  • --sources: Comma-separated list from huggingface, openml, github, papers. Default: all four.
  • --max: Maximum results to return after dedup + ranking. Default: 30.
  • --modality: Filter by modality (image, text, tabular, audio).
  • --workspace: Output directory. Default: ./datasets/discovery/

Optionally also call HF MCP tool hub_repo_search with repo_types: ["dataset"] for semantic search to supplement results.

Phase 3: PRESENT

Show results as a markdown table:

NameSourceDownloadsSizeLicenseTagsURL

Sort by relevance score (highest first).

Phase 4: DETAIL

When the user wants more info on a specific dataset:

bash
python3 scripts/search_ml_datasets.py detail --dataset-id "huggingface:stanfordnlp/imdb" --workspace ./datasets/discovery/

Writes metadata.json and README.md to {workspace}/datasets/{source}_{slug}/.

Phase 5: PULL

When the user wants to preview data:

bash
python3 scripts/search_ml_datasets.py pull --dataset-id "huggingface:stanfordnlp/imdb" --sample-rows 20 --workspace ./datasets/discovery/

Writes sample.jsonl to {workspace}/datasets/{source}_{slug}/.

For full dataset download, confirm with the user first, then use huggingface-cli download or equivalent.

Workspace Layout

{workspace}/                         # default: ./datasets/discovery/
  search-{YYYY-MM-DD}.json           # search results log
  datasets/
    {source}_{slug}/
      metadata.json                  # detailed metadata
      README.md                      # human-readable summary
      sample.jsonl                   # sample rows

Dependencies

  • Python 3.8+
  • requests (stdlib-adjacent, universally available)
  • gh CLI (for GitHub source only)
  • No other packages required

© LigphiDonk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/dataset-discovery of LigphiDonk/Oh-my--paper.

  • SKILL.md
  • scripts/search_ml_datasets.py
  • tests/prompts.txt

Open the folder on GitHubat commit 6baece9

Compare with similar skills

Dataset Discovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Discovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Discovery this skillLigphiDonk/Oh-my--paper739—~1.3kAutomated safety check: PassMIT
Ideer Daily PaperAI45Lab/iDeer416—~2.3kAutomated safety check: NotesAGPL-3.0
Hugging Face Paper Pageshuggingface/skills11k3 repos~2.3kAutomated safety check: PassApache-2.0
Academic AioAperivue/medsci-skills333—~4.8kAutomated safety check: PassMIT
Morning AIdavepoon/buildwithclaude3.6k—~405Automated safety check: PassMIT
ML Dataset DiscoveryOpenLAIR/dr-claw1.2k—~741Automated safety check: PassCustom licence

Similar skills

  • Ideer Daily Paper

    AI45Lab/iDeer

    Daily paper/repo digest where YOU are the reader. An agent skill from AI45Lab/iDeer.

    416 GitHub stars~2.3k tokensUpdated 2 mo ago
    Research & ScienceAuto-check: notes
  • Hugging Face Paper Pages

    huggingface/skills

    Official

    Fetches Hugging Face paper pages as markdown and reads paper metadata through the papers API when you share a paper URL, an arXiv link or an arXiv ID.

    11k GitHub starsUsed in 3 repos~2.3k tokens
    Research & ScienceAuto-check passed
  • Academic Aio

    Aperivue/medsci-skills

    A skill your agent uses when a medical AI paper should be found and cited by AI search engines and RAG tools.

    333 GitHub stars~4.8k tokensUpdated 4 days ago
    Research & ScienceAuto-check passed
  • Morning AI

    davepoon/buildwithclaude

    AI news tracking skill that monitors 80+ entities across 6 free sources (Reddit, HN, GitHub, HuggingFace, arXiv, X/Twitter).

    3.6k GitHub stars~405 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • ML Dataset Discovery

    OpenLAIR/dr-claw

    Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.

    1.2k GitHub stars~741 tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • News Aggregator Skill

    cclank/news-aggregator-skill

    Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…

    1.3k GitHub stars~2.1k tokensUpdated 4 mo ago
    Writing & ContentAuto-check passed

More from LigphiDonk/Oh-my--paper

All 27 skills in this repo
  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    739 GitHub starsUsed in 12 repos~3.7k tokens
    Auto-check passed
  • Literature PDF OCR Library Builder

    LigphiDonk/Oh-my--paper

    Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.

    739 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Inno Code Survey

    LigphiDonk/Oh-my--paper

    Finds and clones missing code repositories for a chosen research idea, then writes a survey that maps academic concepts to their implementations.

    739 GitHub stars~3.6k tokensUpdated 5 mo ago
    Auto-check passed
  • Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.

    739 GitHub stars~3k tokensUpdated 5 mo ago
    Auto-check passed
  • Citation Verification Guide

    LigphiDonk/Oh-my--paper

    Lays out principles for catching fake, mismatched, or inconsistently formatted citations in academic writing, checked through live web search.

    739 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    739 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed

Questions about Dataset Discovery

How do I install Dataset Discovery in Claude Code?

Run `npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a claude-code`. Or copy the skill folder (skills/dataset-discovery in LigphiDonk/Oh-my--paper) into .claude/skills/dataset-discovery in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Discovery in Codex?

Run `npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a codex`. Or copy the skill folder (skills/dataset-discovery in LigphiDonk/Oh-my--paper) into .agents/skills/dataset-discovery in your project. Codex loads it when a task matches its description.

Can I use Dataset Discovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-discovery, .gemini/skills/dataset-discovery, .github/skills/dataset-discovery and .opencode/skills/dataset-discovery in your project.

What does Dataset Discovery need to run?

Going by SKILL.md and its folder, Dataset Discovery needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and huggingface-cli).

Does Dataset Discovery access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dataset Discovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Dataset Discovery use?

Dataset Discovery is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Discovery use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Discovery?

Skills that share tags, products or a category with Dataset Discovery: Ideer Daily Paper (AI45Lab/iDeer, 416 stars), Hugging Face Paper Pages (huggingface/skills, 11k stars), Academic Aio (Aperivue/medsci-skills, 333 stars) and Morning AI (davepoon/buildwithclaude, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Discovery?

LigphiDonk (a GitHub user) maintains it in LigphiDonk/Oh-my--paper, which has 739 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on April 15, 2026.

Source: LigphiDonk/Oh-my--paper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.