Agent skill

Dataset Finder Guide

by wentorai in wentorai/research-plugins

Search and download research datasets from Kaggle, HuggingFace, and repos

MITAuto-check passedAI & LLM Engineering

Install Dataset Finder Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill dataset-finder-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins dataset-finder-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/scraping/dataset-finder-guide .claude/skills/dataset-finder-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-finder-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
491 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Search and download research datasets from Kaggle, HuggingFace, and repos

  • Tasks that involve Model hubs and datasets
  • SKILL.md covers Overview, Dataset Repositories, Searching for Datasets and Dataset Evaluation Checklist, plus 2 more sections
  • Calls pip; reaches datasetsearch.research.google.com and zenodo.org
  • Tasks that involve Web scraping

What it does

Dataset Finder Guide is an agent skill from wentorai/research-plugins. Search and download research datasets from Kaggle, HuggingFace, and repos

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets and Web scraping. It works with Hugging Face and Kaggle. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Model hubs and datasets
  • Tasks that involve Web scraping

Example prompts

  • “/dataset-finder-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • datasetsearch.research.google.com
    • zenodo.org

    Also links to:

    • github.com
    • huggingface.co
    • developers.zenodo.org
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Finder Guide loads about 2.2k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 491 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 491 words, ~2,202 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-finder-guide/SKILL.md (or your agent's skills folder).
name
dataset-finder-guide
description
Search and download research datasets from Kaggle, HuggingFace, and repos

Dataset Finder Guide

Search, evaluate, and download research datasets from major repositories including Kaggle, Hugging Face, Google Dataset Search, Zenodo, UCI Machine Learning Repository, and domain-specific archives. This skill helps researchers locate the right data for their experiments efficiently.

Overview

Finding suitable datasets is often one of the most time-consuming phases of empirical research. Datasets are scattered across dozens of platforms, each with different APIs, licensing terms, download mechanisms, and metadata standards. A single research project might require datasets from Kaggle for benchmarking, Hugging Face for NLP tasks, Zenodo for supplementary materials from published papers, and government open data portals for demographic or economic variables.

This skill provides a unified approach to dataset discovery: formulating search queries, evaluating dataset quality and suitability, understanding licensing implications, and efficiently downloading and organizing data. It covers both general-purpose repositories and domain-specific archives that researchers in various fields need.

The emphasis is on reproducibility -- every dataset used in research should be citable, versioned, and documented. This skill includes patterns for recording dataset provenance, creating data cards, and managing dataset versions across experiments.

Dataset Repositories

General-Purpose Repositories
RepositoryStrengthsAPICitation Support
KaggleML benchmarks, competitions, community kernelsREST + CLIDOI via dataset cards
Hugging Face DatasetsNLP, CV, audio; streaming supportPython libraryBuilt-in citation
ZenodoAny research data, DOI minting, EU-fundedREST APIAutomatic DOI
Google Dataset SearchMeta-search across repositoriesWeb onlyLinks to source
UCI ML RepositoryClassic ML benchmarksDirect downloadBibTeX provided
FigshareFigures, datasets, media, preprintsREST APIDOI per item
DryadEcology, biology, environmental scienceREST APIDOI per dataset
ICPSRSocial science survey dataRestricted APIPersistent IDs
Harvard DataverseMulti-discipline, institutionalREST APIDOI per dataset
Show full SKILL.md (210 more words)Show less
Domain-Specific Archives
DomainRepositoryNotable Datasets
GenomicsNCBI GEO, ENAGene expression, sequencing data
AstronomyNASA archives, SDSSSky surveys, spectral data
EconomicsFRED, World Bank, IMFTime series, macro indicators
ClimateNOAA, CMIP6Temperature, precipitation records
LinguisticsLDC, CLARINCorpora, treebanks
MedicalPhysioNet, MIMICClinical records, ECG/EEG
ChemistryPubChem, ChEMBLMolecular structures, bioassays

Searching for Datasets

Kaggle CLI
bash
# Install and configure
pip install kaggle
# Place kaggle.json in ~/.kaggle/

# Search datasets
kaggle datasets list -s "sentiment analysis" --sort-by votes
kaggle datasets list -s "medical imaging" --file-type csv --min-size 100MB

# Get dataset details
kaggle datasets metadata -d stanford/imdb-review-dataset

# Download dataset
kaggle datasets download -d stanford/imdb-review-dataset -p ./data/
unzip ./data/imdb-review-dataset.zip -d ./data/imdb/

# Download competition data
kaggle competitions download -c titanic -p ./data/
Hugging Face Datasets
python
from datasets import load_dataset, list_datasets

# Search for datasets by task
from huggingface_hub import HfApi
api = HfApi()
datasets = api.list_datasets(
    search="scientific papers",
    sort="downloads",
    direction=-1,
    limit=20
)
for ds in datasets:
    print(f"{ds.id}: {ds.downloads} downloads")

# Load a dataset (with streaming for large datasets)
dataset = load_dataset("scientific_papers", "arxiv", streaming=True)

# Inspect structure
print(dataset["train"].features)
print(f"Number of examples: {dataset['train'].num_rows}")

# Load specific split and subset
validation = load_dataset(
    "scientific_papers", "arxiv",
    split="validation[:1000]"
)
Google Dataset Search (Programmatic)
python
import requests
from bs4 import BeautifulSoup

def search_google_datasets(query, num_results=10):
    """Search Google Dataset Search and extract results."""
    url = f"https://datasetsearch.research.google.com/search"
    params = {"query": query, "docid": ""}
    # Note: Google Dataset Search does not have an official API
    # Use the web interface or alternative approaches
    print(f"Search at: {url}?query={query.replace(' ', '+')}")
    return url
Zenodo API
python
import requests

def search_zenodo(query, resource_type="dataset", size=10):
    """Search Zenodo for research datasets."""
    url = "https://zenodo.org/api/records"
    params = {
        "q": query,
        "type": resource_type,
        "size": size,
        "sort": "mostrecent",
        "access_right": "open"
    }
    response = requests.get(url, params=params)
    results = response.json()

    for hit in results.get("hits", {}).get("hits", []):
        meta = hit["metadata"]
        print(f"Title: {meta['title']}")
        print(f"DOI: {meta.get('doi', 'N/A')}")
        print(f"License: {meta.get('license', {}).get('id', 'N/A')}")
        print(f"Size: {sum(f['size'] for f in hit.get('files', []))/1e6:.1f} MB")
        print("---")

    return results

Dataset Evaluation Checklist

Before using a dataset in research, verify the following:

Quality Assessment
  • Completeness: What percentage of values are missing? Are missing values random or systematic?
  • Accuracy: Are values within expected ranges? Are there obvious errors?
  • Consistency: Are formats uniform (dates, categories, units)?
  • Timeliness: When was the data collected? Is it current enough for your research question?
  • Sample size: Is the dataset large enough for your intended analysis?
Licensing and Ethics
LicenseCommercial UseModificationAttribution Required
CC0YesYesNo
CC-BY 4.0YesYesYes
CC-BY-SA 4.0YesYes (share-alike)Yes
CC-BY-NC 4.0NoYesYes
ODC-ODbLYesYes (share-alike)Yes
Custom/RestrictedVariesVariesVaries
Reproducibility Documentation
markdown
## Data Card

**Dataset**: [Name]
**Source**: [URL]
**Version**: [Version/Date]
**DOI**: [DOI if available]
**License**: [License name]
**Downloaded**: [YYYY-MM-DD]
**Size**: [X rows, Y columns, Z MB]
**Description**: [Brief description]
**Preprocessing**: [Steps applied before use]
**Citation**: [BibTeX entry]

Download and Organization

Project Data Structure
project/
  data/
    raw/              # Original downloaded data (never modify)
      dataset_v1.csv
      README.md       # Data card with provenance
    processed/        # Cleaned and transformed data
      train.csv
      test.csv
    external/         # Third-party reference data
  scripts/
    download_data.py  # Reproducible download script
    preprocess.py     # Data cleaning pipeline
Reproducible Download Script
python
"""download_data.py - Reproducible dataset download."""
import hashlib
from pathlib import Path
import requests

DATASETS = {
    "main_dataset": {
        "url": "https://zenodo.org/record/12345/files/data.csv",
        "sha256": "abc123...",
        "filename": "raw/main_dataset.csv"
    }
}

DATA_DIR = Path("data")

for name, info in DATASETS.items():
    path = DATA_DIR / info["filename"]
    if path.exists():
        print(f"Already downloaded: {name}")
        continue

    path.parent.mkdir(parents=True, exist_ok=True)
    print(f"Downloading {name}...")
    response = requests.get(info["url"])
    path.write_bytes(response.content)

    # Verify integrity
    sha256 = hashlib.sha256(response.content).hexdigest()
    assert sha256 == info["sha256"], f"Checksum mismatch for {name}"
    print(f"Verified: {name}")

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tools/scraping/dataset-finder-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Dataset Finder Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Finder Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Finder Guide this skillwentorai/research-plugins2981 repos~2.2kAutomated safety check: PassMIT
Dataset FinderLeoYeAI/openclaw-master-skills2.2k—~5.4kAutomated safety check: PassProprietary
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT

Similar skills

  • Dataset Finder

    LeoYeAI/openclaw-master-skills

    A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.

    2.2k GitHub stars~5.4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    916 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Dataset Finder Guide

What does Dataset Finder Guide do?

Search and download research datasets from Kaggle, HuggingFace, and repos. Dataset Finder Guide is an agent skill from wentorai/research-plugins.

When should I use Dataset Finder Guide?

Dataset Finder Guide fits situations like: tasks that involve Model hubs and datasets; tasks that involve Web scraping.

How do I install Dataset Finder Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill dataset-finder-guide -a claude-code`. Or copy the skill folder (skills/tools/scraping/dataset-finder-guide in wentorai/research-plugins) into .claude/skills/dataset-finder-guide in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Finder Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill dataset-finder-guide -a codex`. Or copy the skill folder (skills/tools/scraping/dataset-finder-guide in wentorai/research-plugins) into .agents/skills/dataset-finder-guide in your project. Codex loads it when a task matches its description.

Can I use Dataset Finder Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill dataset-finder-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-finder-guide, .gemini/skills/dataset-finder-guide, .github/skills/dataset-finder-guide and .opencode/skills/dataset-finder-guide in your project.

What does Dataset Finder Guide need to run?

Going by SKILL.md and its folder, Dataset Finder Guide needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Dataset Finder Guide access the network?

SKILL.md names 6 domains. In commands or code: datasetsearch.research.google.com and zenodo.org; the agent is likely to contact these when it follows the instructions. As links in the text: github.com, huggingface.co, developers.zenodo.org and arxiv.org. This is read from the text; nothing was executed.

Is Dataset Finder Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dataset Finder Guide use?

Dataset Finder Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Finder Guide use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Finder Guide?

Skills that share tags, products or a category with Dataset Finder Guide: Dataset Finder (LeoYeAI/openclaw-master-skills, 2.2k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Esmfold2 (JimLiu/science-skills, 228 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Finder Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.