Ideer Daily Paper
AI45Lab/iDeer
Daily paper/repo digest where YOU are the reader. An agent skill from AI45Lab/iDeer.
Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install OpenLAIR/dr-claw dataset-discovery --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-discovery .claude/skills/dataset-discovery && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dataset-discovery" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discovery into .claude/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discoveryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install OpenLAIR/dr-claw dataset-discovery --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/dataset-discovery .agents/skills/dataset-discovery && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discovery into .agents/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install OpenLAIR/dr-claw dataset-discovery --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/dataset-discovery .cursor/skills/dataset-discovery && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dataset-discovery" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discovery into .cursor/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/OpenLAIR/dr-claw.git --path skills/dataset-discovery--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install OpenLAIR/dr-claw dataset-discovery --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/dataset-discovery .gemini/skills/dataset-discovery && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discovery into .gemini/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install OpenLAIR/dr-claw dataset-discoveryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/dataset-discovery .github/skills/dataset-discovery && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discovery into .github/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install OpenLAIR/dr-claw dataset-discovery --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OpenLAIR/dr-claw.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/dataset-discovery .opencode/skills/dataset-discovery && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/OpenLAIR/dr-claw/tree/main/skills/dataset-discovery into .opencode/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dataset-discoverySearches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.
The agent first clarifies the research task, the data modality (image, text, tabular, audio or any), a size preference and a licence preference. It then runs `scripts/search_ml_datasets.py search` with your query across four sources, huggingface, openml, github and papers, with options for sources, a maximum result count (30 by default), modality and workspace folder. It can add Hugging Face's `hub_repo_search` MCP tool for semantic search on top.
Results are shown as a markdown table of name, source, downloads, size, licence, tags and URL, sorted by relevance. For a dataset you pick, a `detail` command writes `metadata.json` and a README, and a `pull` command saves sample rows to `sample.jsonl`; a full download needs your confirmation first. Output goes to `./datasets/discovery/` by default, together with a dated search log. The script needs Python 3.8 or newer and `requests`, plus the `gh` CLI for the GitHub source.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d51b64e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3huggingface-cliFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ML Dataset Discovery loads about 741 tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 221 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 221 words (~741 tokens).
“Search multiple ML dataset sources (HuggingFace Hub, OpenML, GitHub, Semantic Scholar) and return a ranked, deduplicated list of relevant datasets.”
SKILL.md and 2 other files (scripts) in skills/dataset-discovery of OpenLAIR/dr-claw.
Open the folder on GitHubat commit d51b64e
ML Dataset Discovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ML Dataset Discovery this skillOpenLAIR/dr-claw | 1.2k | — | ~741 | Automated safety check: Pass | Custom licence | |
| Ideer Daily PaperAI45Lab/iDeer | 416 | — | ~2.3k | Automated safety check: Notes | AGPL-3.0 | |
| Hugging Face Paper Pageshuggingface/skills | 11k | 3 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Academic AioAperivue/medsci-skills | 333 | — | ~4.8k | Automated safety check: Pass | MIT | |
| News Aggregator Skillcclank/news-aggregator-skill | 1.3k | — | ~2.1k | Automated safety check: Pass | None | |
| Morning AIdavepoon/buildwithclaude | 3.6k | — | ~405 | Automated safety check: Pass | MIT |
AI45Lab/iDeer
Daily paper/repo digest where YOU are the reader. An agent skill from AI45Lab/iDeer.
huggingface/skills
Fetches Hugging Face paper pages as markdown and reads paper metadata through the papers API when you share a paper URL, an arXiv link or an arXiv ID.
Aperivue/medsci-skills
A skill your agent uses when a medical AI paper should be found and cited by AI search engines and RAG tools.
cclank/news-aggregator-skill
Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…
davepoon/buildwithclaude
AI news tracking skill that monitors 80+ entities across 6 free sources (Reddit, HN, GitHub, HuggingFace, arXiv, X/Twitter).
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
OpenLAIR/dr-claw
Analyzes reviewer comments and drafts venue-specific rebuttals for AI and computer science conferences, with an issue board, task list and paper edit plan.
OpenLAIR/dr-claw
Turns a research paper into a slide deck and, optionally, a narrated demo video, through script, slide generation, text-to-speech and video assembly stages you control.
OpenLAIR/dr-claw
Clusters the latest news-feed results by topic and writes a briefing of research idea seeds with citations, plus a structured seeds file, without crawling new sources.
OpenLAIR/dr-claw
Runs multi-source web research through Google's Gemini Deep Research Agent with a bundled Python script and saves a structured, cited report as files.
OpenLAIR/dr-claw
Six-phase workflow for writing, revising and adapting grant proposals for NSF, NIH, DOE, DARPA, NASA and China's NSFC, from profiling through simulated peer review.
OpenLAIR/dr-claw
Access Overleaf projects via CLI. An agent skill from OpenLAIR/dr-claw.
Works with
Categories
Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table. The agent first clarifies the research task, the data modality (image, text, tabular, audio or any), a size preference and a licence preference.py search` with your query across four sources, huggingface, openml, github and papers, with options for sources, a maximum result count (30 by default), modality and workspace folder.
ML Dataset Discovery fits situations like: finding training or benchmark datasets for a new research task; comparing candidate datasets by size, licence and download counts; previewing sample rows before committing to a full download.
Run `npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a claude-code`. Or copy the skill folder (skills/dataset-discovery in OpenLAIR/dr-claw) into .claude/skills/dataset-discovery in your project. Claude Code loads it when a task matches its description.
Run `npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a codex`. Or copy the skill folder (skills/dataset-discovery in OpenLAIR/dr-claw) into .agents/skills/dataset-discovery in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenLAIR/dr-claw --skill dataset-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-discovery, .gemini/skills/dataset-discovery, .github/skills/dataset-discovery and .opencode/skills/dataset-discovery in your project.
Going by SKILL.md and its folder, ML Dataset Discovery needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and huggingface-cli). Our summary lists: Python 3.8 or newer with `requests`; The `gh` CLI for the GitHub source; Network access to the dataset sources.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
ML Dataset Discovery has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.
About 741 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with ML Dataset Discovery: Ideer Daily Paper (AI45Lab/iDeer, 416 stars), Hugging Face Paper Pages (huggingface/skills, 11k stars), Academic Aio (Aperivue/medsci-skills, 333 stars) and News Aggregator Skill (cclank/news-aggregator-skill, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
OpenLAIR (a GitHub organization) maintains it in OpenLAIR/dr-claw, which has 1,155 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 17, 2026.
Source: OpenLAIR/dr-claw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.