Ideer Daily Paper
AI45Lab/iDeer
Daily paper/repo digest where YOU are the reader. An agent skill from AI45Lab/iDeer.
“Multi-source ML dataset discovery.”
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LigphiDonk/Oh-my--paper dataset-discovery --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-discovery .claude/skills/dataset-discovery && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dataset-discovery" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discovery into .claude/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discoveryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LigphiDonk/Oh-my--paper dataset-discovery --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/dataset-discovery .agents/skills/dataset-discovery && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discovery into .agents/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LigphiDonk/Oh-my--paper dataset-discovery --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/dataset-discovery .cursor/skills/dataset-discovery && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dataset-discovery" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discovery into .cursor/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LigphiDonk/Oh-my--paper.git --path skills/dataset-discovery--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LigphiDonk/Oh-my--paper dataset-discovery --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/dataset-discovery .gemini/skills/dataset-discovery && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discovery into .gemini/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LigphiDonk/Oh-my--paper dataset-discoveryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/dataset-discovery .github/skills/dataset-discovery && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discovery into .github/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LigphiDonk/Oh-my--paper dataset-discovery --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/dataset-discovery .opencode/skills/dataset-discovery && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dataset-discovery" agent skill from https://github.com/LigphiDonk/Oh-my--paper/tree/main/skills/dataset-discovery into .opencode/skills/dataset-discovery/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-discovery", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dataset-discoveryDataset Discovery is a skill in LigphiDonk/Oh-my--paper (739 stars). Its SKILL.md is about 1.3k tokens, with 2 other files in the folder (scripts). Licence: MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6baece9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3huggingface-cliFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dataset Discovery loads about 1.3k tokens when it runs. Until then it costs about 13 tokens; SKILL.md has 390 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from LigphiDonk/Oh-my--paper at commit 6baece9, republished under its MIT licence (© LigphiDonk). 390 words, ~1,252 tokens.
.claude/skills/dataset-discovery/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Multi-source ML dataset discovery. Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets relevant to a research task. Use when asked to "find datasets for", "search ML datasets", "what datasets exist for", or "dis...
Use this skill when the user request matches its research workflow scope. Prefer the bundled resources instead of recreating templates or reference material. Keep outputs traceable to project files, citations, scripts, or upstream evidence.
scripts/ as optional helpers. Run them only when their dependencies are available, keep outputs in the project workspace, and explain a manual fallback if execution is blocked.Search multiple ML dataset sources (HuggingFace Hub, OpenML, GitHub, Semantic Scholar) and return a ranked, deduplicated list of relevant datasets.
Clarify the user's needs before searching:
Run the search script with the user's query:
python3 scripts/search_ml_datasets.py search --query "<query>" --sources huggingface,openml,github,papers --max 30Options:
--sources: Comma-separated list from huggingface, openml, github, papers. Default: all four.--max: Maximum results to return after dedup + ranking. Default: 30.--modality: Filter by modality (image, text, tabular, audio).--workspace: Output directory. Default: ./datasets/discovery/Optionally also call HF MCP tool hub_repo_search with repo_types: ["dataset"] for semantic search to supplement results.
Show results as a markdown table:
| Name | Source | Downloads | Size | License | Tags | URL |
|---|
Sort by relevance score (highest first).
When the user wants more info on a specific dataset:
python3 scripts/search_ml_datasets.py detail --dataset-id "huggingface:stanfordnlp/imdb" --workspace ./datasets/discovery/Writes metadata.json and README.md to {workspace}/datasets/{source}_{slug}/.
When the user wants to preview data:
python3 scripts/search_ml_datasets.py pull --dataset-id "huggingface:stanfordnlp/imdb" --sample-rows 20 --workspace ./datasets/discovery/Writes sample.jsonl to {workspace}/datasets/{source}_{slug}/.
For full dataset download, confirm with the user first, then use huggingface-cli download or equivalent.
{workspace}/ # default: ./datasets/discovery/
search-{YYYY-MM-DD}.json # search results log
datasets/
{source}_{slug}/
metadata.json # detailed metadata
README.md # human-readable summary
sample.jsonl # sample rowsrequests (stdlib-adjacent, universally available)gh CLI (for GitHub source only)© LigphiDonk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in skills/dataset-discovery of LigphiDonk/Oh-my--paper.
Open the folder on GitHubat commit 6baece9
Dataset Discovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dataset Discovery this skillLigphiDonk/Oh-my--paper | 739 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Ideer Daily PaperAI45Lab/iDeer | 416 | — | ~2.3k | Automated safety check: Notes | AGPL-3.0 | |
| Hugging Face Paper Pageshuggingface/skills | 11k | 3 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Academic AioAperivue/medsci-skills | 333 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Morning AIdavepoon/buildwithclaude | 3.6k | — | ~405 | Automated safety check: Pass | MIT | |
| ML Dataset DiscoveryOpenLAIR/dr-claw | 1.2k | — | ~741 | Automated safety check: Pass | Custom licence |
AI45Lab/iDeer
Daily paper/repo digest where YOU are the reader. An agent skill from AI45Lab/iDeer.
huggingface/skills
Fetches Hugging Face paper pages as markdown and reads paper metadata through the papers API when you share a paper URL, an arXiv link or an arXiv ID.
Aperivue/medsci-skills
A skill your agent uses when a medical AI paper should be found and cited by AI search engines and RAG tools.
davepoon/buildwithclaude
AI news tracking skill that monitors 80+ entities across 6 free sources (Reddit, HN, GitHub, HuggingFace, arXiv, X/Twitter).
OpenLAIR/dr-claw
Searches Hugging Face Hub, OpenML, GitHub and paper references for datasets that fit a research task and returns a ranked, de-duplicated table.
cclank/news-aggregator-skill
Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…
LigphiDonk/Oh-my--paper
Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.
LigphiDonk/Oh-my--paper
Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.
LigphiDonk/Oh-my--paper
Finds and clones missing code repositories for a chosen research idea, then writes a survey that maps academic concepts to their implementations.
LigphiDonk/Oh-my--paper
Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.
LigphiDonk/Oh-my--paper
Lays out principles for catching fake, mismatched, or inconsistently formatted citations in academic writing, checked through live web search.
LigphiDonk/Oh-my--paper
Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.
Works with
Categories
Run `npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a claude-code`. Or copy the skill folder (skills/dataset-discovery in LigphiDonk/Oh-my--paper) into .claude/skills/dataset-discovery in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a codex`. Or copy the skill folder (skills/dataset-discovery in LigphiDonk/Oh-my--paper) into .agents/skills/dataset-discovery in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LigphiDonk/Oh-my--paper --skill dataset-discovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-discovery, .gemini/skills/dataset-discovery, .github/skills/dataset-discovery and .opencode/skills/dataset-discovery in your project.
Going by SKILL.md and its folder, Dataset Discovery needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and huggingface-cli).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Dataset Discovery is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dataset Discovery: Ideer Daily Paper (AI45Lab/iDeer, 416 stars), Hugging Face Paper Pages (huggingface/skills, 11k stars), Academic Aio (Aperivue/medsci-skills, 333 stars) and Morning AI (davepoon/buildwithclaude, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LigphiDonk (a GitHub user) maintains it in LigphiDonk/Oh-my--paper, which has 739 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on April 15, 2026.
Source: LigphiDonk/Oh-my--paper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.