LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.
$ npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills dataset-finder --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-finder .claude/skills/dataset-finder && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dataset-finder" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finder into .claude/skills/dataset-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-finder", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills dataset-finder --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/dataset-finder .agents/skills/dataset-finder && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dataset-finder" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finder into .agents/skills/dataset-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-finder", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills dataset-finder --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/dataset-finder .cursor/skills/dataset-finder && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dataset-finder" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finder into .cursor/skills/dataset-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-finder", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/dataset-finder--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills dataset-finder --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/dataset-finder .gemini/skills/dataset-finder && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dataset-finder" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finder into .gemini/skills/dataset-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-finder", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills dataset-finderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/dataset-finder .github/skills/dataset-finder && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dataset-finder" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finder into .github/skills/dataset-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-finder", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills dataset-finder --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/dataset-finder .opencode/skills/dataset-finder && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dataset-finder" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/dataset-finder into .opencode/skills/dataset-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-finder", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dataset-finderA skill your agent uses when users need to search for datasets, download data files, or explore data repositories.
Dataset Finder is an agent skill from LeoYeAI/openclaw-master-skills. Use this skill when users need to search for datasets, download data files, or explore data repositories. Triggers include: requests to "find datasets", "search for data", "download dataset from Kaggle", "get data from Hugging Face", "find ML datasets", or mentions of data repositories like Kaggle, UCI ML Repository, Data.gov, or Hugging Face. Also use for previewing dataset statistics, generating data cards, or discovering datasets for machine learning projects. Requires OpenClawCLI installation from clawhub.ai.
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `_meta.json`, `references/readme.md` and `scripts/dataset.py`).
It sits in AI & LLM Engineering, covering Model hubs and datasets, Statistics and Machine learning. It works with Hugging Face, Kaggle and Python. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpiphuggingface-cliFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
kaggle.comhuggingface.coarchive.ics.uci.educatalog.data.govAlso links to:
clawhub.aigithub.comdata.govFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dataset Finder loads about 5.4k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 133 tokens; SKILL.md has 872 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Its licence (Proprietary) doesn't allow us to republish the file, so here is its outline and opening line. It has 872 words (~5,357 tokens).
“Search, download, and explore datasets from multiple repositories including Kaggle, Hugging Face, UCI ML Repository, and Data.gov. Preview statistics, generate data cards, and manage datasets for machine learning projects.”
SKILL.md and 4 other files (scripts, references) in skills/dataset-finder of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Dataset Finder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dataset Finder this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~5.4k | Automated safety check: Pass | Proprietary | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Hugging Face Vision Trainerhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face API Tool Builderhuggingface/skills | 11k | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
huggingface/skills
Builds reusable command line scripts that fetch, enrich or process data from the Hugging Face API, aimed at chained, repeated or automated tasks.
Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Works with
Categories
A skill your agent uses when users need to search for datasets, download data files, or explore data repositories. Dataset Finder is an agent skill from LeoYeAI/openclaw-master-skills. Use this skill when users need to search for datasets, download data files, or explore data repositories.
Dataset Finder fits situations like: users need to search for datasets; download data files; explore data repositories; include: requests to find datasets.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a claude-code`. Or copy the skill folder (skills/dataset-finder in LeoYeAI/openclaw-master-skills) into .claude/skills/dataset-finder in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a codex`. Or copy the skill folder (skills/dataset-finder in LeoYeAI/openclaw-master-skills) into .agents/skills/dataset-finder in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill dataset-finder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-finder, .gemini/skills/dataset-finder, .github/skills/dataset-finder and .opencode/skills/dataset-finder in your project.
Going by SKILL.md and its folder, Dataset Finder needs Python for the scripts in its folder, the command-line tools its instructions call (python, pip and huggingface-cli) and credentials named HF_TOKEN. Our summary lists: Python 3.
SKILL.md names 7 domains. In commands or code: kaggle.com, huggingface.co, archive.ics.uci.edu and catalog.data.gov; the agent is likely to contact these when it follows the instructions. As links in the text: clawhub.ai, github.com and data.gov. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Dataset Finder carries a proprietary licence (declared in SKILL.md). It is published on GitHub, but it is not open source: read the licence file before using or sharing it.
About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Dataset Finder: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Hugging Face Vision Trainer (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.