Find and assess datasets for a research question. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

Custom licenceAuto-check passedResearch & Science

Install Data Finder

skills CLI
$ npx skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill data-finder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Auto-Empirical-Research-Skills data-finder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/15-Felpix-Studios-social-science-research/skills/data-finder .claude/skills/data-finder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-finder
GitHub stars
4.6k
Token cost
~1.7k tokens
SKILL.md length
328 words
Files
1
Skills in repo
383
Repo updated
First seen
Licence
Custom licence

At a glance

Find and assess datasets for a research question. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

  • Works in 5 steps: Read Research Context → Dispatch Two Explorer Agents in Parallel → Dispatch Explorer-Critic → …
  • The user wants to identify
  • SKILL.md covers Step 1: Read Research Context, Step 2: Dispatch Two Explorer…, Step 3: Dispatch Explorer-Critic and Step 4: Produce Ranked Output, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Finder is an agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. Find and assess datasets for a research question. Dispatches Explorer agents to search across data source categories, then Explorer-Critic to stress-test each candidate. Produces a ranked list with feasibility grades. Make sure to use this skill whenever the user wants to identify or evaluate data sources — not to search for papers or run analysis. Triggers include: "find data", "what data should I use", "find a dataset for this", "where can I get data on X", "assess datasets", "what datasets exist for", "help me…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Hypothesis generation and Load testing. The repository describes itself as: 🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI…

When your agent uses it

  • The user wants to identify
  • Evaluate data sources — not to search for papers
  • Include: find data
  • What data should I use

Example prompts

  • “find data”
  • “what data should I use”
  • “find a dataset for this”
  • “/data-finder”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Write, WebSearch, WebFetch, Task

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Read Research Context
  2. Dispatch Two Explorer Agents in Parallel
  3. Dispatch Explorer-Critic
  4. Produce Ranked Output
  5. Save Report

What it can do on your machine

Read from SKILL.md and the folder at commit 9fa87d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Write
    • WebSearch
    • WebFetch
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Finder loads about 1.7k tokens when it runs. Until then it costs about 175 tokens; SKILL.md has 328 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~175
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 328 words (~1,748 tokens).

“Find and assess datasets for your research question. Two Explorer agents search in parallel across data source categories; an Explorer-Critic then stress-tests each candidate against the research design.”

— opening of SKILL.md by brycewang-stanford, Custom licence
name
data-finder
allowed-tools
Read, Grep, Glob, Write, WebSearch, WebFetch, Task
argument-hint
[research topic or 'from spec']

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/15-Felpix-Studios-social-science-research/skills/data-finder of brycewang-stanford/Auto-Empirical-Research-Skills.

Open the folder on GitHubat commit 9fa87d8

Compare with similar skills

Data Finder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Finder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Finder this skillbrycewang-stanford/Auto-Empirical-Research-Skills4.6k—~1.7kAutomated safety check: PassCustom licence
Jbv Topic Selectionfranklee16/academic-research-skills2231 repos~1kAutomated safety check: PassNone
Mgsci Topic Selectionfranklee16/academic-research-skills2231 repos~998Automated safety check: PassNone
Research Refineappleweiping/WEIPING_WIKI119—~887Automated safety check: PassMIT
Jms Topic Selectionbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.3kAutomated safety check: PassMIT
Good QuestionRimagination/good-question3051 repos~4.3kAutomated safety check: PassMIT

Similar skills

  • Jbv Topic Selection

    franklee16/academic-research-skills

    A skill your agent uses when shaping or stress-testing a research question for the Journal of Business Venturing (JBV) — confirming the entrepreneurial phenomenon is central, picking a…

    223 GitHub starsUsed in 1 repo~1k tokens
    Research & ScienceAuto-check passed
  • Mgsci Topic Selection

    franklee16/academic-research-skills

    A skill your agent uses when shaping or stress-testing a research question for Management Science (INFORMS) — confirming the decision-relevance bar, choosing the right Department lane (analytical vs…

    223 GitHub starsUsed in 1 repo~998 tokens
    Research & ScienceAuto-check passed
  • Research Refine

    appleweiping/WEIPING_WIKI

    Disciplined idea refinement for research projects. An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~887 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Jms Topic Selection

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when scoping or stress-testing whether a research question fits the Journal of Management Studies (JMS) — a phenomenon-grounded management-theory question, quantitative or…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Research & ScienceAuto-check passed
  • Good Question

    Rimagination/good-question

    A skill your agent uses when a researcher is choosing, framing, refining, or stress-testing a research question, hypothesis, thesis topic, project idea, grant direction, paper angle, or stalled…

    305 GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check passed
  • Academic Grill

    Exekiel179/psyclaw

    Stress-test an academic research question, proposal, study design, analysis plan, manuscript claim, review protocol, or AI research project through a one-question-at-a-time interview until its…

    103 GitHub stars~2k tokensUpdated yesterday
    Research & ScienceAuto-check passed

More from brycewang-stanford/Auto-Empirical-Research-Skills

All 383 skills in this repo
  • Aer Consistency

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when auditing a finished or near-finished AER, AER:Insights, or AEJ manuscript for internal consistency: headline numbers across abstract, introduction, results, and tables…

    4.6k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Latex Paper En

    brycewang-stanford/Auto-Empirical-Research-Skills

    English LaTeX academic paper assistant for existing .tex projects.

    4.6k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Bayesian Workflow

    brycewang-stanford/Auto-Empirical-Research-Skills

    Opinionated Bayesian modeling workflow with PyMC and ArviZ. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

    4.6k GitHub stars~3.5k tokensUpdated 5 days ago
    Auto-check passed
  • Daily Paper Generator

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when the user asks to "generate daily paper", "search arXiv for EEG papers", "find EEG decoding papers", "review brain-computer interface papers", or wants to create paper…

    4.6k GitHub stars~2.2k tokensUpdated 5 days ago
    Auto-check passed
  • Five Questions

    brycewang-stanford/Auto-Empirical-Research-Skills

    Deeply analyze any empirical economics PDF using the five-question framework (五问框架): research question, identification strategy, core estimand, robustness logic, and scholarly contribution.

    4.6k GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check: notes
  • Kaggle Research

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when a research task needs reproducible Kaggle discovery, metadata inspection, bounded public-data downloads, competition or kernel discovery, model discovery, or an…

    4.6k GitHub stars~868 tokensUpdated 5 days ago
    Auto-check passed

Questions about Data Finder

What does Data Finder do?

Find and assess datasets for a research question. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. Data Finder is an agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. Find and assess datasets for a research question.

When should I use Data Finder?

Data Finder fits situations like: the user wants to identify; evaluate data sources — not to search for papers; include: find data; what data should I use.

How do I install Data Finder in Claude Code?

Run `npx skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill data-finder -a claude-code`. Or copy the skill folder (skills/15-Felpix-Studios-social-science-research/skills/data-finder in brycewang-stanford/Auto-Empirical-Research-Skills) into .claude/skills/data-finder in your project. Claude Code loads it when a task matches its description.

How do I install Data Finder in Codex?

Run `npx skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill data-finder -a codex`. Or copy the skill folder (skills/15-Felpix-Studios-social-science-research/skills/data-finder in brycewang-stanford/Auto-Empirical-Research-Skills) into .agents/skills/data-finder in your project. Codex loads it when a task matches its description.

Can I use Data Finder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill data-finder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-finder, .gemini/skills/data-finder, .github/skills/data-finder and .opencode/skills/data-finder in your project.

What does Data Finder need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Finder is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Grep, Glob, Write, WebSearch, WebFetch, Task.

Does Data Finder access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Finder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Finder use?

Data Finder has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Data Finder use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Finder?

Skills that share tags, products or a category with Data Finder: Jbv Topic Selection (franklee16/academic-research-skills, 223 stars), Mgsci Topic Selection (franklee16/academic-research-skills, 223 stars), Research Refine (appleweiping/WEIPING_WIKI, 119 stars) and Jms Topic Selection (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Finder?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Auto-Empirical-Research-Skills, which has 4,556 GitHub stars. The repository holds 383 skills in this directory. The repository was last updated on October 5, 2026.

Source: brycewang-stanford/Auto-Empirical-Research-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.