Agent skill

Scholar Data

by joshzyj in joshzyj/open-scholar-skill

Comprehensive open data directory (100+ datasets across 14 categories) with auto-fetch capability, plus data collection instrument design, variable dictionaries, data management, IRB materials, and…

Custom licenceAuto-check: notesData & Analytics

Install Scholar Data

skills CLI
$ npx skills add joshzyj/open-scholar-skill --skill scholar-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install joshzyj/open-scholar-skill scholar-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/joshzyj/open-scholar-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/scholar-data .claude/skills/scholar-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scholar-data
GitHub stars
168
Token cost
~23k tokens
SKILL.md length
6,670 words
Files
5 (incl. references)
Skills in repo
30
Repo updated
First seen
Licence
Custom licence

At a glance

Comprehensive open data directory (100+ datasets across 14 categories) with auto-fetch capability, plus data collection instrument design, variable dictionaries, data management, IRB materials, and…

  • The user needs to find
  • SKILL.md covers Arguments, Dispatch Table, Step 0 — Data Safety Sidecar… and Pre-Execution Review Duty…, plus 4 more sections
  • Calls bash, git and jq; reaches ghoapi.azureedge.net and worldvaluessurvey.org; needs CENSUS_API_KEY and FRED_API_KEY
  • Download data for a research question

What it does

Scholar Data is an agent skill from joshzyj/open-scholar-skill. Comprehensive open data directory (100+ datasets across 14 categories) with auto-fetch capability, plus data collection instrument design, variable dictionaries, data management, IRB materials, and web/digital data collection for social science studies. Covers GSS, PSID, ACS, CPS, ESS, WVS, Afrobarometer, Eurobarometer, DHS, PISA, OECD, Eurostat, WHO, OpenAlex, Harvard Dataverse, ICPSR, Zenodo, OSF, and many more. Use when the user needs to find or download data for a research question, design a survey or…

Its SKILL.md is about 23k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/data-management.md`, `references/interview-protocols.md` and `references/survey-design.md`).

It sits in Data & Analytics, covering Web scraping, Academic paper search and Hypothesis generation. The repository describes itself as: Open scholar skill, a claude code plugin, for academic research.

When your agent uses it

  • The user needs to find
  • Download data for a research question
  • Design a survey
  • Interview protocol

Example prompts

  • “/scholar-data”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 6e5ac8e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • git
    • jq
    • python3
    • pip
    • playwright

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ghoapi.azureedge.net
    • worldvaluessurvey.org
    • web.archive.org
    • kjhealy.r-universe.dev
    • dataverse.nl
    • opportunityinsights.org
    • data.gdeltproject.org
    • api.openalex.org
    • dataverse.harvard.edu

    Also links to:

    • api.census.gov
    • fred.stlouisfed.org
    • europeansocialsurvey.org
    • dhsprogram.com
    • api.data.gov
    • semanticscholar.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CENSUS_API_KEY
    • FRED_API_KEY
    • FBI_API_KEY
    • S2_API_KEY
    • API_KEY
    • REDDIT_CLIENT_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scholar Data loads about 23k tokens when it runs, and up to ~32k if it reads all its reference files. Until then it costs about 178 tokens; SKILL.md has 6,670 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~178
When it runs · the whole SKILL.md, loaded when a task matches
~23k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~32k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:1206
    .env
  • NoteMentions a .env fileSKILL.md:1217
    entified data files. Use `.Renviron` or `.env` for API keys.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 6,670 words (~22,670 tokens).

“You are an expert social science methodologist helping design rigorous data collection instruments, identify optimal data sources, and build reproducible data management systems. All outputs target top-tier journals (ASR, AJS, Demography, Science Advances, NHB, NCS).”

— opening of SKILL.md by joshzyj, Custom licence
name
scholar-data
tools
Read, WebSearch, Write, Bash, Agent
argument-hint
[dataset|survey|interview|irb|manage|vignette|scrape|web|api|social media] [topic or research question] [optional: population, journal, design]
user-invocable
true

Read the full SKILL.md on GitHub

Files

SKILL.md and 4 other files (references) in .claude/skills/scholar-data of joshzyj/open-scholar-skill.

  • SKILL.md
  • references/data-management.md
  • references/interview-protocols.md
  • references/survey-design.md
  • references/web-scraping.md

Open the folder on GitHubat commit 6e5ac8e

Compare with similar skills

Scholar Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scholar Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scholar Data this skilljoshzyj/open-scholar-skill168—~23kAutomated safety check: NotesCustom licence
Google Scholar Scraperwentorai/research-plugins2981 repos~2.1kAutomated safety check: PassMIT
Paper Radartigerless-labs/paper-radar218—~2.4kAutomated safety check: PassCustom licence
Paper NavigatorEvoScientist/EvoSkills475—~6.3kAutomated safety check: NotesApache-2.0
Superlearnraiyanyahya/Superlearn121—~6.2kAutomated safety check: PassMIT
Research PlatformZS520L/HanakoPro102—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Google Scholar Scraper

    wentorai/research-plugins

    Ethical Google Scholar data collection techniques and best practices

    298 GitHub starsUsed in 1 repo~2.1k tokens
    Data & AnalyticsAuto-check passed
  • Paper Radar

    tigerless-labs/paper-radar

    Scrape AI papers published by 28 big tech companies and AI labs in a given date window, with institutional attribution (lead vs.

    218 GitHub stars~2.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Paper Navigator

    EvoScientist/EvoSkills

    Find and read academic papers (S2 + arXiv). An agent skill from EvoScientist/EvoSkills.

    475 GitHub stars~6.3k tokensUpdated 7 days ago
    Research & ScienceAuto-check: notes
  • Superlearn

    raiyanyahya/Superlearn

    Build an interactive learning board on any topic. An agent skill from raiyanyahya/Superlearn.

    121 GitHub stars~6.2k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Research Platform

    ZS520L/HanakoPro

    全自动科研平台:从论文检索、知识图谱构建、研究缺口分析、假设生成,到实验执行、论文写作、自审修正的完整科研流水线。Agent 按此 Skill 的指令自主推进研究流程。触发场景:做研究、搜论文、找研究缺口、生成假设、跑实验、写论文、文献综述、benchmark对比 / Triggers: research, literature review, paper search, hypothesis…

    102 GitHub stars~1.1k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Literature Reviewer Skill

    Drchronx/ai-agent-research-starter-kit

    Build high-quality literature reviews from a research topic using a 10-phase workflow.

    135 GitHub stars~2.5k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed

More from joshzyj/open-scholar-skill

All 30 skills in this repo
  • Scholar Annotate

    joshzyj/open-scholar-skill

    Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70)…

    168 GitHub stars~4.6k tokensUpdated 19 days ago
    Auto-check passed
  • Scholar Auto Research

    joshzyj/open-scholar-skill

    Stable, deterministic social-science research-paper pipeline from idea or data to verified manuscript, citations, replication package, and final md/docx/tex/pdf outputs.

    168 GitHub stars~21k tokensUpdated 19 days ago
    Auto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 19 days ago
    Auto-check: notes
  • Scholar Causal

    joshzyj/open-scholar-skill

    Comprehensive causal inference toolkit for social science research.

    168 GitHub stars~10k tokensUpdated 19 days ago
    Auto-check passed
  • Scholar Eda

    joshzyj/open-scholar-skill

    Conduct exploratory data analysis (EDA) before hypothesis testing.

    168 GitHub stars~12k tokensUpdated 19 days ago
    Auto-check passed
  • Scholar Idea

    joshzyj/open-scholar-skill

    Explore broad social science ideas and convert them into formal, researchable questions.

    168 GitHub stars~8.8k tokensUpdated 19 days ago
    Auto-check: notes

Questions about Scholar Data

What does Scholar Data do?

Comprehensive open data directory (100+ datasets across 14 categories) with auto-fetch capability, plus data collection instrument design, variable dictionaries, data management, IRB materials, and…. Scholar Data is an agent skill from joshzyj/open-scholar-skill. Comprehensive open data directory (100+ datasets across 14 categories) with auto-fetch capability, plus data collection instrument design, variable dictionaries, data management, IRB materials, and web/digital data collection for social science studies.

When should I use Scholar Data?

Scholar Data fits situations like: the user needs to find; download data for a research question; design a survey; interview protocol.

How do I install Scholar Data in Claude Code?

Run `npx skills add joshzyj/open-scholar-skill --skill scholar-data -a claude-code`. Or copy the skill folder (.claude/skills/scholar-data in joshzyj/open-scholar-skill) into .claude/skills/scholar-data in your project. Claude Code loads it when a task matches its description.

How do I install Scholar Data in Codex?

Run `npx skills add joshzyj/open-scholar-skill --skill scholar-data -a codex`. Or copy the skill folder (.claude/skills/scholar-data in joshzyj/open-scholar-skill) into .agents/skills/scholar-data in your project. Codex loads it when a task matches its description.

Can I use Scholar Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add joshzyj/open-scholar-skill --skill scholar-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scholar-data, .gemini/skills/scholar-data, .github/skills/scholar-data and .opencode/skills/scholar-data in your project.

What does Scholar Data need to run?

Going by SKILL.md and its folder, Scholar Data needs the command-line tools its instructions call (bash, git, jq, python3, pip and playwright) and credentials named CENSUS_API_KEY, FRED_API_KEY, FBI_API_KEY and S2_API_KEY. Our summary lists: Python 3.

Does Scholar Data access the network?

SKILL.md names 15 domains. In commands or code: ghoapi.azureedge.net, worldvaluessurvey.org, web.archive.org, kjhealy.r-universe.dev, dataverse.nl, opportunityinsights.org, data.gdeltproject.org, api.openalex.org and dataverse.harvard.edu; the agent is likely to contact these when it follows the instructions. As links in the text: api.census.gov, fred.stlouisfed.org, europeansocialsurvey.org, dhsprogram.com, api.data.gov and semanticscholar.org. This is read from the text; nothing was executed.

Is Scholar Data safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Scholar Data use?

Scholar Data has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Scholar Data use?

About 23k tokens (SKILL.md is roughly 91k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.8k tokens, read only when the agent opens those files.

What are the alternatives to Scholar Data?

Skills that share tags, products or a category with Scholar Data: Google Scholar Scraper (wentorai/research-plugins, 298 stars), Paper Radar (tigerless-labs/paper-radar, 218 stars), Paper Navigator (EvoScientist/EvoSkills, 475 stars) and Superlearn (raiyanyahya/Superlearn, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scholar Data?

joshzyj (a GitHub user) maintains it in joshzyj/open-scholar-skill, which has 168 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on September 18, 2026.

Source: joshzyj/open-scholar-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.