Agent skill

Scholar Annotate

by joshzyj in joshzyj/open-scholar-skill

Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70)…

Custom licenceAuto-check passedAI & LLM Engineering

Install Scholar Annotate

skills CLI
$ npx skills add joshzyj/open-scholar-skill --skill scholar-annotate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install joshzyj/open-scholar-skill scholar-annotate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/joshzyj/open-scholar-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/scholar-annotate .claude/skills/scholar-annotate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scholar-annotate
GitHub stars
168
Token cost
~4.6k tokens
SKILL.md length
1,843 words
Files
32 (incl. references, assets)
Skills in repo
30
Repo updated
First seen
Licence
Custom licence

At a glance

Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70)…

  • Framing/stance coding
  • SKILL.md covers Arguments and Mode Routing, MODE 0: Setup (all modes), The Execution Engine (assets/) and Modes, plus 3 more sections
  • Runs Python and Shell scripts from its folder; calls python3, bash and claude; needs OPENAI_API_KEY and ANTHROPIC_API_KEY
  • Relevance filtering

What it does

Scholar Annotate is an agent skill from joshzyj/open-scholar-skill. Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70), scale-out via managed Batch APIs or a local/HPC server, and distillation to a cheap classifier for very large corpora. Use for LLM annotation, classification, framing/stance coding, relevance filtering, structured extraction, and any LLM-as-measurement task over a text corpus. Ships a real execution engine (assets/) —…

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 34 other files, including reference files and assets (for example `assets/_lib.sh`, `assets/active_sample.py` and `assets/annotate_engine.py`).

It sits in AI & LLM Engineering, covering Prompt engineering and Document parsing. The repository describes itself as: Open scholar skill, a claude code plugin, for academic research.

When your agent uses it

  • Framing/stance coding
  • Relevance filtering
  • Structured extraction
  • Any LLM-as-measurement task over a text corpus

Example prompts

  • “/scholar-annotate”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 6e5ac8e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • bash
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scholar Annotate loads about 4.6k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 1,843 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,843 words (~4,566 tokens).

“You are an expert computational social scientist who uses LLMs as measurement instruments: converting a text corpus into validated, reproducible variables (classes, frames, stances, scores, extracted fields). This skill ships a real execution engine (assets/annotate_engine.py + friends) that scales from…”

— opening of SKILL.md by joshzyj, Custom licence
name
scholar-annotate
tools
Read, Bash, Write, WebSearch, Agent
argument-hint
[plan|profile|codebook|devset|annotate-gold|optimize|validate|scale|distill|report|full] [corpus path or task description]
user-invocable
true

Read the full SKILL.md on GitHub

Files

SKILL.md and 31 other files (references, assets) in .claude/skills/scholar-annotate of joshzyj/open-scholar-skill.

  • SKILL.md
  • assets/_lib.sh
  • assets/active_sample.py
  • assets/annotate_engine.py
  • assets/codebook_schema.py
  • assets/distill_embed.py
  • assets/dspy_optimize.py
  • assets/dspy_run.py
  • assets/finetune_distill.py
  • assets/gold_reconcile.py
  • assets/hpc/annotate.sbatch
  • assets/hpc/prep-mirror.py
  • assets/hpc/rsync-helper.sh
  • assets/providers.py
  • assets/reduce_corpus.py
  • assets/requirements.txt
  • assets/run-annotate.sh
  • references/codebook-design.md
  • … and 14 more

Open the folder on GitHubat commit 6e5ac8e

Compare with similar skills

Scholar Annotate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scholar Annotate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scholar Annotate this skilljoshzyj/open-scholar-skill168—~4.6kAutomated safety check: PassCustom licence
Create Simple Promptpnp/copilot-prompts892—~2.6kAutomated safety check: PassMIT
Offline Prompt Packagerzhaohui-yang/official-document-drafting140—~433Automated safety check: PassCustom licence
Matlab Recognize Textmatlab/matlab-agentic-toolkit1.1k—~5.2kAutomated safety check: PassCustom licence
Text Analysis BasicDrchronx/ai-agent-research-starter-kit135—~737Automated safety check: PassCustom licence
Prompt Improverseverity1/claude-code-prompt-improver1.9k2 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Create Simple Prompt

    pnp/copilot-prompts

    This skill should be used when the user asks to "create a new prompt sample", "add a new prompt sample", "scaffold a new prompt sample", "create a prompt contribution", "add a prompt", or needs to…

    892 GitHub stars~2.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Offline Prompt Packager

    zhaohui-yang/official-document-drafting

    把基于 prompts/ 主源的 skill 打包成断网单机可用的离线提示词包。Use when the user wants to export/package offline prompts for disconnected hosts (WebUI、Qwen、AnythingLLM、Claude.ai), generate self-contained systemprompt /…

    140 GitHub stars~433 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Matlab Recognize Text

    matlab/matlab-agentic-toolkit

    Build OCR pipelines in MATLAB using the ocr() function. An agent skill from matlab/matlab-agentic-toolkit.

    1.1k GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Text Analysis Basic

    Drchronx/ai-agent-research-starter-kit

    Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity.

    135 GitHub stars~737 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Improver

    severity1/claude-code-prompt-improver

    This skill enriches vague prompts with targeted research and clarification before execution.

    1.9k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineering Patterns

    ynulihao/AgentSkillOS

    Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability in production.

    617 GitHub starsUsed in 15 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from joshzyj/open-scholar-skill

All 30 skills in this repo
  • Scholar Auto Research

    joshzyj/open-scholar-skill

    Stable, deterministic social-science research-paper pipeline from idea or data to verified manuscript, citations, replication package, and final md/docx/tex/pdf outputs.

    168 GitHub stars~21k tokensUpdated 20 days ago
    Auto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 20 days ago
    Auto-check: notes
  • Scholar Causal

    joshzyj/open-scholar-skill

    Comprehensive causal inference toolkit for social science research.

    168 GitHub stars~10k tokensUpdated 20 days ago
    Auto-check passed
  • Scholar Data

    joshzyj/open-scholar-skill

    Comprehensive open data directory (100+ datasets across 14 categories) with auto-fetch capability, plus data collection instrument design, variable dictionaries, data management, IRB materials, and…

    168 GitHub stars~23k tokensUpdated 20 days ago
    Auto-check: notes
  • Scholar Eda

    joshzyj/open-scholar-skill

    Conduct exploratory data analysis (EDA) before hypothesis testing.

    168 GitHub stars~12k tokensUpdated 20 days ago
    Auto-check passed
  • Scholar Idea

    joshzyj/open-scholar-skill

    Explore broad social science ideas and convert them into formal, researchable questions.

    168 GitHub stars~8.8k tokensUpdated 20 days ago
    Auto-check: notes

Questions about Scholar Annotate

What does Scholar Annotate do?

Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70)…. Scholar Annotate is an agent skill from joshzyj/open-scholar-skill.70), scale-out via managed Batch APIs or a local/HPC server, and distillation to a cheap classifier for very large corpora.

When should I use Scholar Annotate?

Scholar Annotate fits situations like: framing/stance coding; relevance filtering; structured extraction; any LLM-as-measurement task over a text corpus.

How do I install Scholar Annotate in Claude Code?

Run `npx skills add joshzyj/open-scholar-skill --skill scholar-annotate -a claude-code`. Or copy the skill folder (.claude/skills/scholar-annotate in joshzyj/open-scholar-skill) into .claude/skills/scholar-annotate in your project. Claude Code loads it when a task matches its description.

How do I install Scholar Annotate in Codex?

Run `npx skills add joshzyj/open-scholar-skill --skill scholar-annotate -a codex`. Or copy the skill folder (.claude/skills/scholar-annotate in joshzyj/open-scholar-skill) into .agents/skills/scholar-annotate in your project. Codex loads it when a task matches its description.

Can I use Scholar Annotate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add joshzyj/open-scholar-skill --skill scholar-annotate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scholar-annotate, .gemini/skills/scholar-annotate, .github/skills/scholar-annotate and .opencode/skills/scholar-annotate in your project.

What does Scholar Annotate need to run?

Going by SKILL.md and its folder, Scholar Annotate needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (python3, bash and claude) and credentials named OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A Bash shell.

Does Scholar Annotate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scholar Annotate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scholar Annotate use?

Scholar Annotate has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Scholar Annotate use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Scholar Annotate?

Skills that share tags, products or a category with Scholar Annotate: Create Simple Prompt (pnp/copilot-prompts, 892 stars), Offline Prompt Packager (zhaohui-yang/official-document-drafting, 140 stars), Matlab Recognize Text (matlab/matlab-agentic-toolkit, 1.1k stars) and Text Analysis Basic (Drchronx/ai-agent-research-starter-kit, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scholar Annotate?

joshzyj (a GitHub user) maintains it in joshzyj/open-scholar-skill, which has 168 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on September 18, 2026.

Source: joshzyj/open-scholar-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.