Agent skill

Paper Labeler

by PurCL in PurCL/ASE

Extracts research papers from .bib and .html venue files, filters them by relevance, labels them with a two-level taxonomy through Claude and rebuilds a website.

No licenceAuto-check: warningsResearch & Science

Install Paper Labeler

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add PurCL/ASE --skill paper-labeler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PurCL/ASE paper-labeler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PurCL/ASE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/paper-labeler .claude/skills/paper-labeler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paper-labeler
GitHub stars
637
Token cost
~5.2k tokens
SKILL.md length
1,830 words
Files
9 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
None found

At a glance

Extracts research papers from .bib and .html venue files, filters them by relevance, labels them with a two-level taxonomy through Claude and rebuilds a website.

  • Works in 12 steps: Extract Papers → Filter and Label → Merge into labeldata → …
  • Processing new conference proceedings from .bib or .html files
  • SKILL.md covers Pipeline Overview, Step-by-Step Workflow, Label Taxonomy and Relevance Criteria, plus 2 more sections
  • Runs Python scripts from its folder; calls python; reaches doi.org and aclanthology.org

What it does

The pipeline runs four scripts in sequence. `extract_papers.py` parses raw venue files into a uniform JSON format, with `fetch_ndss_abstracts.py` added for NDSS to scrape abstracts from paper pages. `label_papers.py` filters by relevance keywords and then labels with the Claude API, `merge_labeldata.py` folds the results into `data/labeldata/labeldata.json`, and `build_site.py` rebuilds the website. It targets LLM-for-code papers from conference proceedings.

For batch work, `process_folder.py` scans a rawdata folder, skips venues already listed in `data/venues.json` and runs the whole pipeline for new ones, grouping files by venue name so that two EMNLP 2024 track files are handled together. Options include `--dry-run`, `--filter-only`, `--model`, `--region`, `--delay` and `--no-rebuild`. `import_original.py` brings in papers from `data/rawdata/original.json` that are missing from the label data, mapping old labels to the current taxonomy.

When your agent uses it

  • Processing new conference proceedings from .bib or .html files
  • Labeling papers with the two-level taxonomy
  • Adding a new venue to the paper database and rebuilding the site

Example prompts

  • “Process the new venue files in data/rawdata and rebuild the website.”
  • “Do a dry run to show which venues in the rawdata folder are not yet processed.”
  • “Process the EMNLP-main2024.html file and add its papers to the label data.”

Requirements

  • Python 3 to run the scripts in .claude/skills/paper-labeler/scripts/
  • AWS credentials with access to a Claude model on Bedrock
  • Raw venue data as .bib or .html files under data/rawdata/

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Extract Papers
  2. Filter and Label
  3. Merge into labeldata
  4. Rebuild Website
  5. Code Generation
  6. Static Analysis
  7. Dynamic Analysis
  8. Model Safety and Security
  9. Agent Safety and Security
  10. Agent Design
  11. Code Model
  12. Other SE Tasks

What it can do on your machine

Read from SKILL.md and the folder at commit 4c7dfc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • doi.org
    • aclanthology.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paper Labeler loads about 5.2k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 1,830 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:318
    (`~/.aws/credentials`, environment variables, or IAM role).

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,830 words (~5,157 tokens).

“Classify LLM-for-code research papers from conference proceedings into a two-level label taxonomy.”

— opening of SKILL.md by PurCL
name
paper-labeler

Read the full SKILL.md on GitHub

Files

SKILL.md and 8 other files (scripts) in .claude/skills/paper-labeler of PurCL/ASE.

  • SKILL.md
  • USAGE.md
  • scripts/build_site.py
  • scripts/extract_papers.py
  • scripts/fetch_ndss_abstracts.py
  • scripts/import_original.py
  • scripts/label_papers.py
  • scripts/merge_labeldata.py
  • scripts/process_folder.py

Open the folder on GitHubat commit 4c7dfc7

Compare with similar skills

Paper Labeler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paper Labeler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paper Labeler this skillPurCL/ASE637—~5.2kAutomated safety check: WarnNone
Preprint Search on bioRxivLigphiDonk/Oh-my--paper73812 repos~3.7kAutomated safety check: PassMIT
Paper Research on arXivXiaomiMiMo/MiMo-Code14k—~1.5kAutomated safety check: PassMIT
Conference Paper Recommenderjuliye2025/evil-read-arxiv1.7k—~2.5kAutomated safety check: PassNone
NSFC Literature Review WriterHuiyuLi-2000/Chinese-Grant-Writer-Skills4341 repos~1.4kAutomated safety check: NotesMIT
Autonomous Researchfedericodeponte/opendraft507—~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    738 GitHub starsUsed in 12 repos~3.7k tokens
    Research & ScienceAuto-check passed
  • Paper Research on arXiv

    XiaomiMiMo/MiMo-Code

    Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.

    14k GitHub stars~1.5k tokensUpdated today
    Research & ScienceAuto-check passed
  • Conference Paper Recommender

    juliye2025/evil-read-arxiv

    Searches DBLP for papers from top conferences such as CVPR, ICLR and NeurIPS, adds Semantic Scholar data, scores them and writes a yearly Obsidian note.

    1.7k GitHub stars~2.5k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • NSFC Literature Review Writer

    HuiyuLi-2000/Chinese-Grant-Writer-Skills

    Writes the research-status literature review and critique section of an NSFC grant proposal, backed by a bundled multi-source literature search.

    434 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check: notes
  • Autonomous Research

    federicodeponte/opendraft

    An 18-agent pipeline that turns one topic line into a drafted research paper, literature review, or thesis chapter.

    507 GitHub stars~8.2k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • Literature Reviewer Skill

    stephenlzc/AI-Powered-Literature-Review-Skills

    根据用户提供的论文主题,进行系统性中英文文献回顾(Literature Survey). An agent skill from stephenlzc/AI-Powered-Literature-Review-Skills.

    174 GitHub starsUsed in 1 repo~4.5k tokens
    Research & ScienceAuto-check passed

Questions about Paper Labeler

What does Paper Labeler do?

Extracts research papers from .bib and .html venue files, filters them by relevance, labels them with a two-level taxonomy through Claude and rebuilds a website. The pipeline runs four scripts in sequence.py` added for NDSS to scrape abstracts from paper pages.

When should I use Paper Labeler?

Paper Labeler fits situations like: processing new conference proceedings from .bib or .html files; labeling papers with the two-level taxonomy; adding a new venue to the paper database and rebuilding the site.

How do I install Paper Labeler in Claude Code?

Run `npx skills add PurCL/ASE --skill paper-labeler -a claude-code`. Or copy the skill folder (.claude/skills/paper-labeler in PurCL/ASE) into .claude/skills/paper-labeler in your project. Claude Code loads it when a task matches its description.

How do I install Paper Labeler in Codex?

Run `npx skills add PurCL/ASE --skill paper-labeler -a codex`. Or copy the skill folder (.claude/skills/paper-labeler in PurCL/ASE) into .agents/skills/paper-labeler in your project. Codex loads it when a task matches its description.

Can I use Paper Labeler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PurCL/ASE --skill paper-labeler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paper-labeler, .gemini/skills/paper-labeler, .github/skills/paper-labeler and .opencode/skills/paper-labeler in your project.

What does Paper Labeler need to run?

Going by SKILL.md and its folder, Paper Labeler needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3 to run the scripts in .claude/skills/paper-labeler/scripts/; AWS credentials with access to a Claude model on Bedrock; Raw venue data as .bib or .html files under data/rawdata/.

Does Paper Labeler access the network?

SKILL.md names 2 domains. In commands or code: doi.org and aclanthology.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Paper Labeler safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Paper Labeler use?

No licence was found for Paper Labeler or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Paper Labeler use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Paper Labeler?

Skills that share tags, products or a category with Paper Labeler: Preprint Search on bioRxiv (LigphiDonk/Oh-my--paper, 738 stars), Paper Research on arXiv (XiaomiMiMo/MiMo-Code, 14k stars), Conference Paper Recommender (juliye2025/evil-read-arxiv, 1.7k stars) and NSFC Literature Review Writer (HuiyuLi-2000/Chinese-Grant-Writer-Skills, 434 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paper Labeler?

PurCL (a GitHub organization) maintains it in PurCL/ASE, which has 637 GitHub stars. The repository was last updated on July 17, 2026.

Source: PurCL/ASE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.