Papers Skill
sickn33/agentic-awesome-skills
Skill for academic research workflows: search Semantic Scholar (200M+ papers), inspect citations, download arXiv PDFs, and extract PDF text.
A skill your agent uses when the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI or PDF input), OR a custom batch of papers (Mode B…
$ npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ltczding-gif/ref-downloader ref-downloader --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ltczding-gif/ref-downloader.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ref-downloader .claude/skills/ref-downloader && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ref-downloader" agent skill from https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloader into .claude/skills/ref-downloader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ref-downloader", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloaderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ltczding-gif/ref-downloader ref-downloader --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ltczding-gif/ref-downloader.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ref-downloader .agents/skills/ref-downloader && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ref-downloader" agent skill from https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloader into .agents/skills/ref-downloader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ref-downloader", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ltczding-gif/ref-downloader ref-downloader --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ltczding-gif/ref-downloader.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ref-downloader .cursor/skills/ref-downloader && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ref-downloader" agent skill from https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloader into .cursor/skills/ref-downloader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ref-downloader", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ltczding-gif/ref-downloader.git --path skills/ref-downloader--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ltczding-gif/ref-downloader ref-downloader --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ltczding-gif/ref-downloader.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ref-downloader .gemini/skills/ref-downloader && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ref-downloader" agent skill from https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloader into .gemini/skills/ref-downloader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ref-downloader", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ltczding-gif/ref-downloader ref-downloaderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ltczding-gif/ref-downloader.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ref-downloader .github/skills/ref-downloader && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ref-downloader" agent skill from https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloader into .github/skills/ref-downloader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ref-downloader", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ltczding-gif/ref-downloader ref-downloader --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ltczding-gif/ref-downloader.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ref-downloader .opencode/skills/ref-downloader && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ref-downloader" agent skill from https://github.com/ltczding-gif/ref-downloader/tree/main/skills/ref-downloader into .opencode/skills/ref-downloader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ref-downloader", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ref-downloaderA skill your agent uses when the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI or PDF input), OR a custom batch of papers (Mode B…
Ref Downloader is an agent skill from ltczding-gif/ref-downloader. Use when the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI or PDF input), OR a custom batch of papers (Mode B: DOI/title/arXiv-PMID list, or abstract query like "Author X's recent papers"). Not for one-off PDFs, paper search, or Zotero import.
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/agent-runbook.md`, `scripts/_config.py` and `scripts/download_refs.py`).
It sits in Research & Science, covering Academic paper search, Citation management and PDF. It works with Zotero, arXiv, Python and Playwright. The repository describes itself as: Batch-download reference PDFs from a DOI or paper PDF using Crossref and your institutional Edge session. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1cc6673. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipplaywrightjustFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
doi.orgdx.doi.orgapi.crossref.orgAlso links to:
github.compypi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ref Downloader loads about 5.9k tokens when it runs, and up to ~9.5k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 2,382 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ltczding-gif/ref-downloader at commit 1cc6673, republished under its MIT licence (© ltczding-gif). 2,382 words, ~5,889 tokens.
.claude/skills/ref-downloader/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Slim entry for agent mode. The full 8-step manual runbook with code snippets for Mode A debug +
PUBLISHER_MAPextension procedure lives in references/agent-runbook.md. Human users see ../../README.md.
<SKILL_DIR> = this folder (skills/ref-downloader in the source repo,
or wherever the user copied this skill — e.g.
~/.claude/skills/ref-downloader/). Python scripts live in
<SKILL_DIR>/scripts/; config files (config.example.toml,
config.local.toml) live at <SKILL_DIR>/.
This skill handles two flows. Pick before running.
Mode A — Reference-list download (original use case). User
provides ONE paper (DOI or local PDF) and wants "all of its
references". Pipeline: extract_refs.py → validate_refs.py →
download_refs.py.
Mode B — Custom batch download. User provides their own batch
of papers — DOIs, paper titles, non-DOI identifiers (arXiv / PMID
/ Semantic Scholar IDs), OR an abstract query ("Smith 在 Google
Scholar 上的文章" / "Nature Energy 2023 papers"). The agent
resolves whatever was given to DOIs, then runs validate_refs.py
→ download_refs.py directly. Skip the wrapper —
run_ref_downloader.py assumes a parent DOI and will fail.
Both modes share install, config, per-publisher strategies, failure modes, output layout, and the CloakBrowser opt-in backend.
| User input shape | Mode | Sub-flow |
|---|---|---|
| One DOI/PDF + "all refs of" / "全部参考文献" / "把这篇引用都下了" | A | — |
≥2 DOIs in input (any wrapping: bare / {} / https://doi.org/… / dx.doi.org/…; ASCII or full-width slashes) | B | B.1 (after canonicalize) |
Non-DOI IDs only: arXiv: / PMID: / S2: / corpusId: | B | B.0 normalize → B.1 |
| Title list ("下载这几篇:title1, title2, …") | B | B.2 |
| Abstract query (author / topic / journal+year / "Google Scholar 上 …") | B | B.3 |
| Mixed (DOIs + titles + IDs + queries) | B | run each, merge |
| Single DOI without "of refs" qualifier | B | B.1 single-item |
| Title + author + year for ONE paper ("Smith 2024 Nature paper on X") | B | B.2 (specific paper, lookup) |
| Open-ended query for a corpus ("Smith 2024 之后所有的 Nature 文章") | B | B.3 (discovery) |
| Insufficient resolvable content ("上次给你的那 5 篇" / pure pronouns) | — | Ask user to repaste / attach file; do NOT guess |
| Genuinely ambiguous A vs B | — | ask user |
Key disambiguators:
Trigger phrases:
10.x/y form) or local PDF path and asks for
"all references" / "全部参考文献"Don't invoke for:
python "<SKILL_DIR>/scripts/run_ref_downloader.py" <DOI_OR_PDF_PATH>The wrapper handles DOI resolution (Zotero → fitz fallback),
output-dir layout, sequential 3-stage pipeline (extract_refs.py →
validate_refs.py → download_refs.py), and end-of-run cleanup.
Useful flags:
--yes — non-interactive (CI/batch), overwrite prompts default-yes--auto — forwarded to download_refs.py: skip "press Enter"
confirm + shorter challenge wait + async retry queue for
manual_pending refs (60s delay, single retry, max 3 concurrent).
Use for CI / overnight runs; not for sessions where you want to
drive captchas yourself.--fail-fast — terminate after first actionable unresolved ref
(useful in CI to surface real failures fast)--output-dir <path> — override default output location--config <path> — alternate TOML config (overrides
config.local.toml)即将下载参考文献:DOI=<doi>msedge.exe processes killed (Task
Manager check). The script claims the user's persistent Edge
profile and needs exclusive access. (Cloak backend skips this.)<SKILL_DIR>/config.local.toml with [crossref].mailto. Missing
config → wrapper prints a WARNING but continues with placeholder
defaults.<cwd>/<project_name>_refs/. For PDF input:
<pdf_dir>/<pdf_stem>_refs/. Override with --output-dir.Input variability is the point. Don't refuse — route. The agent handles whatever shape the user gave (paste, file, prose, BibTeX, RIS, abstract query) and resolves it to a clean DOI list before handing off to the pipeline.
Before any extraction or routing, normalize the input string:
{}, <>, (), [], and quote marks.https://doi.org/, http://doi.org/,
https://dx.doi.org/, http://dx.doi.org/ → bare DOI./ (U+FF0F) → /: (U+FF1A) → :. (U+FF0E) → ."" '' → straight " '.,;)}"' (note }).Apply Step 0 BEFORE the regex pass in B.1 AND before the canonical dedupe compare in step 5 of the main flow.
For each non-DOI identifier the user gave:
| Input | Action |
|---|---|
arXiv:2401.12345 or bare arXiv ID | Use 10.48550/arXiv.<id> (canonical), but prefer a journal DOI if the agent can discover one via Crossref query.bibliographic=<arXiv_id> |
PMID:12345678 or pubmed.ncbi.nlm.nih.gov/12345678 | Hit eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=<pmid> → grab articleids[type=doi] |
Semantic Scholar paper ID (S2:abc... / corpusId:N) | Hit api.semanticscholar.org/graph/v1/paper/<id>?fields=externalIds → grab externalIds.DOI |
| Anything else non-DOI shaped | Leave for B.1 regex pass to ignore; if it survives B.1+B.2 unresolved, drop with skipped (unresolvable_identifier) in the confirm table — do NOT auto-fire requests on garbage |
Network preflight: before firing B.0 lookups, do one cheap
sanity probe (e.g. HEAD https://api.crossref.org/). If it fails,
tell the user "no network — Mode B can't resolve non-DOI identifiers
or do discovery; only direct DOI extraction will work" and let them
decide to proceed with B.1 only.
Semantic Scholar rate limit: unauthenticated ≤ 1 req/sec. Pace
batches; on 429 back off 30s then retry once; on second 429, drop
the entry with skipped (ss_rate_limited).
After Step 0 canonicalization:
10\.\d{4,9}/[^\s,;<>"'{}]+ on the canonicalized input
(or file contents). The character class explicitly excludes {}
so wrappers like BibTeX doi = {10.x/y} don't leak braces..,;)}"'.@article entries but only
5 have ASCII DOIs), send the remaining entries to B.2 title
lookup rather than silently dropping them.For each title:
GET https://api.crossref.org/works?query.title=<urlencode>&rows=5&mailto=<config crossref.mailto>score. Note: Crossref score is unbounded
relevance, NOT 0–100 — absolute thresholds across queries don't
compare. Use relative + content rules:top1.score / top2.score < 1.5 → ambiguous; show user top 3.matched_title vs input_title by
token overlap (or Levenshtein). If overlap < 50%, mark
low-confidence regardless of score ratio.author[].family and issued.date-parts.confidence=high (top1, ratio ≥ 1.5, overlap ≥ 50%, any
author/year match consistent) — included by default.confidence=low (any of: ratio < 1.5, overlap < 50%, no author
match) — excluded by default; user must explicitly pick.unresolved (0 candidates or all rejected) — dropped with
skipped (no_match).Triggered by queries like "Smith 在 Google Scholar 上的文章", "topic Y top 20", "Nature Energy 2023". Do NOT scrape Google Scholar (anti-bot + ToS). Interpret "Scholar" semantically and use the ladder below.
Tool ladder (try in order, use what's available):
api.crossref.org/works?query.author= /
query.bibliographic= / query.container-title=) — always
available, no auth.api.openalex.org/works?search= or
?filter=author.id:A...) — free, no auth, broader coverage than
Crossref author search, returns DOIs directly.api.semanticscholar.org/graph/v1/paper/search?query=...) —
better for topic / abstract search. Rate limit ~1 req/sec
unauthenticated; pace requests, on 429 back off 30s then drop
the query on a second 429.bio-research:pubmed MCP) — if biomedical AND the
MCP is loaded in the host framework.web-search-*
skill is loaded. Highest hallucination risk; agent MUST
round-trip every candidate through Crossref or OpenAlex to verify
the DOI exists before accepting.After discovery: present candidates as a numbered list with
title + first-author + year + DOI + source (which API found it).
Default top 20; ask if user wants more. User strikes out / picks
subset → final list locked.
Open-ended-query clarifier: if the agent's discovery would return more than 50 candidates (e.g. user said "Smith 的所有文章" and the author has 200+ publications), confirm scope with user BEFORE returning — "found 200+; you want all of them, top 20 most cited, or filter by year?".
10.X/ABC} and 10.x/abc must collapse to one entry.找到 N 个唯一 DOI(去重后)。完整列表:
[ 1] doi=10.xxxx/yyy source=B.1 confidence=high
title= ... author= ... year= ...
[ 2] doi=10.zzzz/www source=B.2 confidence=high
matched_title= ... (input: "...")
[ 3] doi=10.aaaa/bbb source=B.3 confidence=high
via=Crossref (query: "...")
[ 4] doi=10.cccc/ddd source=B.2 confidence=LOW
matched_title= ... (input: "...") ← excluded; pick to include
[ 5] (unresolvable) source=B.0 from: "PMID:99999"
← dropped
...
开始下载吗?(y=accept all high-confidence / n=cancel /
include 4 / exclude 1,3 / show <N> / ...)confidence=high rows only. confidence=low
excluded unless user explicitly includes. Unresolvables dropped.groupmtg_<date>; "Smith 综述补充" → smith_review_extras;
nothing topical → custom_<date>). Ask user confirm.<OUTPUT_DIR>/<project_name>/refs_raw.json
exists, ask append / new / rename (default: ask again on any
other input — DO NOT default-append). Append rules:id = max(existing_ids) + 1.validate_refs.py keys
its incremental skip on id, renumbering re-assigns prior
verified metadata to the wrong DOI. Only verified rows are
skipped; failed/pending rows revalidate on re-run.refs_raw.json (heredoc the agent runs):import json
from datetime import datetime
dois = [...] # finalized canonical-lowercase list
start_id = 1 # or max(existing_ids)+1 in append mode
data = {
"parent_doi": "", # empty string for clean report labels;
# validate_refs.py reads as raw JSON, null
# would also work but "" is preferred.
"parent_title": f"Custom batch — {user_label}",
"extracted_at": datetime.now().isoformat(timespec="seconds"),
"total": len(dois),
"with_doi": len(dois),
"without_doi": 0,
"references": [
{"id": i, "doi": d,
"key": "", "unstructured": "",
"author": "", "year": "", "journal": "",
"volume": "", "first_page": ""}
for i, d in enumerate(dois, start=start_id)
],
}
with open(f"{project_name}/refs_raw.json", "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=2)validate_refs.py fills them
from Crossref per DOI on success. Rows whose DOI Crossref can't
resolve become status=failed with empty metadata (not partially
enriched).cd <OUTPUT_DIR>
python <SKILL_DIR>/scripts/validate_refs.py <project_name>
python <SKILL_DIR>/scripts/download_refs.py <project_name> [--auto] [--fail-fast][crossref].mailto for polite-pool latency.api.crossref.org HEAD probe).--auto: works with Mode B (manual_pending refs go to async
retry queue same as Mode A).--fail-fast: works with Mode B (stops on first actionable
unresolved ref).[user].verified_no_si_dois: works with Mode B (matches by
lowercase DOI — independent of how refs_raw.json was produced).REF_DOWNLOADER_BROWSER=cloak): works
with Mode B (browser backend is decided by env var,
independent of input mode).The skill protocol can't manage Python deps. If
python -c "import playwright" fails, the user needs:
cd "<SKILL_DIR>"
pip install playwright pymupdf
playwright install msedge # downloads Edge driver
cp config.example.toml config.local.toml # then user edits [crossref].mailtoIf the user is developing from the source repo instead of an installed
skill copy, they can also install from the repo root with
pip install -r requirements.txt -r requirements-dev.txt.
Default backend is Microsoft Edge. For sites that keep blocking
ordinary Playwright (Cloudflare Turnstile, Radware, persistent
Just a moment / 安全验证 pages), switch to the CloakBrowser stealth
Chromium backend.
What CloakBrowser is. Third-party MIT-licensed Python package by
CloakHQ (github.com/CloakHQ/CloakBrowser,
pypi:cloakbrowser). Ships a
patched Chromium build with anti-fingerprint changes. Its
launch_persistent_context_async() is Playwright-API-compatible,
which is why ref-downloader can swap it in with one env var. NOT a
dependency of ref-downloader — if the user doesn't
pip install cloakbrowser, it's never imported and the default Edge
path runs as normal. Beta software; user installs it at their own
discretion.
# One-time setup (separate from ref-downloader's `pip install playwright pymupdf`)
pip install cloakbrowser
# Switch backend (env vars; no CLI flag changes)
$env:REF_DOWNLOADER_BROWSER = "cloak"
$env:REF_DOWNLOADER_CLOAK_HUMAN_PRESET = "careful" # optional: slower mouse/scroll
# Optional overrides:
# $env:REF_DOWNLOADER_CLOAK_PROFILE = "<custom path>" # default: ~/.local/cloakbrowser/profiles/ref-downloader
# $env:REF_DOWNLOADER_CLOAK_PROXY = "http://..."
# $env:REF_DOWNLOADER_CLOAK_GEOIP = "1"
# $env:CLOAKBROWSER_PYTHONPATH = "<dev source>" # sys.path hint if cloakbrowser is checked out, not pip-installed
python "<SKILL_DIR>/scripts/download_refs.py" <PROJECT_NAME>Caveats:
REF_DOWNLOADER_CLOAK_PROFILE and finishing
any verification before running the downloader.human_preset=careful lowers behavior-detection trigger rates but
is not a captcha solver.Same for both modes:
<OUTPUT_DIR>/
├── <PROJECT_NAME>/
│ ├── refs_raw.json # extract_refs.py output (Mode A) or
│ │ # hand-built JSON (Mode B)
│ ├── refs_validated.json # validate_refs.py output
│ ├── download_report.csv # per-ref status (only on graceful
│ │ # completion; OVERWRITTEN each run —
│ │ # NOT historical truth)
│ ├── *.pdf # reference PDFs
│ └── *_SI.pdf # supplementary files (where supported)
└── runs/<timestamp>-round-03/
└── events.jsonl # full event trace per ref
# (append-only across runs;
# THIS is the authoritative history)Interruption note: if the run is interrupted (Ctrl+C / Edge
crash / VPN drop), the root download_report.csv may be stale.
Trust the latest runs/<timestamp>/events.jsonl + actual files in
<PROJECT_NAME>/.
| Status / symptom | Meaning | Action |
|---|---|---|
manual_pending (auth_redirect) | Bounced to institution SSO | User signs in via live Edge tab; re-run (incremental skips done refs) |
manual_pending (challenge_timeout) | Cloudflare / publisher challenge unsolved in time | Re-run interactively; solve captcha when prompted |
manual_pending (elsevier_crasolve_shell) | Elsevier viewer stuck in transition | In --auto mode the async retry queue picks it up ~60s later; in interactive mode the hot-session retry usually catches it, else manual click in live page |
failed (auto) | Generic auto path failed | Check events.jsonl for that ref; may need a publisher-specific patch |
ignored (ignored_institution_access) | DOI listed in [institution].ignored_access_dois | Skip-by-design; remove from config to retry |
| Edge won't launch | Background msedge.exe still holding profile | Kill all msedge.exe in Task Manager, re-run (cloak backend skips this) |
ModuleNotFoundError: playwright | Install prereqs not done | See "Install prerequisites" section above |
WARNING: crossref.mailto is the placeholder | First-run config uncustomized | Edit <SKILL_DIR>/config.local.toml → set [crossref].mailto to a real email (Crossref polite pool) |
| Mode B: Step 0 left full-width slash unconverted | Canonicalization bug | Verify Step 0 ran before regex; flag for design fix |
Mode B: BibTeX doi = {10.x/y} left trailing } in refs_raw.json | Step 0 bypassed | Step 0 MUST run before B.1 regex |
| Mode B: Crossref title query 0 hits for a B.2 row | Title couldn't match | Drop entry as unresolved; suggest user provide author/year/journal |
| Mode B: B.3 abstract query returns 0 across all ladder steps | No matches found | Suggest user narrow (add author / year / journal); or accept that no papers match |
| Mode B: B.3 discovery returns 200+ candidates | Query too broad | Ask user to scope (year range / top-N by citations / specific journal) BEFORE listing |
Mode B: run_ref_downloader.py invoked accidentally | It assumes parent DOI — will fail | Direct validate_refs.py + download_refs.py invocation only |
| Mode B: Semantic Scholar 429 | Unauthenticated rate limit hit | Back off 30s, retry once; on second 429, drop with skipped (ss_rate_limited) |
Mode B: PMID lookup returns no articleids[type=doi] | PubMed has no DOI for this entry | Entry has no DOI; tell user, suggest alternative identifier |
| Mode B: Input is purely conversational ("上次那 5 篇") | No resolvable content | Refuse; ask user to repaste / attach file |
| Mode B: No network reachable | B.0 / B.2 / B.3 all need network | Tell user before starting; only B.1 viable |
If the wrapper fails partway (Mode A), run the 3 scripts standalone for partial re-execution:
python <SKILL_DIR>/scripts/extract_refs.py <DOI> # → refs_raw.json
python <SKILL_DIR>/scripts/validate_refs.py <PROJECT> # → refs_validated.json
python <SKILL_DIR>/scripts/download_refs.py <PROJECT> # → PDFs + download_report.csvMode B uses the same standalone invocation, just skipping
extract_refs.py (the agent builds refs_raw.json directly).
Full 8-step manual flow with code snippets, DOI-resolution fallback
chain (Zotero query → fitz text → user prompt), and the procedure
for extending PUBLISHER_MAP when encountering an unknown DOI
prefix → references/agent-runbook.md.
© ltczding-gif, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in skills/ref-downloader of ltczding-gif/ref-downloader.
Open the folder on GitHubat commit 1cc6673
Ref Downloader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ref Downloader this skillltczding-gif/ref-downloader | 139 | — | ~5.9k | Automated safety check: Pass | MIT | |
| Papers Skillsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Literature Review Toolsbrycewang-stanford/Auto-Empirical-Research-Skills | 4.6k | — | ~2.5k | Automated safety check: Notes | Custom licence | |
| Citation ManagementK-Dense-AI/claude-scientific-writer | 2.4k | 2 repos | ~3.9k | Automated safety check: Notes | MIT | |
| Paper2codePrathamLearnsToCode/paper2code | 1.5k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Paper Research on arXivXiaomiMiMo/MiMo-Code | 14k | — | ~1.5k | Automated safety check: Pass | MIT |
sickn33/agentic-awesome-skills
Skill for academic research workflows: search Semantic Scholar (200M+ papers), inspect citations, download arXiv PDFs, and extract PDF text.
brycewang-stanford/Auto-Empirical-Research-Skills
Recommend AND run open-source AI tools, agents, Claude Code / Codex skills, and MCP servers for any stage of a literature review — searching, reading, extracting, synthesizing, screening…
K-Dense-AI/claude-scientific-writer
Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.
PrathamLearnsToCode/paper2code
Converts an arxiv paper into a minimal, citation-anchored Python implementation.
XiaomiMiMo/MiMo-Code
Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.
917Dhj/DeepPaperNote
Generate a high-quality deep-reading note for a single paper and write it into an Obsidian-style vault.
Works with
Categories
A skill your agent uses when the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI or PDF input), OR a custom batch of papers (Mode B…. Ref Downloader is an agent skill from ltczding-gif/ref-downloader. Use when the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI or PDF input), OR a custom batch of papers (Mode B: DOI/title/arXiv-PMID list, or abstract query like "Author X's recent papers").
Ref Downloader fits situations like: the user asks to batch-download academic PDFs with ref-downloader — either ALL references of one paper (Mode A: DOI; A custom batch of papers (Mode B: DOI/title/arXiv-PMID list; abstract query like Author Xs recent papers).
Run `npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a claude-code`. Or copy the skill folder (skills/ref-downloader in ltczding-gif/ref-downloader) into .claude/skills/ref-downloader in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a codex`. Or copy the skill folder (skills/ref-downloader in ltczding-gif/ref-downloader) into .agents/skills/ref-downloader in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ltczding-gif/ref-downloader --skill ref-downloader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ref-downloader, .gemini/skills/ref-downloader, .github/skills/ref-downloader and .opencode/skills/ref-downloader in your project.
Going by SKILL.md and its folder, Ref Downloader needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip, playwright and just). Our summary lists: Python 3.
SKILL.md names 5 domains. In commands or code: doi.org, dx.doi.org and api.crossref.org; the agent is likely to contact these when it follows the instructions. As links in the text: github.com and pypi.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ref Downloader is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ref Downloader: Papers Skill (sickn33/agentic-awesome-skills, 47k stars), Literature Review Tools (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars) and Paper2code (PrathamLearnsToCode/paper2code, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ltczding-gif (a GitHub user) maintains it in ltczding-gif/ref-downloader, which has 139 GitHub stars. The repository was last updated on June 1, 2026.
Source: ltczding-gif/ref-downloader on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.