Agent skill

Reproduce

by fcakyon in fcakyon/phd-skills

End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

MITAuto-check passedResearch & Science

Install Reproduce

skills CLI
$ npx skills add fcakyon/phd-skills --skill reproduce -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/phd-skills reproduce --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/phd-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/reproduce .claude/skills/reproduce && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reproduce
GitHub stars
415
Token cost
~1.1k tokens
SKILL.md length
393 words
Files
8 (incl. references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

  • The user asks to reproduce
  • SKILL.md covers When to run, Workflow, Working directory layout and Cross-references, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Re-run a paper from scratch

What it does

Reproduce is an agent skill from fcakyon/phd-skills. End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments. Handles missing or partial official code, missing training scripts, missing hyperparameters, and private datasets via similar-public-dataset substitution. Use when the user asks to reproduce, implement, replicate, or re-run a paper from scratch, or pastes an arxiv URL with reproduction intent.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/01-paper-fetch.md`, `references/02-code-clone.md` and `references/03-gap-analysis.md`).

It sits in Research & Science, covering Academic paper search and Database administration. It works with arXiv. The repository describes itself as: PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more. The licence is MIT.

When your agent uses it

  • The user asks to reproduce
  • Re-run a paper from scratch
  • Pastes an arxiv URL with reproduction intent

Example prompts

  • “/reproduce”

What it can do on your machine

Read from SKILL.md and the folder at commit 67acd61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reproduce loads about 1.1k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 393 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/phd-skills at commit 67acd61, republished under its MIT licence (© fcakyon). 393 words, ~1,123 tokens.

Download SKILL.mdSave it as .claude/skills/reproduce/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
reproduce
description
End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments. Handles missing or partial official code, missing training scripts, missing hyperparameters, and private datasets via similar-public-dataset substitution. Use when the user asks to reproduce, implement, replicate, or re-run a paper from scratch, or pastes an arxiv URL with reproduction intent.

Reproduce: paper reproduction from scratch

Reproducing an ML paper often means filling gaps the authors didn't ship, training scripts, hyperparameter tables, augmentation specifics, exact dataset splits. This skill walks seven stages from "I have an arxiv link" to "I have a replication run with measurable delta vs the paper's number."

Each stage has a separate reference file under references/ so this overview stays scannable.

When to run

The user just said any of:

  • "reproduce / implement / replicate / re-run paper X"
  • pasted an arxiv URL with reproduction intent ("can you redo this", "let's try this approach")
  • pointed at an OpenReview / proceedings link with the same intent
  • said "the paper has no code, can we build it"

Workflow

StageWhatReference
1Paper acquisition (arxiv HTML → structured extract)references/01-paper-fetch.md
2Existing code discovery + inventoryreferences/02-code-clone.md
3Gap analysis (extract every missing hyperparam from the prose)references/03-gap-analysis.md
4Implementation (uv venv, fill gaps, commit per gap)references/04-implement.md
5Dataset acquisition (HF datasets first; substitute if private)references/05-dataset.md
6Smoke runs (forward pass → 1 step → 20 iters)references/06-smoke.md
7Replication runs + comparison at paper's reported epochsreferences/07-replicate.md

Walk them in order. Each stage has its own success criteria; do not advance to the next until the current one passes.

Working directory layout

For each paper reproduction, set up a dedicated workspace:

repro/<paper-arxiv-id>/
├── paper.md              # structured extract from stage 1
├── inventory.md          # what exists / missing from stage 2
├── gaps_filled.md        # hyperparam table with provenance from stage 3
├── code/                 # implementation from stage 4 (or cloned + extended)
├── data/                 # dataset symlinks or actual data from stage 5
├── dataset_substitution.md  # if a public dataset stood in for a private one
├── smoke_logs/           # outputs from stage 6
└── results.md            # replication outcomes from stage 7

This keeps reproductions self-contained and easy to revisit later.

Show full SKILL.md (168 more words)Show less

Cross-references

  • After stage 3, hand the gap analysis off to the paper-verification skill for a round-trip check ("did I really capture every hyperparam the paper mentions").
  • Stage 4 implementation should be committed in small, reviewable pieces: each commit references the paper section that justified the filled value.
  • Stage 6 smoke failures route to the /phd-skills:debug skill, not to ad-hoc fixes.
  • Stage 7 launches go through the /phd-skills:launch checklist before any multi-hour run.
  • Stage 7 comparisons go through the /phd-skills:compare skill at the paper's reported epochs (never current-vs-final).

Output

For each reproduction, the final artifact is results.md with absolute deltas (not just %) and one of three labels per metric:

  • [matched within 0.X pp]: within the paper's reported variance
  • [gap, hypothesis: ...]: measurable underperformance, with a stated hypothesis for the cause
  • [fundamental disagreement, see X]: the result and the paper's claim are inconsistent in a way that needs investigation, not just more compute

If the workspace is on a public repo, link the workspace README from the project's main reproduction-tracking doc.

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in plugin/skills/reproduce of fcakyon/phd-skills.

  • SKILL.md
  • references/01-paper-fetch.md
  • references/02-code-clone.md
  • references/03-gap-analysis.md
  • references/04-implement.md
  • references/05-dataset.md
  • references/06-smoke.md
  • references/07-replicate.md

Open the folder on GitHubat commit 67acd61

Compare with similar skills

Reproduce next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reproduce compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reproduce this skillfcakyon/phd-skills415—~1.1kAutomated safety check: PassMIT
Read arXiv Paperkarpathy/nanochat59k1 repos~494Automated safety check: PassMIT
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Openalex Databaseneflibata-feng/MyArxiv-Agent12612 repos~3kAutomated safety check: PassCustom licence
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Citation Managementneflibata-feng/MyArxiv-Agent12619 repos~8.1kAutomated safety check: NotesMIT

Similar skills

  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    59k GitHub starsUsed in 1 repo~494 tokens
    Research & ScienceAuto-check passed
  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Openalex Database

    neflibata-feng/MyArxiv-Agent

    Query and analyze scholarly literature using the OpenAlex database.

    126 GitHub starsUsed in 12 repos~3k tokens
    Research & ScienceAuto-check passed
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Citation Management

    neflibata-feng/MyArxiv-Agent

    Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.

    126 GitHub starsUsed in 19 repos~8.1k tokens
    Research & ScienceAuto-check: notes
  • Searches arXiv across many papers on one topic, extracts each paper's methodology and findings in parallel, and synthesizes a cited literature review.

    84k GitHub starsUsed in 2 repos~4.3k tokens
    Research & ScienceAuto-check passed

More from fcakyon/phd-skills

All 12 skills in this repo
  • Compare

    fcakyon/phd-skills

    Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

    415 GitHub stars~1.2k tokensUpdated 24 days ago
    Auto-check passed
  • Debug

    fcakyon/phd-skills

    Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

    415 GitHub stars~1.3k tokensUpdated 24 days ago
    Auto-check passed
  • Experiment Design

    fcakyon/phd-skills

    A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

    415 GitHub stars~987 tokensUpdated 24 days ago
    Auto-check passed
  • Latex Setup

    fcakyon/phd-skills

    A skill your agent uses when the user wants to set up or troubleshoot a LaTeX environment, choose between biber and bibtex, install packages for a specific venue template, or configure compilation.

    415 GitHub stars~1.1k tokensUpdated 24 days ago
    Auto-check: notes
  • Launch

    fcakyon/phd-skills

    Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

    415 GitHub stars~1.4k tokensUpdated 24 days ago
    Auto-check passed
  • Literature Research

    fcakyon/phd-skills

    A skill your agent uses when the user wants to find related work, survey a research area, identify literature gaps, or discover open-source implementations.

    415 GitHub stars~999 tokensUpdated 24 days ago
    Auto-check passed

Works with

Questions about Reproduce

What does Reproduce do?

End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments. Reproduce is an agent skill from fcakyon/phd-skills. End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

When should I use Reproduce?

Reproduce fits situations like: the user asks to reproduce; re-run a paper from scratch; pastes an arxiv URL with reproduction intent.

How do I install Reproduce in Claude Code?

Run `npx skills add fcakyon/phd-skills --skill reproduce -a claude-code`. Or copy the skill folder (plugin/skills/reproduce in fcakyon/phd-skills) into .claude/skills/reproduce in your project. Claude Code loads it when a task matches its description.

How do I install Reproduce in Codex?

Run `npx skills add fcakyon/phd-skills --skill reproduce -a codex`. Or copy the skill folder (plugin/skills/reproduce in fcakyon/phd-skills) into .agents/skills/reproduce in your project. Codex loads it when a task matches its description.

Can I use Reproduce in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/phd-skills --skill reproduce -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reproduce, .gemini/skills/reproduce, .github/skills/reproduce and .opencode/skills/reproduce in your project.

What does Reproduce need to run?

SKILL.md names no scripts, command-line tools or credentials: Reproduce is instructions for the agent only.

Does Reproduce access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Reproduce safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Reproduce use?

Reproduce is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reproduce use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.

What are the alternatives to Reproduce?

Skills that share tags, products or a category with Reproduce: Read arXiv Paper (karpathy/nanochat, 59k stars), Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Openalex Database (neflibata-feng/MyArxiv-Agent, 126 stars) and Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reproduce?

fcakyon (a GitHub user) maintains it in fcakyon/phd-skills, which has 415 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 16, 2026.

Source: fcakyon/phd-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.