Agent skill

Sigmod Reproducibility

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when hardening the reproducibility story of a SIGMOD submission, covering PACMMOD's expectation that code, data, scripts, and notebooks be shared, experiment provenance from…

MITAuto-check passedResearch & Science

Install Sigmod Reproducibility

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill sigmod-reproducibility -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills sigmod-reproducibility --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/SIGMOD-Skills/skills/sigmod-reproducibility .claude/skills/sigmod-reproducibility && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sigmod-reproducibility
GitHub stars
1.2k
Token cost
~1.5k tokens
SKILL.md length
612 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when hardening the reproducibility story of a SIGMOD submission, covering PACMMOD's expectation that code, data, scripts, and notebooks be shared, experiment provenance from…

  • Hardening the reproducibility story of a SIGMOD submission
  • SKILL.md covers Provenance chain, not vibes, Systems numbers need…, Fairness to baselines is… and Repro debt ledger, plus 5 more sections
  • Calls git
  • Covering PACMMODs expectation that code

What it does

Sigmod Reproducibility is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when hardening the reproducibility story of a SIGMOD submission, covering PACMMOD's expectation that code, data, scripts, and notebooks be shared, experiment provenance from config to figure, dataset and workload disclosure, variance reporting for systems numbers, and alignment with later ARI badging.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Reproducible research. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Hardening the reproducibility story of a SIGMOD submission
  • Covering PACMMODs expectation that code
  • Notebooks be shared
  • Experiment provenance from config to figure

Example prompts

  • “/sigmod-reproducibility”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sigmod Reproducibility loads about 1.5k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 612 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 612 words, ~1,499 tokens.

Download SKILL.mdSave it as .claude/skills/sigmod-reproducibility/SKILL.md (or your agent's skills folder).
name
sigmod-reproducibility
description
Use when hardening the reproducibility story of a SIGMOD submission, covering PACMMOD's expectation that code, data, scripts, and notebooks be shared, experiment provenance from config to figure, dataset and workload disclosure, variance reporting for systems numbers, and alignment with later ARI badging.

SIGMOD Reproducibility

PACMMOD's author guidelines state the culture plainly: sharing research artifacts should be the norm, and papers are expected to make code, data, scripts, and notebooks available where possible — encouraged rather than mandatory for acceptance, but reviewers read availability as a credibility signal. This skill covers reproducibility as engineered into the paper; the post-acceptance badge process lives in sigmod-artifact-evaluation.

Provenance chain, not vibes

A database paper's result is a function of code version, configuration, dataset, workload, and hardware. Reproducibility means the paper pins all five for every number it prints.

LayerMust be recoverable from paper + artifactWhere it usually hides
Code versionCommit or tag behind each experiment"latest" at submission time
ConfigurationBuffer sizes, thread counts, compaction/GC settings, flagsDefaults nobody recorded
DatasetSource, version, generator seed, scale factor"standard benchmark data"
WorkloadQuery mix, arrival pattern, skew parameters, warm/cold stateThe harness script
HardwareCPU model, cores, RAM, storage class, network, OS/kernelA single sentence, if any

The practical test: could a competent stranger fill this whole table from your materials? Every empty cell is a review question waiting to be asked in the feedback phase.

Systems numbers need distributions

Throughput and latency are noisy. The SIGMOD-credible floor:

  • Repeated runs with the repetition count stated; medians or means with a spread measure, never a single lucky run.
  • Latency reported at named percentiles (p50/p95/p99), since tail behavior is often the actual contribution.
  • Warm-up policy stated: what was cached, JIT-compiled, or pre-compacted before measurement started.
  • Cross-run controls named: same machine, isolated tenancy, pinned cores, disabled turbo — whatever was actually done, said explicitly.

Fairness to baselines is reproducibility too

The most contested numbers in a data-systems paper are the competitor's. Record and disclose: which version of each baseline, who tuned it and how, which of its features were enabled, and whether its authors' recommended configuration was used. An artifact that reproduces your system but ships an untuned strawman baseline reproduces the wrong thing.

Repro debt ledger

Track gaps while writing rather than reconstructing at deadline:

text
# repro-ledger.md (kept in the paper repo)
| Paper item | Script | Data pinned | Config pinned | Runs/variance | Status |
|-----------|--------|-------------|---------------|---------------|--------|
| Fig 6     | exp/f6.sh | yes (sf=100, seed 41) | yes | 5 runs, p50/p99 | OK |
| Fig 7     | exp/f7.sh | yes | NO — flags undocumented | 1 run | DEBT |
| Tab 2     | manual | partial | yes | n/a | DEBT: script it |

Rule: nothing with status DEBT appears in the submitted PDF. The ledger later becomes the ARI claims map almost verbatim.

Show full SKILL.md (256 more words)Show less

Pinning, mechanically

The cheapest insurance is captured at run time, not reconstructed later:

bash
# stamped into every experiment's log directory by the harness
git -C "$ENGINE_DIR" rev-parse HEAD          > meta/engine.commit
"$ENGINE" --dump-effective-config             > meta/config.effective
sha256sum data/*.bin                          > meta/datasets.sha256
lscpu; free -h; uname -r; nvme list          >> meta/hardware.txt
echo "$WORKLOAD_SEED $QUERY_MIX $SKEW"        > meta/workload.params

Five commands, and the provenance table fills itself for every figure. The --dump-effective-config line matters most: defaults you never set are still part of the experiment, and engines change defaults between versions.

Data you cannot publish

Industrial collaborations often involve proprietary workloads. The accepted pattern at SIGMOD: characterize the private data (sizes, distributions, skew, schema shape), provide a public or synthetic stand-in that exhibits the same phenomena, run headline experiments on both, and say which conclusions are supported by the public path alone. A paper whose every claim requires private data cannot be badged and will be read skeptically.

Multi-round consistency

Because SIGMOD reviewing spans rounds, reproducibility discipline has a second job: your own revision. Reviewers may demand new experiments with a one-month window — regenerating the whole evaluation under a new flag is only survivable if the original runs were scripted, seeded, and logged. Teams that hand-ran their plots discover this during the revision, at the worst time.

What reviewers can check without running anything

Even reviewers who never open the artifact test reproducibility passively: do the numbers in the abstract, the results section, and the conclusion agree; do figure axes and caption units match the prose; does the claimed hardware plausibly fit the claimed dataset in memory; do percentages in tables sum sensibly. Internal inconsistency is read as evidence that the pipeline is hand-operated — and it usually is. A final numeric-consistency pass over the PDF is reproducibility work, not copyediting.

Output format

text
[Sharing posture] full artifact / partial / withheld with stated reason
[Provenance table] complete cells vs. gaps, per figure and table
[Variance floor] repetitions, spread measures, percentile reporting
[Baseline fairness] versions, tuning provenance, config disclosure
[Private-data plan] public stand-in coverage of headline claims
[Debt ledger] items that must close before the round deadline

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in SIGMOD-Skills/skills/sigmod-reproducibility of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Sigmod Reproducibility next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sigmod Reproducibility compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sigmod Reproducibility this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Peer ReviewK-Dense-AI/claude-scientific-writer2.4k2 repos~3.1kAutomated safety check: NotesMIT
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Compute Environment Setupaipoch/open-science5.5k—~2.6kAutomated safety check: PassApache-2.0
Figure Styleaipoch/open-science5.5k—~5.1kAutomated safety check: PassApache-2.0
Add Bactopia Toolbactopia/bactopia522—~4.1kAutomated safety check: PassMIT

Similar skills

  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Compute Environment Setup

    aipoch/open-science

    Prepares setup instructions and a named activation file for a user-managed software environment on an Open-Science SSH or Slurm compute host.

    5.5k GitHub stars~2.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Figure Style

    aipoch/open-science

    Publication-grade correctness and legibility rules for final-deliverable scientific figures, not exploratory plots.

    5.5k GitHub stars~5.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Add Bactopia Tool

    bactopia/bactopia

    Scaffold a complete Bactopia Tool across all three tiers -- module, subworkflow, and workflow entry point under workflows/bactopia-tools/.

    522 GitHub stars~4.1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Modeling Code and Result Contracts

    yushui2022/MathModel-Skill

    Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.

    454 GitHub stars~1.4k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 13 days ago
    Auto-check passed

Questions about Sigmod Reproducibility

What does Sigmod Reproducibility do?

A skill your agent uses when hardening the reproducibility story of a SIGMOD submission, covering PACMMOD's expectation that code, data, scripts, and notebooks be shared, experiment provenance from…. Sigmod Reproducibility is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when hardening the reproducibility story of a SIGMOD submission, covering PACMMOD's expectation that code, data, scripts, and notebooks be shared, experiment provenance from config to figure, dataset and workload disclosure, variance reporting for systems numbers, and alignment with later ARI badging.

When should I use Sigmod Reproducibility?

Sigmod Reproducibility fits situations like: hardening the reproducibility story of a SIGMOD submission; covering PACMMODs expectation that code; notebooks be shared; experiment provenance from config to figure.

How do I install Sigmod Reproducibility in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill sigmod-reproducibility -a claude-code`. Or copy the skill folder (SIGMOD-Skills/skills/sigmod-reproducibility in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/sigmod-reproducibility in your project. Claude Code loads it when a task matches its description.

How do I install Sigmod Reproducibility in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill sigmod-reproducibility -a codex`. Or copy the skill folder (SIGMOD-Skills/skills/sigmod-reproducibility in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/sigmod-reproducibility in your project. Codex loads it when a task matches its description.

Can I use Sigmod Reproducibility in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill sigmod-reproducibility -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sigmod-reproducibility, .gemini/skills/sigmod-reproducibility, .github/skills/sigmod-reproducibility and .opencode/skills/sigmod-reproducibility in your project.

What does Sigmod Reproducibility need to run?

Going by SKILL.md and its folder, Sigmod Reproducibility needs the command-line tools its instructions call (git).

Does Sigmod Reproducibility access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sigmod Reproducibility safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sigmod Reproducibility use?

Sigmod Reproducibility is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sigmod Reproducibility use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sigmod Reproducibility?

Skills that share tags, products or a category with Sigmod Reproducibility: Peer Review (K-Dense-AI/claude-scientific-writer, 2.4k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Compute Environment Setup (aipoch/open-science, 5.5k stars) and Figure Style (aipoch/open-science, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sigmod Reproducibility?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.