Agent skill

Webconf Reproducibility

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls, platform APIs, live systems, or user logs — covering dataset decay…

MITAuto-check passedResearch & Science

Install Webconf Reproducibility

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill webconf-reproducibility -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills webconf-reproducibility --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/The-Web-Conference-Skills/skills/webconf-reproducibility .claude/skills/webconf-reproducibility && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
webconf-reproducibility
GitHub stars
1.2k
Token cost
~1.7k tokens
SKILL.md length
767 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls, platform APIs, live systems, or user logs — covering dataset decay…

  • Works in 3 steps: Main 8 pages: one reproducibility… → Appendix (within the 12): hyperparameter… → Artifact (see…
  • Hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls
  • SKILL.md covers Three reproducibility regimes, Pinning the decaying regime, Determinism audit for the… and Temporal honesty, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Webconf Reproducibility is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls, platform APIs, live systems, or user logs — covering dataset decay, temporal snapshots, seed and environment reporting, the reproducibility appendix inside the 12-page PDF, and honest claims when the Web itself cannot be replayed.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Reproducible research. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls
  • User logs — covering dataset decay
  • Temporal snapshots
  • Seed and environment reporting

Example prompts

  • “/webconf-reproducibility”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Main 8 pages: one reproducibility paragraph — regime classification,
  2. Appendix (within the 12): hyperparameter tables, environment details,
  3. Artifact (see webconf-artifact-evaluation): everything executable, the

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Webconf Reproducibility loads about 1.7k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 767 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 767 words, ~1,734 tokens.

Download SKILL.mdSave it as .claude/skills/webconf-reproducibility/SKILL.md (or your agent's skills folder).
name
webconf-reproducibility
description
Use when hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls, platform APIs, live systems, or user logs — covering dataset decay, temporal snapshots, seed and environment reporting, the reproducibility appendix inside the 12-page PDF, and honest claims when the Web itself cannot be replayed.

Web Conference Reproducibility

Reproducibility at this venue has a problem no offline-ML venue has: the object of study mutates. Pages die, APIs close, ranking systems retrain, platform policies change what may be collected at all. A Web Conference paper is reproducible to the degree that it pins what can be pinned and measures what cannot. The 2026 CFP's sanctioned home for this material is the optional appendix — "details on reproducibility, proofs, pseudo-code" — inside the same 12-page PDF, which reviewers are not obliged to read; so the reproducibility claims go in the main 8 pages and the reproducibility mechanics go in the appendix.

Three reproducibility regimes

RegimeExample evidenceWhat "reproducible" meansYour obligation
FrozenPublic benchmark, released crawlRe-run → same numbersSeeds, versions, exact splits
DecayingYour own crawl, API pullsRe-collect → quantifiably similar corpusSnapshot, checksums, collection code, date stamps
UnreplayableLive A/B test, production traffic, human subjectsIndependent teams can audit the protocolFull protocol, power analysis, aggregate release

Most reviews go wrong when a paper claims regime-1 language ("fully reproducible") for regime-2 or regime-3 evidence. Classify every experiment in the paper into a regime and phrase its claim accordingly; the honest sentence "results on the live platform are audit-reproducible but not replay-reproducible" has never sunk a strong paper.

Pinning the decaying regime

  • Date-stamp everything: crawl window, API version, model checkpoints of any third-party service used (an LLM API call in 2025 is not the same function in 2026 — record model name and version string).
  • Checksum the corpus at collection time; publish per-file SHA-256 in a manifest so drift is detectable, not just suspected.
  • Prefer citable snapshots: Common Crawl snapshot IDs, Internet Archive captures, Wikipedia dumps with dates — these convert a decaying corpus into a frozen one for everyone downstream.
  • Separate collection from analysis in the codebase, so a future team can re-run analysis on your frozen snapshot even when re-collection is impossible.

Determinism audit for the modeling stack

python
# Repro header every experiment script in the artifact should share
import os, random, numpy as np, torch

SEED = int(os.environ.get("RUN_SEED", 17))
random.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)
torch.use_deterministic_algorithms(True)          # surfaces nondeterministic ops
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8" # required by some CUDA GEMMs

# Log the things people forget to log:
#   graph/dataloader shuffling seeds, negative-sampling seeds,
#   train/val/test split hash, library versions, GPU model, wall-clock.

Web-specific nondeterminism deserves explicit lines in the appendix: crawl ordering, deduplication thresholds, timezone normalization of timestamps, and — for graph papers — node ID remapping, which silently reorders neighbor sampling.

Temporal honesty

Web data is time-indexed, and the venue's reviewers increasingly check for temporal leakage: random splits over user-item interactions or evolving graphs let the model train on the future. The reproducibility appendix should state the split rule (e.g., "train < 2025-06-01 ≤ test"), not just percentages, and the artifact should ship the split-generation code rather than opaque index files alone. If the paper uses a random split on temporal data for comparability with prior work, say so and add one temporal split as a robustness check — this one-sentence-plus-one-table addition preempts the most common modern objection.

Show full SKILL.md (324 more words)Show less

Vignette: the benchmark that dissolved

A 2024-vintage misinformation dataset distributes tweet IDs for rehydration. By the time a team builds on it for a WWW submission, 38% of the tweets are deleted, suspended, or geo-blocked — and deletion is not random: the most-reported content vanishes first. Naively rehydrating and comparing against the original paper's numbers silently changes both the task and the class balance. The regime-honest handling, which fits in four appendix sentences plus one table column: report the rehydration date and survival rate, compare label distributions between the original and surviving corpus, rerun the strongest baseline on the surviving subset so all comparisons share one corpus, and phrase cross-paper comparisons as indicative rather than head-to-head. Reviewers do not penalize decay — it is the field's shared condition — but they increasingly penalize pretending it did not happen.

What goes where

  1. Main 8 pages: one reproducibility paragraph — regime classification, snapshot/DOI pointers, seed policy, split rule. Claims a reviewer must weigh cannot hide past page 8.
  2. Appendix (within the 12): hyperparameter tables, environment details, collection protocol, per-dataset manifests, negative results of tuning.
  3. Artifact (see webconf-artifact-evaluation): everything executable, the manifest with checksums, and the recrawl/dead-link accounting script.

A placement corollary for review strategy: because the appendix is optional reading, a reviewer who doubts reproducibility may score the doubt without opening Appendix B. The main-text paragraph therefore needs one forward pointer with content — "seeds, environment, and the full collection protocol are in App. B; the artifact reproduces Table 2 with one script" — so the doubt has an address before it becomes a score.

Pre-submission reproducibility gate

  • Every experiment classified frozen / decaying / unreplayable, phrased to match.
  • Collection dates, API versions, and third-party model versions stated.
  • Split rules explicit; temporal data has a temporal split somewhere.
  • Seeds and repetition counts reported; variance shown where runs repeat.
  • A stranger with the appendix plus artifact could rebuild the headline table or, failing that, audit the protocol end to end.

Output format

text
[Regimes] frozen=<experiments> decaying=<...> unreplayable=<...>
[Pinning] dates/checksums/snapshots: complete / gaps <where>
[Temporal] split rule stated? leakage risk? robustness split present?
[Determinism] seed policy + environment logged: yes/no
[Placement] claims in main text, mechanics in appendix: verified
[Honesty edits] <sentences whose reproducibility claim overshoots the regime>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in The-Web-Conference-Skills/skills/webconf-reproducibility of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Webconf Reproducibility next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Webconf Reproducibility compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Webconf Reproducibility this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
Peer ReviewK-Dense-AI/claude-scientific-writer2.4k2 repos~3.1kAutomated safety check: NotesMIT
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Compute Environment Setupaipoch/open-science5.5k—~2.6kAutomated safety check: PassApache-2.0
Figure Styleaipoch/open-science5.5k—~5.1kAutomated safety check: PassApache-2.0
Add Bactopia Toolbactopia/bactopia522—~4.1kAutomated safety check: PassMIT

Similar skills

  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Compute Environment Setup

    aipoch/open-science

    Prepares setup instructions and a named activation file for a user-managed software environment on an Open-Science SSH or Slurm compute host.

    5.5k GitHub stars~2.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Figure Style

    aipoch/open-science

    Publication-grade correctness and legibility rules for final-deliverable scientific figures, not exploratory plots.

    5.5k GitHub stars~5.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Add Bactopia Tool

    bactopia/bactopia

    Scaffold a complete Bactopia Tool across all three tiers -- module, subworkflow, and workflow entry point under workflows/bactopia-tools/.

    522 GitHub stars~4.1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Modeling Code and Result Contracts

    yushui2022/MathModel-Skill

    Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.

    453 GitHub stars~1.4k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Webconf Reproducibility

What does Webconf Reproducibility do?

A skill your agent uses when hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls, platform APIs, live systems, or user logs — covering dataset decay…. Webconf Reproducibility is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls, platform APIs, live systems, or user logs — covering dataset decay, temporal snapshots, seed and environment reporting, the reproducibility appendix inside the 12-page PDF, and honest claims when the Web itself cannot be replayed.

When should I use Webconf Reproducibility?

Webconf Reproducibility fits situations like: hardening the reproducibility of a Web Conference (WWW) paper whose evidence rests on crawls; user logs — covering dataset decay; temporal snapshots; seed and environment reporting.

How do I install Webconf Reproducibility in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill webconf-reproducibility -a claude-code`. Or copy the skill folder (The-Web-Conference-Skills/skills/webconf-reproducibility in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/webconf-reproducibility in your project. Claude Code loads it when a task matches its description.

How do I install Webconf Reproducibility in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill webconf-reproducibility -a codex`. Or copy the skill folder (The-Web-Conference-Skills/skills/webconf-reproducibility in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/webconf-reproducibility in your project. Codex loads it when a task matches its description.

Can I use Webconf Reproducibility in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill webconf-reproducibility -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/webconf-reproducibility, .gemini/skills/webconf-reproducibility, .github/skills/webconf-reproducibility and .opencode/skills/webconf-reproducibility in your project.

What does Webconf Reproducibility need to run?

SKILL.md names no scripts, command-line tools or credentials: Webconf Reproducibility is instructions for the agent only. Our summary lists: Python 3.

Does Webconf Reproducibility access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Webconf Reproducibility safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Webconf Reproducibility use?

Webconf Reproducibility is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Webconf Reproducibility use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Webconf Reproducibility?

Skills that share tags, products or a category with Webconf Reproducibility: Peer Review (K-Dense-AI/claude-scientific-writer, 2.4k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Compute Environment Setup (aipoch/open-science, 5.5k stars) and Figure Style (aipoch/open-science, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Webconf Reproducibility?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.