Agent skill

Kdd Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when packaging code, datasets, configs, and deployment evidence for a KDD paper, where the repository cited in the submission is the only artifact reviewers can reach because…

MITAuto-check passed

Install Kdd Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill kdd-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills kdd-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/KDD-Skills/skills/kdd-artifact-evaluation .claude/skills/kdd-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kdd-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.7k tokens
SKILL.md length
747 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when packaging code, datasets, configs, and deployment evidence for a KDD paper, where the repository cited in the submission is the only artifact reviewers can reach because…

  • Deployment evidence for a KDD paper
  • SKILL.md covers Artifact strategy by track, Building the anonymized…, Scale claims need scale… and ADS evidence packaging without…, plus 4 more sections
  • Calls docker and bash
  • Where the repository cited in the submission is the only artifact reviewers can reach because rebuttals ban links

What it does

Kdd Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging code, datasets, configs, and deployment evidence for a KDD paper, where the repository cited in the submission is the only artifact reviewers can reach because rebuttals ban links. Covers anonymized repo construction, scale-claim harnesses, ADS evidence without production data, and post-acceptance release.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Deployment evidence for a KDD paper
  • Where the repository cited in the submission is the only artifact reviewers can reach because rebuttals ban links

Example prompts

  • “/kdd-artifact-evaluation”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kdd Artifact Evaluation loads about 1.7k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 747 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 747 words, ~1,736 tokens.

Download SKILL.mdSave it as .claude/skills/kdd-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
kdd-artifact-evaluation
description
Use when packaging code, datasets, configs, and deployment evidence for a KDD paper, where the repository cited in the submission is the only artifact reviewers can reach because rebuttals ban links. Covers anonymized repo construction, scale-claim harnesses, ADS evidence without production data, and post-acceptance release.

KDD Artifact Evaluation

Use this while the submission is being assembled — not after. KDD's review mechanics create one hard constraint that reorders all artifact work: the rebuttal phase does not allow hyperlinks, so the anonymized repository referenced inside the submitted PDF is the complete and final artifact channel for the whole review. There is no "we'll share code if reviewers ask"; asking happens in a phase where you cannot answer with a link.

Artifact strategy by track

TrackPrimary artifactWhat reviewers actually probeNon-shippable core, and its substitute
ResearchAnonymized code + configs + data loadersCan the headline table be regenerated? Does the scale claim have a runnable path?Massive datasets → downsampled slice + full-scale download script
ADSMeasurement definitions + pipeline skeletonAre post-launch metrics precisely defined? Is the eval window stated?Production data/code → metric spec, schema, synthetic replay generator
Datasets & BenchmarksThe dataset itself + loaders + baseline harnessLicense, provenance, documentation, versioningNothing — the artifact is the paper

Building the anonymized repository

  • Create a fresh repository, never a scrubbed clone: git history, CI configs, issue templates, and .git metadata leak identity that no README edit removes.
  • Sweep for: usernames in paths, institutional cluster names, cloud bucket names, internal package registries, license headers, notebook execution metadata, and dataset names that only one company uses.
  • Cite the repository in the paper body (and optionally the abstract); after the deadline this reference is unchangeable, so double-check the URL resolves logged out.
  • Structure for a 10-minute reviewer, because that is the realistic inspection budget:
bash
anonymous-artifact/
├── README.md            # 1 screen: claim -> command -> expected output table
├── env/                 # lockfile or container spec, exact versions
├── configs/             # one config per reported table/figure row
├── data/
│   ├── get_data.sh      # public downloads, checksums
│   └── sample/          # small slice so the pipeline runs in minutes
├── run.sh               # regenerates the smallest headline result end-to-end
└── results/expected/    # committed reference outputs for diffing
# smoke-test on a clean machine:
docker run --rm -v $PWD:/w -w /w python:3.11 bash -c "pip install -r env/requirements.txt && bash run.sh --sample"

Scale claims need scale artifacts

KDD reviewers read "scales to billions of edges" as a checkable claim, not marketing:

  • Ship the throughput/memory harness, not only accuracy scripts; a scalability figure without its benchmark script is the least-trusted plot in the paper.
  • Record hardware, dataset cardinalities (rows/edges/events), and wall-clock per run in a machine-readable manifest, so the repo answers "on what?" precisely.
  • Where the full-scale run needs a cluster no reviewer has, provide the mid-scale run that completes on one machine plus the exact extrapolation logic the paper uses.

ADS evidence packaging without leaking production

The ADS track requires quantified post-launch performance, but production data almost never ships. Reviewers accept that trade when the package contains:

  • Exact metric definitions (numerator, denominator, unit, aggregation window) for every post-launch number in the paper.
  • The measurement design: A/B split or pre/post window, traffic share, duration, and what guardrail metrics were tracked.
  • A synthetic or replayed data generator that exercises the same pipeline shape, so the method is runnable even though the evidence is observational.
  • An honesty note distinguishing measured production numbers from simulated ones — mixing them silently is a rejection pattern.
Show full SKILL.md (311 more words)Show less

After acceptance

  • Replace the anonymous mirror with the public, licensed, citable repository; the camera-ready must carry the durable link (see kdd-camera-ready).
  • Prefer an archival deposit (DOI-stamped release) over a bare git URL for the version that the proceedings reference.
  • Check the current cycle for artifact-badging or code-availability initiatives before assuming none exists (待核实 — such programs change year to year).

Vignette: packaging a billion-edge graph paper

A Research Track paper claims its sampler trains GNNs on a 3B-edge graph on one machine. What the artifact must contain for that claim to survive contact with a skeptical reviewer:

  • get_data.sh that downloads the public 3B-edge graph (or constructs it from public parts) with checksums — a scale claim on an unfetchable graph is attested, not rerunnable, and should be labeled accordingly (kdd-reproducibility).
  • A --scale small|medium|full switch: small finishes on a laptop in minutes and validates the pipeline; medium reproduces one main-table row on a single GPU overnight; full documents the exact hardware used for the headline.
  • The peak-memory tracker wired into every run, because the paper's "one machine" claim is really a memory claim.
  • results/expected/ with per-scale reference outputs, so a reviewer's partial rerun has something to diff against.

What it must not contain: the 40GB of intermediate artifacts (regenerable), the authors' cluster submission scripts (identity leak), or a README promising "full instructions after acceptance" — that sentence tells reviewers the artifact is theater.

Common artifact failures at this venue

FailureWhy it is fatal at KDD specifically
Repo created but never cited in the PDFThe link ban makes it undiscoverable during rebuttal
Git history preserved from the lab repoIdentity leak → desk-level anonymity problem
Accuracy scripts only, no efficiency harnessThe paper's scale/efficiency axis becomes unverifiable
Sample data missing, full data gatedReviewer's 10-minute budget ends at the download wall
ADS package with raw production extractsConfidentiality violation risk transferred to reviewers

Output format

text
[Artifact channel] repo cited in PDF: yes/no (if no: unrecoverable after deadline)
[Track register] research-repro / ads-deployment-evidence / dataset-release
[Regeneration level] one-command sample / scripted / descriptive only
[Scale evidence] throughput+memory harness: present / missing
[Anonymity sweep] <paths/history/metadata findings>
[Post-acceptance plan] <public repo, license, archival DOI>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in KDD-Skills/skills/kdd-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Kdd Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kdd Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kdd Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
DatasetsArize-ai/phoenix12k—~1.6kAutomated safety check: PassCustom licence
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
Ccs Artifact Evaluationbrycewang-stanford/Awesome-Journal-Skills1.2k—~969Automated safety check: PassMIT
Micro Artifact Evaluationbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
Mobisys Artifact Evaluationbrycewang-stanford/Awesome-Journal-Skills1.2k—~1kAutomated safety check: PassMIT

Similar skills

  • Datasets

    Arize-ai/phoenix

    Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.

    12k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ccs Artifact Evaluation

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when packaging ACM CCS artifacts for the artifact-evaluation committee and the ACM badges — Artifacts Available, Artifacts Evaluated Functional, Artifacts Evaluated Reusable…

    1.2k GitHub stars~969 tokensUpdated 13 days ago
    Auto-check passed
  • Micro Artifact Evaluation

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when preparing a MICRO artifact for post-acceptance evaluation — packaging simulators, configs, traces, and scripts so evaluators can regenerate the paper's figures…

    1.2k GitHub stars~1.7k tokensUpdated 13 days ago
    Auto-check passed
  • Mobisys Artifact Evaluation

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when packaging a MobiSys artifact for the Artifact Evaluation Committee — choosing among the three independent ACM badges (Available, Evaluated–Functional, Results…

    1.2k GitHub stars~1k tokensUpdated 13 days ago
    Auto-check passed
  • Fast Artifact Evaluation

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when packaging a USENIX FAST artifact for the USENIX Artifact Evaluation scheme (Artifacts Available, Artifacts Functional, Results Reproduced), covering what a storage AEC…

    1.2k GitHub stars~1.6k tokensUpdated 13 days ago
    Auto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 13 days ago
    Auto-check passed

Questions about Kdd Artifact Evaluation

What does Kdd Artifact Evaluation do?

A skill your agent uses when packaging code, datasets, configs, and deployment evidence for a KDD paper, where the repository cited in the submission is the only artifact reviewers can reach because…. Kdd Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging code, datasets, configs, and deployment evidence for a KDD paper, where the repository cited in the submission is the only artifact reviewers can reach because rebuttals ban links.

When should I use Kdd Artifact Evaluation?

Kdd Artifact Evaluation fits situations like: deployment evidence for a KDD paper; where the repository cited in the submission is the only artifact reviewers can reach because rebuttals ban links.

How do I install Kdd Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill kdd-artifact-evaluation -a claude-code`. Or copy the skill folder (KDD-Skills/skills/kdd-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/kdd-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Kdd Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill kdd-artifact-evaluation -a codex`. Or copy the skill folder (KDD-Skills/skills/kdd-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/kdd-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Kdd Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill kdd-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kdd-artifact-evaluation, .gemini/skills/kdd-artifact-evaluation, .github/skills/kdd-artifact-evaluation and .opencode/skills/kdd-artifact-evaluation in your project.

What does Kdd Artifact Evaluation need to run?

Going by SKILL.md and its folder, Kdd Artifact Evaluation needs the command-line tools its instructions call (docker and bash). Our summary lists: Python 3; Docker.

Does Kdd Artifact Evaluation access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Kdd Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kdd Artifact Evaluation use?

Kdd Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kdd Artifact Evaluation use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kdd Artifact Evaluation?

Skills that share tags, products or a category with Kdd Artifact Evaluation: Datasets (Arize-ai/phoenix, 12k stars), Arize Evaluator (github/awesome-copilot, 40k stars), Ccs Artifact Evaluation (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars) and Micro Artifact Evaluation (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kdd Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.