Agent skill

Wsdm Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when packaging code, data, and models as evidence for a WSDM paper - anonymous repositories cited in the PDF, the proprietary-log dilemma of web-scale research…

MITAuto-check passedDocuments & Office

Install Wsdm Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill wsdm-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills wsdm-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/WSDM-Skills/skills/wsdm-artifact-evaluation .claude/skills/wsdm-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wsdm-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.7k tokens
SKILL.md length
746 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when packaging code, data, and models as evidence for a WSDM paper - anonymous repositories cited in the PDF, the proprietary-log dilemma of web-scale research…

  • Models as evidence for a WSDM paper - anonymous repositories cited in the PDF
  • SKILL.md covers The reviewer-facing artifact, The proprietary-log dilemma, Public substitutes worth knowing and Model and prompt artifacts, plus 3 more sections
  • Calls make
  • The proprietary-log dilemma of web-scale research

What it does

Wsdm Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging code, data, and models as evidence for a WSDM paper - anonymous repositories cited in the PDF, the proprietary-log dilemma of web-scale research, public-benchmark substitution tiers, WSDM Cup datasets, and what credible artifact release looks like at a venue without a formal badge process.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Models as evidence for a WSDM paper - anonymous repositories cited in the PDF
  • The proprietary-log dilemma of web-scale research
  • Public-benchmark substitution tiers
  • WSDM Cup datasets

Example prompts

  • “/wsdm-artifact-evaluation”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wsdm Artifact Evaluation loads about 1.7k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 746 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 746 words, ~1,667 tokens.

Download SKILL.mdSave it as .claude/skills/wsdm-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
wsdm-artifact-evaluation
description
Use when packaging code, data, and models as evidence for a WSDM paper - anonymous repositories cited in the PDF, the proprietary-log dilemma of web-scale research, public-benchmark substitution tiers, WSDM Cup datasets, and what credible artifact release looks like at a venue without a formal badge process.

WSDM Artifact Evaluation

Package artifacts for a venue where they are persuasion, not process. The pack found no formal artifact-evaluation track or badge system for current WSDM editions (待核实 each cycle) - the CFP-level expectation is the community norm of "practical yet principled": reviewers reward submissions whose claims a skeptic could re-derive. That means the artifact's job is to be inspectable during review and usable after publication, with no committee to certify it.

The reviewer-facing artifact

Because appendices count against WSDM's page budget and there is no rebuttal in which to hand over materials later, the anonymous repository referenced in the PDF is the only expansion space you get. Build it to be skimmed in ten minutes:

text
anon-artifact/
├── README.md            # 1 screen: claim -> script -> expected output table
├── LICENSE              # anonymized placeholder license during review
├── env/                 # lockfile or container spec, exact versions
├── data/
│   ├── public/          # download scripts for public benchmarks
│   └── PROPRIETARY.md   # honest statement of what cannot be shared and why
├── src/                 # training/ranking/mining code, no company paths
├── configs/             # one config per reported table row
└── results/
    └── seeds_1-5/       # raw metric dumps behind every table in the paper

Anonymity requirements are stricter than habit: fresh remote with no commit history, no author handles in configs or notebook metadata, no internal package registries, no dataset paths like /nfs/companyname/.... An Associate Chair can see who you are; your reviewers must not.

The proprietary-log dilemma

WSDM's core subject matter - query logs, click data, user graphs, ad interactions - is often legally unshareable. The venue's reviewers know this; what they punish is unverifiable work, not industrial work. Choose a rung and state it explicitly in the paper:

RungWhat is releasedWhat the paper must then carry
Full releaseData + code + configsJust the pointers
Sampled/anonymized releaseA privacy-scrubbed sample + full codeThe scrubbing procedure and how sample results track full-data results
Public-benchmark mirrorCode + experiments re-run on public datasetsBoth result sets, with the deltas discussed, not hidden
Code-onlyPipeline code, no dataEnough dataset statistics that a platform-holder could replicate
Nothing sharable-A candid limitation; expect a proportional credibility discount

The public-benchmark mirror rung is the WSDM workhorse: pair the proprietary result with the same method on public data so at least one full row of the evidence is end-to-end reproducible by anyone.

Public substitutes worth knowing

  • Standard IR/rec benchmarks with public tooling (MS MARCO, TREC collections, MovieLens/Amazon review dumps, KuaiRec-style log releases) - verify current licensing before citing a release plan.
  • WSDM Cup datasets: the conference's own annual challenge (the 2027 site already carries a WSDM Cup call) periodically produces real industrial datasets; using one signals community alignment. Check the specific cup's redistribution terms.
  • Graph/social benchmarks (SNAP collections and successors) for the network side of the venue.

Anything here can go stale or change license - re-verify at packaging time.

Model and prompt artifacts

Foundation-model-era WSDM papers ("Search with Foundation Models" is in-scope per the 2026 CFP) add artifact surfaces older guidance misses:

  • Pin model identity precisely: provider, model string, version/date, and decoding parameters; an unpinned API model makes results unreproducible by construction - say what you did about that (snapshotting outputs, local checkpoints, fixed eval windows).
  • Ship prompts verbatim in the repository even when the page budget forces the paper to summarize them.
  • Cache and release model outputs used in evaluation when terms permit; that lets others audit the judgment layer without paying for inference.
Show full SKILL.md (257 more words)Show less

Credibility signals at a badge-free venue

With no committee stamping artifacts, reviewers use fast proxies to decide whether the repository is real or decorative. Engineer the proxies:

SignalReads asCost to provide
Config file per reported table rowThe numbers came from these runsMinutes
Raw seed-level metric dumpsVariance claims are checkableMinutes
A make reproduce-table2 entry pointSomeone actually reran thisAn hour
Download script for each public datasetThe pipeline is end-to-endAn hour
Honest PROPRIETARY.mdThe authors know what they cannot proveMinutes
Commit dated one hour before deadline as the only commitDecorative repo- avoid

The last row is about the squeeze: a repository assembled on deadline night tends to mismatch the paper's numbers, and one mismatch found by a reviewer discounts the entire artifact. Freeze the repo when experiments freeze (wsdm-workflow sets this at one week out), not when the PDF does.

Post-acceptance conversion

At camera-ready time the anonymous mirror becomes the citable artifact: move to the real organization/account, add the actual license, tag the release that matches the camera-ready numbers, and archive a snapshot (e.g., a DOI-issuing archive) so the URL in the ACM DL version outlives your CI. Update the README's claim-to-script table against final camera-ready table numbers - drift between repo and proceedings is the most common post-publication complaint.

One more conversion detail: if the paper used the public-benchmark-mirror rung, keep both pipelines in the public repo permanently - the mirror is what future papers will actually build on, and its issues page becomes your citation engine.

Output format

text
[Artifact tier] full / sampled / public-mirror / code-only / none + stated in paper? 
[Ten-minute test] README claim->script->output table present: yes / no
[Anonymity sweep] history / handles / paths / registries: clean or leaks listed
[FM pins] model string+date+decoding params recorded: yes / no / n-a
[Post-acceptance plan] public home, license, tagged release, archival snapshot

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in WSDM-Skills/skills/wsdm-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Wsdm Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wsdm Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wsdm Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
Markdown Article FormatterJimLiu/baoyu-skills27k6 repos~3.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Obsidian MarkdownAtmosphere/atmosphere3.8k20 repos~1.3kAutomated safety check: PassApache-2.0
DOCXrvdbreemen/OTGW-firmware20733 repos~4.3kAutomated safety check: PassProprietary
Word Document Reader and WriterHKUDS/DeepTutor41k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    27k GitHub starsUsed in 6 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Obsidian Markdown

    Atmosphere/atmosphere

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    3.8k GitHub starsUsed in 20 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • DOCX

    rvdbreemen/OTGW-firmware

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).

    207 GitHub starsUsed in 33 repos~4.3k tokens
    Documents & OfficeAuto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • Crossposting

    wasp-lang/wasp

    Crosspost Wasp blog articles (MDX) to DEV.to and Medium. An agent skill from wasp-lang/wasp.

    19k GitHub stars~1.1k tokensUpdated today
    Documents & OfficeAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 14 days ago
    Auto-check passed

Questions about Wsdm Artifact Evaluation

What does Wsdm Artifact Evaluation do?

A skill your agent uses when packaging code, data, and models as evidence for a WSDM paper - anonymous repositories cited in the PDF, the proprietary-log dilemma of web-scale research…. Wsdm Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging code, data, and models as evidence for a WSDM paper - anonymous repositories cited in the PDF, the proprietary-log dilemma of web-scale research, public-benchmark substitution tiers, WSDM Cup datasets, and what credible artifact release looks like at a venue without a formal badge process.

When should I use Wsdm Artifact Evaluation?

Wsdm Artifact Evaluation fits situations like: models as evidence for a WSDM paper - anonymous repositories cited in the PDF; the proprietary-log dilemma of web-scale research; public-benchmark substitution tiers; WSDM Cup datasets.

How do I install Wsdm Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill wsdm-artifact-evaluation -a claude-code`. Or copy the skill folder (WSDM-Skills/skills/wsdm-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/wsdm-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Wsdm Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill wsdm-artifact-evaluation -a codex`. Or copy the skill folder (WSDM-Skills/skills/wsdm-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/wsdm-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Wsdm Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill wsdm-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wsdm-artifact-evaluation, .gemini/skills/wsdm-artifact-evaluation, .github/skills/wsdm-artifact-evaluation and .opencode/skills/wsdm-artifact-evaluation in your project.

What does Wsdm Artifact Evaluation need to run?

Going by SKILL.md and its folder, Wsdm Artifact Evaluation needs the command-line tools its instructions call (make).

Does Wsdm Artifact Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Wsdm Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wsdm Artifact Evaluation use?

Wsdm Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wsdm Artifact Evaluation use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wsdm Artifact Evaluation?

Skills that share tags, products or a category with Wsdm Artifact Evaluation: Markdown Article Formatter (JimLiu/baoyu-skills, 27k stars), Markitdown (ImCa0/just-laws, 781 stars), Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars) and DOCX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wsdm Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.