Agent skill

Osdi Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline, the 2026 narrowing to a single Artifacts Available badge, Zenodo-grade…

MITAuto-check passedDevOps & Cloud

Install Osdi Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills osdi-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/OSDI-Skills/skills/osdi-artifact-evaluation .claude/skills/osdi-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
osdi-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.7k tokens
SKILL.md length
742 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline, the 2026 narrowing to a single Artifacts Available badge, Zenodo-grade…

  • Works in 2 steps: Meet the letter: a permanent, resolvable… → Build to the historical bar anyway.…
  • Preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline
  • SKILL.md covers How OSDI AE is wired, The 2026 badge narrowing, Artifact anatomy and Systems-specific honesty, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Osdi Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline, the 2026 narrowing to a single Artifacts Available badge, Zenodo-grade permanent archiving, the AE-committee runbook, and the two-page Artifact Appendix that documents the result.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Runbooks and postmortems. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline
  • The 2026 narrowing to a single Artifacts Available badge
  • Zenodo-grade permanent archiving
  • The AE-committee runbook

Example prompts

  • “/osdi-artifact-evaluation”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Meet the letter: a permanent, resolvable archive — not a personal homepage, not
  2. Build to the historical bar anyway. Functional/Reproduced may return in later

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Osdi Artifact Evaluation loads about 1.7k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 742 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 742 words, ~1,720 tokens.

Download SKILL.mdSave it as .claude/skills/osdi-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
osdi-artifact-evaluation
description
Use when preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline, the 2026 narrowing to a single Artifacts Available badge, Zenodo-grade permanent archiving, the AE-committee runbook, and the two-page Artifact Appendix that documents the result.

OSDI Artifact Evaluation

Package the system for the artifact evaluation committee (AEC). Everything cycle-specific below is OSDI '26 as verified 2026-07-08 — and 2026 changed the rules, so check the live Call for Artifacts (usenix.org/conference/osdi26/call-for-artifacts pattern) before optimizing for the wrong target.

How OSDI AE is wired

  • Who: an AEC organized with the sysartifacts community (sysartifacts.github.io), which runs artifact evaluation across the systems conferences — reviewers are systems students and researchers on machines that are not yours.
  • When: artifacts are submitted after the paper is (conditionally) accepted — the '26 deadline was May 8, 2026, 8:59 pm PDT, about six weeks after the March 26 notification. The chairs explicitly encouraged preparing the artifact while the paper was still under review; teams that ignored this spent April in a panic.
  • Stakes: badges appear on the published paper. AE is optional but has become an expectation for systems papers claiming practical relevance — and it is the prerequisite for the final paper's two-page Artifact Appendix.

The 2026 badge narrowing

CycleBadges evaluated
OSDI '21–'25 (sysartifacts calls)Artifacts Available, Artifacts Functional, Results Reproduced
OSDI '26Artifacts Available only

This is the pack's sharpest example of cycle volatility. In 2026 the AEC judged one thing: that the artifacts have been made available for retrieval, permanently and publicly. Zenodo was the encouraged host; institutional repositories and third-party archives (FigShare, Dryad, Software Heritage, GitHub, GitLab) were acceptable.

Two disciplined responses:

  1. Meet the letter: a permanent, resolvable archive — not a personal homepage, not a repo you might rename. Mint the DOI/permanent identifier before the deadline and put that identifier in the paper.
  2. Build to the historical bar anyway. Functional/Reproduced may return in later cycles, your readers will attempt reproduction regardless (the proceedings are open access), and a runnable artifact is what makes the badge worth having. The structure below targets reproducibility even when only availability is graded.

Artifact anatomy

text
osdi26-artifact/                    # archived as one versioned unit (e.g. Zenodo)
├── README.md                       # THE document: claims map + kick-the-tires + full runs
├── LICENSE                         # explicit; AEC members must be allowed to run it
├── system/                         # source at the paper's commit, build instructions
├── baselines/                      # versions/tags, tuning configs, rebuild notes
├── workloads/                      # traces or regeneration scripts + provenance/licensing
├── experiments/                    # one driver per paper claim: run_fig7.sh, run_tab3.sh
├── expected/                       # reference outputs + tolerance notes (variance!)
├── ledger/                         # provenance records from osdi-reproducibility
└── environment/                    # container/VM image or exact setup script; hw requirements

The README's claims map is the heart: a table from paper claim (Fig. 7, Table 3, "recovery under 900 ms") to driver script, expected output, tolerance, and runtime. An evaluator should reach a first signal — the "kick-the-tires" path — in under 30 minutes on commodity hardware, even if headline experiments need a testbed.

Systems-specific honesty

  • Hardware walls: if a claim needs 64 nodes or specific NICs, say so at the top and provide a scaled-down mode that exercises every code path with reduced magnitudes; state which numbers the small mode can and cannot reproduce.
  • Variance: publish tolerance windows per claim ("p99 within ±10% across 10 runs"), sourced from your own ledger — an evaluator who gets 3.1x where the paper says 3.2x should find reassurance, not confusion.
  • Licensing traps: production-derived traces often cannot be redistributed; include the generator plus provenance documentation, and say so plainly rather than shipping a mystery synthetic substitute.
  • De-anonymization timing: AE runs post-acceptance, so the artifact need not be anonymous in 2026's flow — but verify this against the current call rather than assuming it, since submission-time AE exists at other systems venues.
Show full SKILL.md (248 more words)Show less

Working with the AEC

Artifact evaluation is a cooperative review, and the committee's constraints shape what a good artifact looks like:

  • Evaluators are typically graduate students and postdocs handling several artifacts in a fixed window — assume limited time per artifact and front-load the payoff: the fastest path from download to a verifiable signal wins goodwill that carries through the harder experiments.
  • Expect a communication channel (the AE submission system) for questions during evaluation; answer fast and patch the README rather than replying inline, so the archived artifact ends evaluation better than it started. Where the archive is immutable (a minted DOI version), publish a new version rather than editing history.
  • Provide fallbacks per failure point: a prebuilt container beside the build-from-source path, a small dataset beside the trace generator, expected logs beside every driver so a partial failure is diagnosable from output alone.
  • If the badge set is availability-only (as in 2026), evaluators still need to retrieve and inspect the archive — a 40 GB tarball with no manifest can fail even that bar in practice. Ship a top-level manifest with sizes and checksums.

From badge to Artifact Appendix

Papers passing evaluation may add up to two pages of Artifact Appendix to the final paper (June 9 deadline, osdi-camera-ready): a standard-format description of scope, requirements, and reproduction steps. Write it by condensing the README the AEC just validated — it is the only reviewed-adjacent supplement OSDI has, and it outlives the conference as the artifact's citable front door.

Output format

text
[Cycle badge set] confirmed from live Call for Artifacts: <list> (待核实 if unread)
[Archive] permanent host + identifier minted? <where / missing>
[Claims map] paper claims covered by drivers: <n/m>; kick-the-tires < 30 min? yes/no
[Hardware honesty] testbed requirements + scaled-down mode documented? yes/no
[Timeline] days to artifact deadline; README/expected-output gaps: <list>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in OSDI-Skills/skills/osdi-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Osdi Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Osdi Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Osdi Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
Trader Memory Coretradermonty/claude-trading-skills3k2 repos~4.3kAutomated safety check: PassMIT
Author Migrationnrwl/nx29k—~12kAutomated safety check: NotesMIT
Write Notes Like Deepseekczm15053/write-notes-like-deepseek508—~2kAutomated safety check: PassNone
OpenRig Upgrade Proceduremvschwarz/openrig6.8k—~2.9kAutomated safety check: PassApache-2.0
GreptimeDB Release RunbookGreptimeTeam/greptimedb6.7k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Trader Memory Core

    tradermonty/claude-trading-skills

    Track investment theses across their lifecycle — from screening idea to closed position with postmortem.

    3k GitHub starsUsed in 2 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Author or scope a first-party Nx migration. An agent skill from nrwl/nx.

    29k GitHub stars~12k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Write Notes Like Deepseek

    czm15053/write-notes-like-deepseek

    A skill your agent uses when a change is non-trivial by DSH standards (behavior, architecture, cross-file contracts, process/tooling, testing strategy, or on-disk/wire/config formats), when choosing…

    508 GitHub stars~2k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • OpenRig Upgrade Procedure

    mvschwarz/openrig

    Walks an agent through upgrading the OpenRig CLI and daemon one observed step at a time, keeping live seats alive and reconciling managed plugin files.

    6.8k GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • GreptimeDB Release Runbook

    GreptimeTeam/greptimedb

    Runbook for publishing a GreptimeDB version: pick the release branch, verify the Cargo version, then tag, create the GitHub release and open the docs note PR.

    6.7k GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Statem

    henryqin1997/statem

    A skill your agent uses when a long coding or research task should be managed with statem state-machine runbooks, including creating specs, starting or resuming runs, checking current state…

    1.3k GitHub stars~1.2k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 14 days ago
    Auto-check passed

Categories

Questions about Osdi Artifact Evaluation

What does Osdi Artifact Evaluation do?

A skill your agent uses when preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline, the 2026 narrowing to a single Artifacts Available badge, Zenodo-grade…. Osdi Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline, the 2026 narrowing to a single Artifacts Available badge, Zenodo-grade permanent archiving, the AE-committee runbook, and the two-page Artifact Appendix that documents the result.

When should I use Osdi Artifact Evaluation?

Osdi Artifact Evaluation fits situations like: preparing an OSDI artifact for sysartifacts-run evaluation — the post-acceptance timeline; the 2026 narrowing to a single Artifacts Available badge; zenodo-grade permanent archiving; the AE-committee runbook.

How do I install Osdi Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-artifact-evaluation -a claude-code`. Or copy the skill folder (OSDI-Skills/skills/osdi-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/osdi-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Osdi Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-artifact-evaluation -a codex`. Or copy the skill folder (OSDI-Skills/skills/osdi-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/osdi-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Osdi Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/osdi-artifact-evaluation, .gemini/skills/osdi-artifact-evaluation, .github/skills/osdi-artifact-evaluation and .opencode/skills/osdi-artifact-evaluation in your project.

What does Osdi Artifact Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Osdi Artifact Evaluation is instructions for the agent only.

Does Osdi Artifact Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Osdi Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Osdi Artifact Evaluation use?

Osdi Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Osdi Artifact Evaluation use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Osdi Artifact Evaluation?

Skills that share tags, products or a category with Osdi Artifact Evaluation: Trader Memory Core (tradermonty/claude-trading-skills, 3k stars), Author Migration (nrwl/nx, 29k stars), Write Notes Like Deepseek (czm15053/write-notes-like-deepseek, 508 stars) and OpenRig Upgrade Procedure (mvschwarz/openrig, 6.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Osdi Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.