A skill your agent uses when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions…

MITAuto-check passedResearch & Science

Install Osdi Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills osdi-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/OSDI-Skills/skills/osdi-experiments .claude/skills/osdi-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
osdi-experiments
GitHub stars
1.2k
Token cost
~1.5k tokens
SKILL.md length
682 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions…

  • Works in 4 steps: Does the idea work end to end? — the… → Where does the benefit come from? —… → What does it cost? — the overheads the… → …
  • Auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads
  • SKILL.md covers Research questions first, Baselines that fight back, Workload realism and Measurement discipline, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Osdi Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions, measuring scalability and tail behavior, quantifying the design's costs, and fitting the evidence into the 12-page reviewed body.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Hypothesis generation. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads
  • Structuring the section around research questions
  • Measuring scalability and tail behavior
  • Quantifying the designs costs

Example prompts

  • “/osdi-experiments”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Does the idea work end to end? — the headline comparison on a realistic
  2. Where does the benefit come from? — component breakdown attributing the win to
  3. What does it cost? — the overheads the design admits (memory, write
  4. When does it break? — scalability limits, adversarial workloads, failure and

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Osdi Experiments loads about 1.5k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 682 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 682 words, ~1,549 tokens.

Download SKILL.mdSave it as .claude/skills/osdi-experiments/SKILL.md (or your agent's skills folder).
name
osdi-experiments
description
Use when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions, measuring scalability and tail behavior, quantifying the design's costs, and fitting the evidence into the 12-page reviewed body.

OSDI Experiments

Design the evaluation as the paper's proof obligation. The page constraints referenced here are OSDI '26 rules (12 reviewed pages, no appendices at submission — verified 2026-07-08); the evidence standards are the durable expectations of systems PCs.

Research questions first

Write the evaluation's research questions before running anything, and derive the experiment set from them. Every OSDI evaluation ultimately answers versions of:

  1. Does the idea work end to end? — the headline comparison on a realistic workload against the strongest baseline.
  2. Where does the benefit come from? — component breakdown attributing the win to the named design idea rather than to incidental engineering.
  3. What does it cost? — the overheads the design admits (memory, write amplification, CPU, complexity), measured, not estimated.
  4. When does it break? — scalability limits, adversarial workloads, failure and recovery behavior.

An evaluation organized as RQ1–RQ4 with one experiment cluster each reads as an argument; a tour of every benchmark you happened to run reads as padding, which the OSDI '26 CFP explicitly invites reviewers to down-rank.

Baselines that fight back

The baseline question decides more OSDI reviews than any other. Standards:

  • Compare against the strongest deployed or published system for the problem, at its tuned best — not a strawman configuration. Reviewers who built the baseline will recognize a sandbagged setup instantly.
  • If the nearest competitor is unavailable, re-implement from its paper and say so, with the re-implementation validated against published numbers where possible.
  • Include the do-nothing baseline where it is honest: sometimes the existing system plus more hardware is the real alternative, and the paper is stronger for pricing it.
  • Version, configuration, and tuning of every baseline belong in the paper. Under the no-appendix rule this is body text — budget for it.

Workload realism

Workload tierRole in the argumentTrap
MicrobenchmarksIsolate a mechanism; explain why the end-to-end effect existsAs the only evidence: workshop-grade
Standard suites (e.g., YCSB-class)Comparability with prior papersDefaults nobody runs in production
Trace-driven / production-derivedThe claim's load-bearing evidenceProvenance undocumented (see osdi-reproducibility)
Adversarial / stressAnswers RQ4 honestlyOmitted, leaving reviewers to imagine worse

Systems reviewers read workload sections looking for the flattering-choice smell: the one skew setting, working-set size, or thread count where the design shines. Sweep the parameter, show the crossover point, and say where the baseline wins — a visible crossover is credibility, not weakness.

Show full SKILL.md (294 more words)Show less

Measurement discipline

  • Report distributions, not averages, wherever latency matters: medians and p99s diverge exactly where systems papers live. State run counts and variance.
  • Separate warm and cold behavior; state measurement windows and what was discarded.
  • Attribute wins: a profile or counter-level breakdown connecting the speedup to the mechanism ("the win disappears when we disable per-tenant ordering") beats any additional benchmark.
  • Failure experiments need injected faults with stated injection method and timing, not prose about what recovery "would" do.
text
Experiment matrix skeleton (freeze ~8 weeks before the December deadline):
RQ | workload (tier + provenance) | baselines (version, tuning) | metric
   | scale points | runs x seeds | expected figure/table | status
Freeze the matrix, then let deadline pressure cut rows, never redefine them —
redefinition under pressure is how flattering choices happen.

Fitting evidence into 12 pages

With no appendix at submission, the evaluation must be self-sufficient and compact:

  • One figure or table per research question, each with a caption stating its takeaway.
  • Cut experiments that do not serve an RQ, even if they took weeks — the accepted-paper 14-page budget (plus appendices) can resurrect them later (osdi-camera-ready).
  • Keep the setup paragraph brutal: hardware, topology, versions, workloads, runs — in one place, once (osdi-reproducibility owns the full ledger).

Reporting grid

Match each metric class to its honest presentation before making figures:

Metric classReport asNot as
ThroughputCurve vs offered load, to saturationSingle peak number
LatencyMedian + p99 (p999 if claimed), distribution across runsMean ± nothing
Recovery/failoverTimeline from fault injection, per scale point"Fast recovery" prose
Overhead (the design's cost)Same rigor as the win, same tableFootnote estimate
ScalabilityEfficiency vs ideal at each point"Near-linear" unquantified

One convention repays its cost: keep the baseline's color/marker identical across every figure, so the skim (osdi-review-process) reads the comparison correctly without consulting legends.

Review-time exposure

The evaluation objections you cannot rebut (no response period in 2026) are the predictable ones: weak baseline, unrealistic workload, missing cost measurement, and average-only latency. Audit for exactly these four before submission; each unaddressed one is a review point conceded silently.

Output format

text
[RQ coverage] RQ1-4 each mapped to experiments? gaps: <list>
[Baseline verdict] strongest opponent present + tuned? <one-line judgment>
[Workload realism] tiers present; flattering-choice risks: <list>
[Cost honesty] design's costs measured? <which, where>
[Tail discipline] distributions + variance reported? <yes/no + fix>
[Page fit] evaluation length vs 12-page budget; cut candidates: <list>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in OSDI-Skills/skills/osdi-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Osdi Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Osdi Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Osdi Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: NotesMIT
Nature Paper CardYuan1z0825/nature-skills47k2 repos~2.1kAutomated safety check: PassApache-2.0
Hypothesis GenerationK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: PassMIT
Good QuestionRimagination/good-question3051 repos~4.3kAutomated safety check: PassMIT
High Stakes Analytics Decision Lablimingrui679-design/high-stakes-analytics-decision-lab1k—~2.2kAutomated safety check: PassMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    47k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Hypothesis Generation

    K-Dense-AI/claude-scientific-writer

    Formulate evidence-bounded scientific questions, candidate hypotheses, rival explanations, causal or associational claims, discriminating predictions, measurements, and preregistration-ready…

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check passed
  • Good Question

    Rimagination/good-question

    A skill your agent uses when a researcher is choosing, framing, refining, or stress-testing a research question, hypothesis, thesis topic, project idea, grant direction, paper angle, or stalled…

    305 GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check passed
  • High Stakes Analytics Decision Lab

    limingrui679-design/high-stakes-analytics-decision-lab

    Build or review source-backed descriptive, diagnostic, predictive, and prescriptive analysis for consequential decisions.

    1k GitHub stars~2.2k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 6 repos~2.3k tokens
    Research & ScienceAuto-check: notes

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Osdi Experiments

What does Osdi Experiments do?

A skill your agent uses when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions…. Osdi Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions, measuring scalability and tail behavior, quantifying the design's costs, and fitting the evidence into the 12-page reviewed body.

When should I use Osdi Experiments?

Osdi Experiments fits situations like: auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads; structuring the section around research questions; measuring scalability and tail behavior; quantifying the designs costs.

How do I install Osdi Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-experiments -a claude-code`. Or copy the skill folder (OSDI-Skills/skills/osdi-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/osdi-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Osdi Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-experiments -a codex`. Or copy the skill folder (OSDI-Skills/skills/osdi-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/osdi-experiments in your project. Codex loads it when a task matches its description.

Can I use Osdi Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill osdi-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/osdi-experiments, .gemini/skills/osdi-experiments, .github/skills/osdi-experiments and .opencode/skills/osdi-experiments in your project.

What does Osdi Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Osdi Experiments is instructions for the agent only.

Does Osdi Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Osdi Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Osdi Experiments use?

Osdi Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Osdi Experiments use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Osdi Experiments?

Skills that share tags, products or a category with Osdi Experiments: Hypothesis Generation (spacering-net/codeg, 3.9k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars), Hypothesis Generation (K-Dense-AI/claude-scientific-writer, 2.4k stars) and Good Question (Rimagination/good-question, 305 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Osdi Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.