A skill your agent uses when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking…

MITAuto-check passedDevOps & Cloud

Install Nsdi Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill nsdi-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills nsdi-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/NSDI-Skills/skills/nsdi-experiments .claude/skills/nsdi-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nsdi-experiments
GitHub stars
1.2k
Token cost
~1.6k tokens
SKILL.md length
718 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking…

  • Works in 3 steps: Synthetic microbenchmarks — isolate… → Trace-driven or production-derived… → Real deployment — production or…
  • Auditing the evaluation of an NSDI submission — choosing traces
  • SKILL.md covers The evidence ladder, Questions before benchmarks, Baselines that fight back and Tails, variance, and time, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nsdi Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking baselines that fight back, and reporting tail behavior so networked-systems reviewers can see where the design wins, loses, and breaks.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Chaos engineering. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the evaluation of an NSDI submission — choosing traces
  • Deployment evidence
  • Sizing scale and failure-injection experiments
  • Picking baselines that fight back

Example prompts

  • “/nsdi-experiments”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Synthetic microbenchmarks — isolate mechanism costs. Necessary, never
  2. Trace-driven or production-derived workloads on a testbed — the venue's
  3. Real deployment — production or long-running operational use. This is

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nsdi Experiments loads about 1.6k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 718 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 718 words, ~1,620 tokens.

Download SKILL.mdSave it as .claude/skills/nsdi-experiments/SKILL.md (or your agent's skills folder).
name
nsdi-experiments
description
Use when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking baselines that fight back, and reporting tail behavior so networked-systems reviewers can see where the design wins, loses, and breaks.

NSDI Experiments

NSDI's phrase is "practical evaluation," and its reviewer culture decodes that as: realistic traffic, honest baselines, visible tails, and at least one experiment where the system is hurt on purpose. Design the evaluation as a set of questions the paper must answer, then build the smallest experiment matrix that answers them.

The evidence ladder

Climb as high as the project honestly can, and say plainly which rung you are on:

  1. Synthetic microbenchmarks — isolate mechanism costs. Necessary, never sufficient; a microbenchmark-only evaluation reads as a workshop draft.
  2. Trace-driven or production-derived workloads on a testbed — the venue's workhorse rung. Provenance of the trace matters as much as its size: say where the call graphs, flow mixes, or arrival processes come from and what was scaled.
  3. Real deployment — production or long-running operational use. This is operational-track territory and the strongest evidence NSDI recognizes; do not imply it with wording if you are on rung 2 (nsdi-topic-selection).

Questions before benchmarks

Write the evaluation section's subsection titles as questions first — Does the lease mechanism help under regional congestion? What does it cost at baseline? When does it misfire? — then design one experiment per question. The inverted approach (run everything, narrate survivors) produces the benchmark tour that NSDI reviews call unfocused.

A minimal matrix for a design-track paper:

QuestionExperimentMetrics that answer it
Does it work under the motivating pain?replay of the incident-class workloadp99/p99.9 latency, goodput during events
What does it cost when the pain is absent?baseline weeks, no faultsmedian latency, CPU/memory/bandwidth overhead
Why does it work?component breakdown / ablationper-mechanism contribution
Does it scale?node / connection / load sweepsknee location, per-node cost curve
When does it break?adversarial or boundary regimesthe regime where baselines win
Does it survive failures?injected partitions, crashes, stragglersrecovery time, correctness under churn

Baselines that fight back

  • Compare against the deployed incumbent (the mature system operators actually run), not only the nearest academic prototype — and tune it the way its own documentation says to. An untuned baseline is the most common credibility wound in systems reviewing.
  • Include the do-less baseline: the trivial fix (bigger timeout, more replicas, overprovisioning). If the paper cannot beat overprovisioning at equal cost, that is the finding.
  • When a competitor cannot be run (closed source, unavailable hardware), reimplement the mechanism and label it a reimplementation, or compare on published numbers with the configuration deltas stated.
Show full SKILL.md (318 more words)Show less

Tails, variance, and time

Networked-systems phenomena live in distributions and in time:

  • Report p99 and above for any latency claim; medians alone are treated as hiding something. Show distributions (CDFs) for headline results.
  • Congestion and failover results are time-dependent — plot the event timeline (before/during/after), not just aggregates over the run.
  • Repeat runs across enough trace segments or days to show run-to-run spread; state what varies between repeats (seeds, trace slices, background load) and report ranges, not single best runs.
  • "Up to N×" without the distribution behind it should not survive review — or your own audit (nsdi-writing-style).

Testbed and trace hygiene

text
Per-experiment provenance record (kept as the runs happen, machine-readable):
  topology: nodes, cores/RAM, NIC speed, switch model, RTT matrix
  software: kernel, framework versions, config diffs from defaults
  workload: trace id + collection context + scaling transform
  fault schedule: what was injected, when, by what tool
  outputs: raw logs location + commit hash of analysis scripts

This record is simultaneously the reproducibility ledger (nsdi-reproducibility), the artifact-evaluation seed (nsdi-artifact-evaluation), and the insurance policy for a one-shot revision that demands re-running experiments months later on the same setup (nsdi-author-response).

Budgeting machine time against the two-deadline calendar

Evaluation scope should be sized to the gate being targeted (nsdi-workflow): the question map above, costed in machine-days, tells you whether the fall gate is reachable or the plan is quietly a spring plan. Two rules of thumb from systems deadline archaeology: trace-replay pipelines take twice as long to stabilize as to run, and the break-it experiments — the ones reviewers value most — are always the ones cut when the schedule slips. Protect them by running the failure-injection matrix before the final scale sweeps, not after.

Audit checklist

  • Every evaluation subsection answers a named question.
  • Trace/workload provenance stated; scaling transforms disclosed.
  • Incumbent-grade baseline present and tuned; do-less baseline present.
  • Tail percentiles + distributions for latency claims; event timelines for dynamic behavior.
  • At least one experiment the design does not win, discussed rather than buried.
  • Failure injection covers the failure modes the design claims to handle.
  • Numbers in abstract/intro regenerate from the recorded runs.
  • Downscaled variant of the headline experiment exists (artifact evaluators and revision reviewers will not have your topology).
  • Overheads reported at baseline, not only under the motivating stress.

Output format

text
[Evidence rung] micro / trace+testbed / deployment (claimed vs actual)
[Question map] question -> experiment -> metric (gaps flagged)
[Baseline audit] incumbent? tuned? do-less alternative?
[Tail report] percentiles + distribution figures present? y/n per claim
[Break experiment] regime where the design loses: <named or MISSING>
[Priority additions] ordered by review-risk reduction per machine-week

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in NSDI-Skills/skills/nsdi-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Nsdi Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nsdi Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nsdi Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.6kAutomated safety check: PassMIT
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Audit Reviewtestflows/TestFlows-GitHub-Hetzner-Runners102—~2.1kAutomated safety check: PassCustom licence
Chaos EngineerJeffallan/claude-skills12k—~1.8kAutomated safety check: PassMIT
Chaos Dr Testharness/harness-skills115—~2.6kAutomated safety check: PassApache-2.0
Chaos Experimentharness/harness-skills115—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Audit Review

    testflows/TestFlows-GitHub-Hetzner-Runners

    Perform deep feature audits with transition-matrix and logical fault-injection validation.

    102 GitHub stars~2.1k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Chaos Engineer

    Jeffallan/claude-skills

    Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings.

    12k GitHub stars~1.8k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Chaos Dr Test

    harness/harness-skills

    A skill your agent uses when working with Chaos Engineering steps inside a Harness pipeline.

    115 GitHub stars~2.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Chaos Experiment

    harness/harness-skills

    A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…

    115 GitHub stars~1.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Categories

Questions about Nsdi Experiments

What does Nsdi Experiments do?

A skill your agent uses when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking…. Nsdi Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking baselines that fight back, and reporting tail behavior so networked-systems reviewers can see where the design wins, loses, and breaks.

When should I use Nsdi Experiments?

Nsdi Experiments fits situations like: auditing the evaluation of an NSDI submission — choosing traces; deployment evidence; sizing scale and failure-injection experiments; picking baselines that fight back.

How do I install Nsdi Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill nsdi-experiments -a claude-code`. Or copy the skill folder (NSDI-Skills/skills/nsdi-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/nsdi-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Nsdi Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill nsdi-experiments -a codex`. Or copy the skill folder (NSDI-Skills/skills/nsdi-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/nsdi-experiments in your project. Codex loads it when a task matches its description.

Can I use Nsdi Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill nsdi-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nsdi-experiments, .gemini/skills/nsdi-experiments, .github/skills/nsdi-experiments and .opencode/skills/nsdi-experiments in your project.

What does Nsdi Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Nsdi Experiments is instructions for the agent only.

Does Nsdi Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nsdi Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nsdi Experiments use?

Nsdi Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nsdi Experiments use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nsdi Experiments?

Skills that share tags, products or a category with Nsdi Experiments: Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Audit Review (testflows/TestFlows-GitHub-Hetzner-Runners, 102 stars), Chaos Engineer (Jeffallan/claude-skills, 12k stars) and Chaos Dr Test (harness/harness-skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nsdi Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.