A skill your agent uses when designing or auditing IROS experiments — real-robot trial counts, success criteria set in advance, reset procedures, failure taxonomies, baseline fairness on matched…

MITAuto-check passedData & Analytics

Install Iros Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill iros-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills iros-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/IROS-Skills/skills/iros-experiments .claude/skills/iros-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
iros-experiments
GitHub stars
1.2k
Token cost
~1k tokens
SKILL.md length
464 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing IROS experiments — real-robot trial counts, success criteria set in advance, reset procedures, failure taxonomies, baseline fairness on matched…

  • Auditing IROS experiments — real-robot trial counts
  • SKILL.md covers Experiment audit, The evidence ladder, Small-n statistics for real… and Vignette: a manipulation…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Success criteria set in advance

What it does

Iros Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing IROS experiments — real-robot trial counts, success criteria set in advance, reset procedures, failure taxonomies, baseline fairness on matched hardware, sim-to-real gap reporting, small-n statistics, and the claim-to-evidence ladder that embodied-systems reviewers apply before they trust a demo.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Statistics. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing IROS experiments — real-robot trial counts
  • Success criteria set in advance
  • Reset procedures
  • Failure taxonomies

Example prompts

  • “/iros-experiments”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Iros Experiments loads about 1k tokens when it runs. Until then it costs about 86 tokens; SKILL.md has 464 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 464 words, ~1,012 tokens.

Download SKILL.mdSave it as .claude/skills/iros-experiments/SKILL.md (or your agent's skills folder).
name
iros-experiments
description
Use when designing or auditing IROS experiments — real-robot trial counts, success criteria set in advance, reset procedures, failure taxonomies, baseline fairness on matched hardware, sim-to-real gap reporting, small-n statistics, and the claim-to-evidence ladder that embodied-systems reviewers apply before they trust a demo.

IROS Experiments

Use this before submission when the empirical story is not yet locked. At IROS, experiments exist to prove a robot did something reliably, under stated conditions — not to top a leaderboard.

Experiment audit

  • Map each claim to the evidence that supports it: a real-robot trial set, a simulation study, an ablation, or a field deployment. A claim with no matched evidence is an overclaim.
  • Define the success criterion before running, in one sentence a skeptic would accept, and hold to it; a criterion invented after seeing results is a reviewer red flag.
  • Report trial counts and resets: how many attempts, how start conditions were randomized, and what reset happened between trials. Hidden resets inflate reliability.
  • Publish a failure taxonomy with counts, not just a success rate; the failures are where reviewers calibrate trust.
  • Make baselines fair: run them on the same platform, sensors, and route, with comparable tuning effort, or state precisely why not.
  • Separate simulation from real and report the sim-to-real gap as a number; an implied zero gap is the fastest way to lose a reviewer.

The evidence ladder

Claim altitudeEvidence IROS expectsReject pattern avoided
"The system works"Trials with n, success interval, and resets stated"One hero run shown as if typical"
"It is reliable"Failure taxonomy with counts across conditions"Success rate with no failures reported"
"It transfers"Real-robot numbers plus the measured sim-to-real gap"Sim results implying real performance"
"It beats prior work"Same-hardware baseline on the same task"Comparison against a weaker or re-tuned baseline"
"It runs onboard"Measured rate and power under the real compute budget"Real-time asserted, never measured"
Show full SKILL.md (193 more words)Show less

Small-n statistics for real robots

Real trials are expensive, so n is small — but small n does not excuse a bare mean. Report a success count as a proportion with a confidence interval (a Wilson interval behaves better than normal approximation at small n), and for paired system-vs-baseline comparisons on the same trials, prefer a paired test over independent means. State n every time; "usually succeeds" is not a measurement.

Vignette: a manipulation reliability study

Suppose the system claims reliable grasping of unseen objects. The matching plan: fix an object set and a success criterion (lifted and held 3 seconds), run a stated number of trials per object with randomized poses, log every failure by cause (slip, mis-localization, collision), and report the per-object and pooled success rates with intervals. A simulation sweep over object mass then maps where the grasp model degrades, and the real-vs-sim gap is stated — every panel tied to a specific claim.

text
Trial-logging template (one row per attempt):
  trial_id, object/scenario, start_pose_seed, outcome{success|fail},
  failure_cause, reset_type, wall_time, notes
Aggregate to: success rate + interval per condition, failure histogram, sim-to-real delta.

Reporting floor

  • Every reliability figure carries n and a criterion; every timing claim carries a measured rate and the compute/power budget it ran under.
  • Report the compute actually consumed on the robot, not a desktop proxy.

Output format

text
[Experiment readiness] strong / adequate / weak
[Claim -> evidence map] <claim: trials/sim/ablation/deployment>
[Missing evidence] <trials/resets/failures/baseline/transfer>
[Statistics] criterion set? interval reported? paired where paired?
[Decision-critical next run] <one experiment on the robot>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in IROS-Skills/skills/iros-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Iros Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Iros Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Iros Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1kAutomated safety check: PassMIT
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.9k3 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone
Statistical Powerspacering-net/codeg3.9k1 repos~3.6kAutomated safety check: NotesMIT

Similar skills

  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.9k GitHub starsUsed in 3 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.9k GitHub starsUsed in 1 repo~3.6k tokens
    Data & AnalyticsAuto-check: notes
  • Agent Session Monitor

    higress-group/higress

    Real-time agent conversation monitoring - monitors Higress access logs, aggregates conversations by session, tracks token usage.

    9.5k GitHub stars~3.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Iros Experiments

What does Iros Experiments do?

A skill your agent uses when designing or auditing IROS experiments — real-robot trial counts, success criteria set in advance, reset procedures, failure taxonomies, baseline fairness on matched…. Iros Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing IROS experiments — real-robot trial counts, success criteria set in advance, reset procedures, failure taxonomies, baseline fairness on matched hardware, sim-to-real gap reporting, small-n statistics, and the claim-to-evidence ladder that embodied-systems reviewers apply before they trust a demo.

When should I use Iros Experiments?

Iros Experiments fits situations like: auditing IROS experiments — real-robot trial counts; success criteria set in advance; reset procedures; failure taxonomies.

How do I install Iros Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill iros-experiments -a claude-code`. Or copy the skill folder (IROS-Skills/skills/iros-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/iros-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Iros Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill iros-experiments -a codex`. Or copy the skill folder (IROS-Skills/skills/iros-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/iros-experiments in your project. Codex loads it when a task matches its description.

Can I use Iros Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill iros-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/iros-experiments, .gemini/skills/iros-experiments, .github/skills/iros-experiments and .opencode/skills/iros-experiments in your project.

What does Iros Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Iros Experiments is instructions for the agent only.

Does Iros Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Iros Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Iros Experiments use?

Iros Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Iros Experiments use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Iros Experiments?

Skills that share tags, products or a category with Iros Experiments: Sandbox Bench (vercel/next.js, 143k stars), Statistical Analysis (spacering-net/codeg, 3.9k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars) and AI Daily Digest (vigorX777/ai-daily-digest, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Iros Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.