Agent skill

Interspeech Experiments

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when designing or auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER, MOS/CMOS, EER/minDCF, PESQ/STOI), baselines reviewers accept…

MITAuto-check passed

Install Interspeech Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill interspeech-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills interspeech-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/INTERSPEECH-Skills/skills/interspeech-experiments .claude/skills/interspeech-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
interspeech-experiments
GitHub stars
1.2k
Token cost
~1.6k tokens
SKILL.md length
657 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER, MOS/CMOS, EER/minDCF, PESQ/STOI), baselines reviewers accept…

  • Auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER
  • SKILL.md covers Metric-task law, Baselines that count, Significance: over what… and Condition coverage — the…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Baselines reviewers accept

What it does

Interspeech Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER, MOS/CMOS, EER/minDCF, PESQ/STOI), baselines reviewers accept, significance testing over utterances and seeds, condition coverage across speakers/noise/languages, and data hygiene for speech corpora.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER
  • Baselines reviewers accept
  • Significance testing over utterances and seeds
  • Condition coverage across speakers/noise/languages

Example prompts

  • “/interspeech-experiments”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Interspeech Experiments loads about 1.6k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 657 words, ~1,624 tokens.

Download SKILL.mdSave it as .claude/skills/interspeech-experiments/SKILL.md (or your agent's skills folder).
name
interspeech-experiments
description
Use when designing or auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER, MOS/CMOS, EER/minDCF, PESQ/STOI), baselines reviewers accept, significance testing over utterances and seeds, condition coverage across speakers/noise/languages, and data hygiene for speech corpora.

INTERSPEECH Experiments

"Not convincing" is the standard Interspeech rejection, and it almost always means the experimental design — not the idea — failed. Speech evaluation has decades of conventions per task; an experiment section that ignores them is illegible to the reviewer pool regardless of how good the numbers are.

Metric-task law

Task familyPrimary metricsConvention notes
ASRWER / CERnormalization + scorer disclosed; CER for unsegmented scripts
TTS / VCMOS, CMOS (+ objective proxies)panel protocol reported; CMOS for close systems
Speaker verificationEER, minDCFofficial trial lists; DCF prior/costs stated
DiarizationDER / JERcollar and overlap handling stated
Enhancement / separationPESQ, STOI/ESTOI, SI-SDR (+ DNSMOS-style proxies)wideband vs narrowband named
SLU / speech translationintent acc / F1, BLEU/COMET on ASR outputcascaded vs end-to-end made explicit
Paralinguistics / healthUAR, F1speaker-disjoint splits are mandatory

Using a proxy where the community expects the primary (e.g., only neural MOS predictors for a TTS claim) needs an explicit defense sentence.

Baselines that count

  • The stock recipe of a public toolkit on the same corpus (an ESPnet/ SpeechBrain recipe number is a shared, checkable baseline).
  • The latest challenge baseline if your task has a running challenge — reviewers know those numbers by heart.
  • Your method's ablated self — at Interspeech, one clean ablation of the single proposed component often persuades more than two extra datasets.
  • Reimplemented prior work must be validated: show your reimplementation matches its published number before showing you beat it.

Significance: over what randomness?

State which variation your statistics cover — the two are routinely conflated:

  • Test-set variation: bootstrap over utterances (or speakers, if claims are speaker-level) → CI on the metric difference between systems.
  • Training variation: multiple seeds → mean ± sd; a 0.2 WER gain with sd 0.3 across seeds is not a result.
  • Matched-pairs tests (paired bootstrap; the classic MAPSSWE-style segment test for ASR) for A-vs-B claims on the same test set.
  • Subjective scores: CIs over raters and stimuli; ±0.1 MOS is panel noise under most protocols (see interspeech-reproducibility).

Condition coverage — the speech-specific axis

A speech claim is implicitly quantified over speakers, acoustic conditions, and often languages. Reviewers probe the quantifier:

  • Speaker-disjointness: train/test speaker overlap invalidates verification, paralinguistic, and health claims outright.
  • Condition breakdown: report clean vs noisy, near- vs far-field, read vs spontaneous where the corpus offers them — an average hides the regression your method causes in one condition.
  • Language scope: "multilingual" needs a per-language table; English-only results support English-only claims.
  • Domain leakage: SSL pretraining data overlapping the test corpus (the LibriSpeech-descendant problem) must be checked and stated for foundation-model work.
Show full SKILL.md (242 more words)Show less

Data hygiene

  • Official partitions only, or published manifests for custom splits.
  • No tuning on test: LM weights, thresholds, and checkpoint selection all happen on dev — say so in one line.
  • License and consent status of every corpus stated (see interspeech-artifact-evaluation); leaked or scraped audio can sink an otherwise strong paper on ethics review.

Designing inside 4 pages

Budget roughly one column for the decisive comparison, half for ablation, half for analysis. The analysis half is what separates accepted Interspeech papers: one error-pattern finding (where the gains live — short utterances, overlapping speech, a phone class) converts a benchmark delta into a scientific statement.

Worked micro-example: is 4.9 vs 5.6 WER real?

text
Claim: proposed 4.9% vs baseline 5.6% WER on test-other (2939 utts).
1. Paired per-utterance errors → paired bootstrap, 1000 resamples.
2. Δ WER 95% CI: [-0.9, -0.5] — excludes 0 → test-set variation covered.
3. Across 3 seeds: 4.9, 5.0, 4.8 (sd 0.1) vs 5.6, 5.7, 5.6 (sd 0.06)
   → training variation does not swallow the gap.
4. Report: "−0.7 abs. WER (95% CI [−0.9, −0.5], paired bootstrap;
   consistent across 3 seeds)."

Two randomness sources, two checks, one sentence in the paper. If step 2's CI had straddled zero, the honest paper reports the trend and softens the verb — and usually survives review better than the inflated version.

Negative results and regressions

Interspeech's mixed jury respects a disclosed regression far more than a suspicious clean sweep. If the method loses on clean speech while winning on noisy, print both numbers and make the trade-off the story — condition-dependent behavior is a finding in a field about acoustic variability, and hiding it is the reviewer-trust equivalent of a failed significance test.

Pre-submission experiment audit

text
[ ] Primary metric matches task convention; ruler disclosed
[ ] Baseline set includes a public-recipe or challenge anchor
[ ] Each A>B claim carries a CI or matched-pairs test
[ ] Seeds: n stated; variance reported or single-run admitted
[ ] Speaker-disjoint splits verified where required
[ ] Condition/language breakdown present; regressions named
[ ] Dev-only tuning stated; test touched once
[ ] One analysis finding, not just deltas

Output format

text
[Claim inventory] each claim → metric → evidence status
[Metric-law check] conventions met / violations
[Baseline verdict] anchored / self-referential / stale
[Statistics] randomness covered (test-set / seeds / raters) per claim
[Coverage gaps] speaker / condition / language / leakage
[Cheapest decisive fix] <one experiment that most raises conviction>

Metric conventions are community law rather than CFP text and move slowly, but challenge editions and recipe baselines roll every year — re-anchor at design time (sources logged in resources/official-source-map.md, 2026-07-08).

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in INTERSPEECH-Skills/skills/interspeech-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Interspeech Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Interspeech Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Interspeech Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.6kAutomated safety check: PassMIT
Socialcoreyhaines31/marketingskills54k4 repos~4.5kAutomated safety check: PassMIT
Geo Fundamentalswasp-lang/wasp19k9 repos~861Automated safety check: PassMIT
Banner Design Systemnextlevelbuilder/ui-ux-pro-max-skill134k1 repos~1.8kAutomated safety check: PassMIT
Ab Testingcoreyhaines31/marketingskills54k3 repos~3.1kAutomated safety check: PassMIT
Agent ReachPanniantong/Agent-Reach94k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Social

    coreyhaines31/marketingskills

    When the user wants help creating, scheduling, or optimizing social media content for LinkedIn, Twitter/X, Instagram, TikTok, or Facebook, or wants to do social listening and engagement triage.

    54k GitHub starsUsed in 4 repos~4.5k tokens
    Writing & ContentAuto-check passed
  • Geo Fundamentals

    wasp-lang/wasp

    Generative Engine Optimization for AI search engines (ChatGPT, Claude, Perplexity).

    19k GitHub starsUsed in 9 repos~861 tokens
    Marketing & SEOAuto-check passed
  • Banner Design System

    nextlevelbuilder/ui-ux-pro-max-skill

    Walks through designing a banner for social media, ads, a website hero or print, from gathering requirements to building 2 or 3 art-direction options in HTML and CSS.

    134k GitHub starsUsed in 1 repo~1.8k tokens
    Frontend & DesignAuto-check passed
  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~3.1k tokens
    Marketing & SEOAuto-check passed
  • Agent Reach

    Panniantong/Agent-Reach

    Routes web research and platform lookups across 16 sites, including Twitter, Reddit, YouTube, Bilibili, Xiaohongshu and GitHub, through one command-line tool.

    94k GitHub stars~1.4k tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Hreflang and International SEO

    AgriciDaniel/claude-seo

    Audits, validates and generates hreflang tags for multi-language and multi-region sites in HTML, HTTP headers or XML sitemaps, flagging common code and return-tag mistakes.

    19k GitHub starsUsed in 5 repos~3.4k tokens
    Marketing & SEOAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Interspeech Experiments

What does Interspeech Experiments do?

A skill your agent uses when designing or auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER, MOS/CMOS, EER/minDCF, PESQ/STOI), baselines reviewers accept…. Interspeech Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER, MOS/CMOS, EER/minDCF, PESQ/STOI), baselines reviewers accept, significance testing over utterances and seeds, condition coverage across speakers/noise/languages, and data hygiene for speech corpora.

When should I use Interspeech Experiments?

Interspeech Experiments fits situations like: auditing the experimental evidence for an INTERSPEECH paper — task-correct metrics (WER/CER; baselines reviewers accept; significance testing over utterances and seeds; condition coverage across speakers/noise/languages.

How do I install Interspeech Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill interspeech-experiments -a claude-code`. Or copy the skill folder (INTERSPEECH-Skills/skills/interspeech-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/interspeech-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Interspeech Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill interspeech-experiments -a codex`. Or copy the skill folder (INTERSPEECH-Skills/skills/interspeech-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/interspeech-experiments in your project. Codex loads it when a task matches its description.

Can I use Interspeech Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill interspeech-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/interspeech-experiments, .gemini/skills/interspeech-experiments, .github/skills/interspeech-experiments and .opencode/skills/interspeech-experiments in your project.

What does Interspeech Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Interspeech Experiments is instructions for the agent only.

Does Interspeech Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Interspeech Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Interspeech Experiments use?

Interspeech Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Interspeech Experiments use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Interspeech Experiments?

Skills that share tags, products or a category with Interspeech Experiments: Social (coreyhaines31/marketingskills, 54k stars), Geo Fundamentals (wasp-lang/wasp, 19k stars), Banner Design System (nextlevelbuilder/ui-ux-pro-max-skill, 134k stars) and Ab Testing (coreyhaines31/marketingskills, 54k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Interspeech Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.