A skill your agent uses when designing or auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims, keeping cross-lingual comparisons budget-fair…

MITAuto-check passed

Install Naacl Experiments

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/NAACL-Skills/skills/naacl-experiments .claude/skills/naacl-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
naacl-experiments
GitHub stars
1.2k
Token cost
~1.5k tokens
SKILL.md length
729 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims, keeping cross-lingual comparisons budget-fair…

  • Works in 5 steps: The scope probe: claim's language list… → The fairness probe: was the strongest… → The contamination probe: could the eval… → …
  • Auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims
  • SKILL.md covers Match the design to the…, Budget-fair comparison rules, Variance and significance floor and The probes NAACL reviewers run, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Naacl Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims, keeping cross-lingual comparisons budget-fair, testing on natively authored rather than translated data where the claim requires it, and reporting variance that survives reviewer probing.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims
  • Keeping cross-lingual comparisons budget-fair
  • Testing on natively authored rather than translated data where the claim requires it
  • Reporting variance that survives reviewer probing

Example prompts

  • “/naacl-experiments”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. The scope probe: claim's language list vs. tables' language list.
  2. The fairness probe: was the strongest baseline given a real chance?
  3. The contamination probe: could the eval data sit in pretraining?
  4. The mechanism probe: does any analysis show why the gain occurs
  5. The human-eval probe: who annotated, in which language variety, paid

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Naacl Experiments loads about 1.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 729 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 729 words, ~1,499 tokens.

Download SKILL.mdSave it as .claude/skills/naacl-experiments/SKILL.md (or your agent's skills folder).
name
naacl-experiments
description
Use when designing or auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims, keeping cross-lingual comparisons budget-fair, testing on natively authored rather than translated data where the claim requires it, and reporting variance that survives reviewer probing.

NAACL Experiments

Design experiments backwards from the sentence you want the meta-review to contain. For NAACL-bound work that sentence almost always has a language scope in it, so the experimental program's first duty is to make the scope claim measurable — and its second duty is to make every comparison fair enough that no single reviewer probe collapses it.

Match the design to the coverage claim

Claim you want to makeMinimum design that supports itDesign that fakes it
"Works for language X"Natively authored X test data, native-speaker error reviewMachine-translated English benchmark relabeled as X
"Works across the Americas' languages"Typologically spread sample (e.g., analytic + agglutinative + polysynthetic)Three Romance languages standing in for a continent
"Robust to dialectal variation"Variety-labeled eval sets, per-variety breakdownOne standard variety plus vibes
"Better than baseline B"B re-run under equal tuning/compute budget, same prompts regimeB's two-year-old published number
"Model-agnostic"≥3 model families, sizes reportedTwo checkpoints of one family

Translationese deserves its own line of caution: test sets translated from English preserve English information structure and topic distribution, so gains measured on them may be gains at modeling translation artifacts. Where natively authored data cannot be had, say so and mark the limitation — do not let the table caption imply nativeness the data lacks.

Budget-fair comparison rules

  • Every system in a table gets the same tuning attempts, the same prompt engineering effort, and the same test-set exposure; report the attempt counts.
  • Separate "our method with our infrastructure" from "baseline as published" — mixing the two in one column is the most common fairness foul.
  • When a hosted API model is a baseline, record query dates and note that the comparison is against a moving target.

Variance and significance floor

  • Report mean and spread over ≥3 seeds (or bootstrap over test items when training is deterministic); a caption states which.
  • Pair headline comparisons with a significance or effect-size statement appropriate to the metric; per-language results additionally need per-language n, because a 4-point gain on a 200-item Quechua set and on a 20,000-item Spanish set are different objects.
  • Small multilingual test sets make single-run deltas noise; if the per-language sample is small, aggregate honestly and show the intervals.

The probes NAACL reviewers run

  1. The scope probe: claim's language list vs. tables' language list.
  2. The fairness probe: was the strongest baseline given a real chance?
  3. The contamination probe: could the eval data sit in pretraining? State your screening method even when the answer is "cannot rule out."
  4. The mechanism probe: does any analysis show why the gain occurs (error classes, ablations), or only that it occurs?
  5. The human-eval probe: who annotated, in which language variety, paid how, agreeing how much?

Design so each probe has a prepared landing spot in the paper.

Show full SKILL.md (266 more words)Show less

Experiment ledger

text
# One row per run, committed with the code
run_id, task, lang, variety, model@rev, prompt_id, seed,
train_data@hash, test_data@hash, metric, value, gpu_h, date
# Tables in the paper are views over this ledger — nothing enters
# a table that lacks a row, and per-language n comes along free.

The ledger sounds bureaucratic until the response window, when "R1 asks for es-MX vs es-AR breakdown" becomes a ten-minute query instead of a lost weekend.

Vignette: a dialect-identification study, probe by probe

Fictional setup: classifying Caribbean vs. Andean vs. Rioplatense Spanish in social media text, claiming "robust dialect ID across Latin American varieties."

  • Scope probe: three varieties do not license "across Latin American varieties" — either add Central American and Mexican data or rescope the claim to the three tested.
  • Fairness probe: the strongest baseline is a fine-tuned multilingual encoder from a prior paper; it gets re-tuned on this study's training data with the same search budget, and both budgets appear in a footnote.
  • Contamination probe: the test tweets postdate every evaluated model's training cutoff where cutoffs are published; where they are not, the caption says so.
  • Mechanism probe: a confusion analysis shows the Caribbean-Andean errors concentrate in short, lexically neutral posts — evidence the model reads topic and orthography, not dialect, in those cases.
  • Human-eval probe: variety labels came from annotator self-identified L1 region plus a validation round; agreement and the adjudication rule are reported next to the label counts.

The vignette's lesson: every probe answer becomes a sentence or a caption in the paper, not a private reassurance.

Audit sequence

  1. Write the target meta-review sentence; extract its claims.
  2. Map each claim to the table/figure that carries it; kill or downscope orphans.
  3. Run the five probes adversarially against your own draft.
  4. Check every comparison for budget fairness; annotate exceptions.
  5. Verify variance reporting exists for every number bolded anywhere.

Output format

text
[Target sentence] <the meta-review sentence>
[Claim -> evidence map] <claim: table/figure/analysis>
[Probe results] scope / fairness / contamination / mechanism / human-eval
[Variance status] <where missing>
[Next decisive run] <one experiment, why it changes the decision>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in NAACL-Skills/skills/naacl-experiments of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Naacl Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Naacl Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Naacl Experiments this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Design Audit Against Rams' Principlesthedotmack/claude-mem99k—~4.6kAutomated safety check: PassApache-2.0
Experiment Designeralirezarezvani/claude-skills28k1 repos~783Automated safety check: PassMIT
Experimental Designaiming-lab/AutoResearchClaw15k—~286Automated safety check: PassMIT
Experiment Auditwanshuiyin/Auto-claude-code-research-in-sleep17k1 repos~2.7kAutomated safety check: NotesMIT
Experiment Auditwanshuiyin/Auto-claude-code-research-in-sleep17k—~3.2kAutomated safety check: NotesMIT

Similar skills

  • Audits a design against Dieter Rams' ten principles of good design, scores each with evidence, and hands off a make-plan prompt for a new, refined or redesigned outcome.

    99k GitHub stars~4.6k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Experiment Designer

    alirezarezvani/claude-skills

    A skill your agent uses when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

    28k GitHub starsUsed in 1 repo~783 tokens
    Research & ScienceAuto-check passed
  • Experimental Design

    aiming-lab/AutoResearchClaw

    Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw.

    15k GitHub stars~286 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub starsUsed in 1 repo~2.7k tokens
    DatabasesAuto-check: notes
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub stars~3.2k tokensUpdated 2 days ago
    DatabasesAuto-check: notes
  • OpenClaw Design Audit

    openclaw/clawhub

    Audits OpenClaw frontend code and rendered pages for token misuse, reimplemented primitives, accessibility and responsive defects and off-brand copy, with an evidence-based report.

    9.5k GitHub stars~498 tokensUpdated yesterday
    Frontend & DesignAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed

Questions about Naacl Experiments

What does Naacl Experiments do?

A skill your agent uses when designing or auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims, keeping cross-lingual comparisons budget-fair…. Naacl Experiments is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when designing or auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims, keeping cross-lingual comparisons budget-fair, testing on natively authored rather than translated data where the claim requires it, and reporting variance that survives reviewer probing.

When should I use Naacl Experiments?

Naacl Experiments fits situations like: auditing the experimental program of a NAACL submission — matching evidence to language-coverage claims; keeping cross-lingual comparisons budget-fair; testing on natively authored rather than translated data where the claim requires it; reporting variance that survives reviewer probing.

How do I install Naacl Experiments in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-experiments -a claude-code`. Or copy the skill folder (NAACL-Skills/skills/naacl-experiments in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/naacl-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Naacl Experiments in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-experiments -a codex`. Or copy the skill folder (NAACL-Skills/skills/naacl-experiments in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/naacl-experiments in your project. Codex loads it when a task matches its description.

Can I use Naacl Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/naacl-experiments, .gemini/skills/naacl-experiments, .github/skills/naacl-experiments and .opencode/skills/naacl-experiments in your project.

What does Naacl Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Naacl Experiments is instructions for the agent only.

Does Naacl Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Naacl Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Naacl Experiments use?

Naacl Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Naacl Experiments use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Naacl Experiments?

Skills that share tags, products or a category with Naacl Experiments: Design Audit Against Rams' Principles (thedotmack/claude-mem, 99k stars), Experiment Designer (alirezarezvani/claude-skills, 28k stars), Experimental Design (aiming-lab/AutoResearchClaw, 15k stars) and Experiment Audit (wanshuiyin/Auto-claude-code-research-in-sleep, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Naacl Experiments?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,228 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.