Agent skill

Naacl Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact…

MITAuto-check passedAI & LLM Engineering

Install Naacl Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .claude/skills/naacl-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
naacl-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.5k tokens
SKILL.md length
622 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact…

  • Packaging datasets
  • SKILL.md covers What the artifact must let a…, Language-data documentation…, Anonymous packaging that… and Vignette: a Quechua-Spanish…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklists artifact questions

What it does

Naacl Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact questions, documenting provenance and licensing for language data, and handling community-owned or Indigenous-language resources correctly.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Packaging datasets
  • Annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklists artifact questions
  • Documenting provenance and licensing for language data
  • Handling community-owned

Example prompts

  • “/naacl-artifact-evaluation”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Naacl Artifact Evaluation loads about 1.5k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 622 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 622 words, ~1,462 tokens.

Download SKILL.mdSave it as .claude/skills/naacl-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
naacl-artifact-evaluation
description
Use when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact questions, documenting provenance and licensing for language data, and handling community-owned or Indigenous-language resources correctly.

NAACL Artifact Evaluation

NAACL has no separate artifact-badging track; artifacts are judged inside the ARR review itself, through the supplement upload and section B of the Responsible NLP checklist ("scientific artifacts"). That placement matters: your artifact documentation is not an optional extra but a set of sworn answers reviewers cross-examine against the PDF.

What the artifact must let a reviewer do

  • Trace provenance. Where did every corpus come from, under what license, and does your use match the terms and the creators' intent?
  • Inspect the instrument. Prompts, annotation guidelines, interface screenshots, and pay rates are artifacts too — a human-evaluation claim without its instrument is unverifiable.
  • Rerun the cheap parts. Scoring scripts and metric code should execute from the archive alone; nobody will retrain your model, but everybody can re-score your outputs if you include them.
  • Audit the data card. Language varieties, dialect coverage, speaker demographics where relevant, known gaps, and intended use.

Language-data documentation ladder

Data situationMinimum documentation for a NAACL reviewerExtra step
Standard public benchmarkVersion, split, license, citationNote any known contamination reports
Web-scraped textCollection dates, filtering rules, deduplication, license basisPII handling statement
New annotated corpusGuidelines, annotator recruitment and pay, agreement scoresRelease the guidelines verbatim in the supplement
Dialectal / code-switched dataVariety labels and how they were assignedNative-speaker validation description
Indigenous or community-owned language dataConsent and partnership terms, community approval for releaseVerify whether public release is permitted at all

The last row is a NAACL signature concern. Work on languages of the Americas increasingly follows community-controlled data norms: some corpora may be used but not redistributed, some require named attribution (which conflicts with anonymous review — use a placeholder and restore at camera-ready), and some communities set conditions on derived models. "We release everything" is not automatically the ethical high ground here; the checklist rewards accuracy about constraints, not maximal openness.

Anonymous packaging that survives inspection

text
artifact.zip
├── README.md          # one-screen orientation: what, how, how long
├── data/
│   ├── data_card.md   # provenance, license, varieties, gaps
│   └── samples/       # enough rows to judge quality, not the corpus
├── prompts/           # exact strings, all variants tried
├── eval/
│   ├── score.py       # runs on outputs/ with no network access
│   └── outputs/       # raw model outputs backing the main tables
└── annotation/
    └── guidelines.pdf # the instrument, scrubbed of institution marks

Scrub before zipping: repository history, notebook execution metadata, absolute paths with usernames, license headers naming the lab, and any consent form carrying institutional letterhead (replace with a redacted copy; note that the original exists).

Show full SKILL.md (272 more words)Show less

Vignette: a Quechua-Spanish parallel corpus package

A submission introduces a 40k-pair Quechua-Spanish parallel corpus built with two community organizations, plus MT baselines. The packaging calls that follow from this skill:

  • The corpus itself does not ship in the review archive — the partnership terms permit research use but defer public release to a community decision. The data card states this, and section B of the checklist answers "no, with reason" for artifact release.
  • What does ship: 200 sample pairs cleared for review purposes, the collection protocol, annotator recruitment and payment description, the cleaning scripts, and the full MT evaluation pipeline with outputs.
  • The consent-form template ships with the letterhead redacted and a note that the original is held by the partner organizations.
  • The README's first paragraph tells the reviewer exactly which claims the archive can and cannot let them verify — pre-empting the "authors refuse to release data" misreading with a governance explanation instead.

The result is an artifact that scores as honest and inspectable rather than incomplete, which is the realistic best outcome for community-governed data.

Cycle-volatile mechanics

Archive size caps, accepted formats, and whether supplements upload as one file or several are OpenReview-form details that shift between ARR cycles; read the live submission form before building the final zip, and never reverse-engineer the limits from a previous cycle's folklore.

Post-acceptance conversion

At camera-ready, the anonymous bundle becomes the public record: move it to a persistent host with a DOI or a tagged release, apply the real license, restore attribution the community partnership requires, and update the checklist-facing statements in the paper if the release scope changed between review and publication.

Output format

text
[Artifact inventory] <data / code / prompts / guidelines / outputs>
[Checklist B alignment] <each B answer -> where the artifact proves it>
[Provenance gaps] <unlicensed, undocumented, or unclear-consent items>
[Community constraints] <redistribution / attribution / approval terms>
[Anonymity sweep] clean / issues found
[Release plan] <anonymous now -> public form at camera-ready>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in NAACL-Skills/skills/naacl-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Naacl Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Naacl Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Naacl Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
OpenMed Model Card Writermaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 14 days ago
    Auto-check passed

Questions about Naacl Artifact Evaluation

What does Naacl Artifact Evaluation do?

A skill your agent uses when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact…. Naacl Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact questions, documenting provenance and licensing for language data, and handling community-owned or Indigenous-language resources correctly.

When should I use Naacl Artifact Evaluation?

Naacl Artifact Evaluation fits situations like: packaging datasets; annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklists artifact questions; documenting provenance and licensing for language data; handling community-owned.

How do I install Naacl Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a claude-code`. Or copy the skill folder (NAACL-Skills/skills/naacl-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/naacl-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Naacl Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a codex`. Or copy the skill folder (NAACL-Skills/skills/naacl-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/naacl-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Naacl Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/naacl-artifact-evaluation, .gemini/skills/naacl-artifact-evaluation, .github/skills/naacl-artifact-evaluation and .opencode/skills/naacl-artifact-evaluation in your project.

What does Naacl Artifact Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Naacl Artifact Evaluation is instructions for the agent only.

Does Naacl Artifact Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Naacl Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Naacl Artifact Evaluation use?

Naacl Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Naacl Artifact Evaluation use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Naacl Artifact Evaluation?

Skills that share tags, products or a category with Naacl Artifact Evaluation: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and Andrej Karpathy (K-Dense-AI/mimeo, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Naacl Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.