Agent skill

Emnlp Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when packaging the artifacts of an EMNLP paper — datasets, annotation guidelines, prompts, evaluation code, and model outputs — as anonymous review-time evidence or public…

MITAuto-check passedAI & LLM Engineering

Install Emnlp Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill emnlp-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills emnlp-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/EMNLP-Skills/skills/emnlp-artifact-evaluation .claude/skills/emnlp-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
emnlp-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.7k tokens
SKILL.md length
786 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when packaging the artifacts of an EMNLP paper — datasets, annotation guidelines, prompts, evaluation code, and model outputs — as anonymous review-time evidence or public…

  • Packaging the artifacts of an EMNLP paper — datasets
  • SKILL.md covers What NLP reviewers open, in…, Dataset packaging: the data…, Prompts and outputs as… and Review-time anonymity vs…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Annotation guidelines

What it does

Emnlp Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging the artifacts of an EMNLP paper — datasets, annotation guidelines, prompts, evaluation code, and model outputs — as anonymous review-time evidence or public post-acceptance releases, with licensing, data statements, and the inspection order NLP reviewers actually follow.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Packaging the artifacts of an EMNLP paper — datasets
  • Annotation guidelines
  • Evaluation code
  • Model outputs — as anonymous review-time evidence

Example prompts

  • “/emnlp-artifact-evaluation”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Emnlp Artifact Evaluation loads about 1.7k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 786 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 786 words, ~1,690 tokens.

Download SKILL.mdSave it as .claude/skills/emnlp-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
emnlp-artifact-evaluation
description
Use when packaging the artifacts of an EMNLP paper — datasets, annotation guidelines, prompts, evaluation code, and model outputs — as anonymous review-time evidence or public post-acceptance releases, with licensing, data statements, and the inspection order NLP reviewers actually follow.

EMNLP Artifact Evaluation

Use this for evidence packaging around an EMNLP submission. EMNLP has no separate artifact-badging committee; the artifact is part of the scientific claim, filed under the Responsible NLP checklist and inspected at reviewer discretion. In NLP the artifact surface is unusually broad — data, labels, prompts, outputs, and scoring code are all first-class — and each has its own failure mode.

What NLP reviewers open, in order

Reviewers with thirty minutes and suspicion follow a predictable path:

OrderArtifactWhat they are checkingCheap failure
1Data sampleDo instances look like the paper's description?Examples contradict claimed label definitions
2Prompt filesDo prompts match the paper's claimed setup?Prompt contains hints the paper never mentioned
3Scoring scriptIs the metric computed the standard way?Custom normalization inflates the headline metric
4Annotation guidelinesCould these instructions produce these labels?Guidelines answer a different question than the task
5Output dumpsAre generations as good as the excerpted ones?Body examples are the best 5 of 500

Package for this order: a top-level README that routes to each artifact in one hop, a data sample small enough to open in a text editor, and prompts stored as files rather than embedded in code.

Dataset packaging: the data statement standard

A released corpus travels with documentation or it travels badly:

  • Provenance — where text came from, collection dates, and the license chain that permits redistribution (scraped ≠ redistributable; say what you verified).
  • Annotation record — who labeled (recruitment, qualifications, compensation), under what guidelines (included verbatim), with what agreement (statistic and value), and how disagreements were resolved.
  • Composition — languages, domains, demographic scope of speakers and annotators where relevant; the boundaries here feed the Limitations section directly.
  • Intended use — the task the labels are valid for, and foreseeable misuses; a toxicity corpus without a misuse note reads as naive in 2026.
  • Splits and hygiene — split construction, dedup across splits, and any overlap audit against common pretraining corpora.

Prompts and outputs as archival objects

Treat prompts like code and outputs like data:

text
artifacts/
├── prompts/
│   ├── task_nli_zeroshot_v3.txt      # verbatim, one file per reported condition
│   └── CHANGELOG.md                  # v1→v3: what changed and which tables use which
├── outputs/
│   ├── model=llama3-8b_seed=42.jsonl # every generation behind every reported number
│   └── sample_100_stratified.jsonl   # reviewer-sized random sample, stratified by error class
└── scoring/
    └── score.py                      # reads outputs/, emits the paper's tables

The outputs/ directory is the underrated one: releasing full generations lets a reviewer (and later, the field) re-score with better metrics without rerunning models — the single highest-leverage reproducibility gift an NLP paper can give.

Review-time anonymity vs release-time findability

The same package lives two lives. At review: anonymized hosting, no usernames in paths, no lab-identifying data sources, license files present but grant numbers absent. After acceptance: permanent home, citable version (tag or DOI), README pointing at the Anthology entry, model cards / data statements filled with the identity-bearing detail review forbade. Build the package once with an anonymize.sh that strips the delta, rather than maintaining two diverging trees.

Show full SKILL.md (341 more words)Show less

The checklist asks about licenses in both directions — what you used and what you release — and most failures are ignorance, not malice:

  • Inbound: a benchmark's license travels with its derivatives; a corpus built by filtering a research-only dataset cannot ship under a commercial-friendly license. Record the license of every ingredient at collection time, because reconstructing the chain in deadline week fails.
  • Model outputs: generations from API models may carry terms-of-use restrictions on redistribution and on training-data use; releasing an output corpus without checking the provider's terms is the current decade's version of scraping-and- hoping.
  • Outbound: pick a standard license (CC for data, a common OSS license for code), state it in the README and in the paper, and resist inventing bespoke "research only, email us" terms that make the artifact uncitable in practice.
  • People in the data: licensing does not settle privacy; text containing personal information needs handling and redaction notes independent of copyright status.

Restricted and unreleasable material

When data cannot ship (privacy, license, platform terms): release the collection pipeline and filters instead of the corpus; provide a synthetic or public-subset sample demonstrating the format; state retention terms and whom to contact for research access. For closed API models, ship exact prompts, dates, decoding parameters, and full outputs — the model is closed; your measurement of it need not be. The checklist rewards documented honesty over silent omission in every one of these cases.

Versioning after release

An NLP artifact that succeeds gets used, and use creates obligations the paper never mentioned:

  • Tag the exact version the camera-ready numbers came from and never move that tag; fixes go in new versions with a changelog explaining what results they might shift.
  • When label errors are found post-release (they will be), publish errata in the repository rather than silently rewriting files other papers already evaluated on — benchmark drift breaks comparability for everyone downstream.
  • Keep an issue tracker open; the questions users file are a free audit of your documentation's gaps and occasionally of your paper's claims.

Output format

text
[Package role] review-time evidence / camera-ready release / both
[Inspection path] <README -> data sample -> prompts -> scoring -> outputs: intact?>
[Data statement] <provenance / annotation / composition / use / splits — gaps>
[Anonymity delta] <what anonymize.sh strips; leaks found>
[Unreleasable items] <artifact -> documented workaround>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in EMNLP-Skills/skills/emnlp-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Emnlp Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Emnlp Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Emnlp Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
OpenMed Model Card Writermaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 14 days ago
    Auto-check passed

Questions about Emnlp Artifact Evaluation

What does Emnlp Artifact Evaluation do?

A skill your agent uses when packaging the artifacts of an EMNLP paper — datasets, annotation guidelines, prompts, evaluation code, and model outputs — as anonymous review-time evidence or public…. Emnlp Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging the artifacts of an EMNLP paper — datasets, annotation guidelines, prompts, evaluation code, and model outputs — as anonymous review-time evidence or public post-acceptance releases, with licensing, data statements, and the inspection order NLP reviewers actually follow.

When should I use Emnlp Artifact Evaluation?

Emnlp Artifact Evaluation fits situations like: packaging the artifacts of an EMNLP paper — datasets; annotation guidelines; evaluation code; model outputs — as anonymous review-time evidence.

How do I install Emnlp Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill emnlp-artifact-evaluation -a claude-code`. Or copy the skill folder (EMNLP-Skills/skills/emnlp-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/emnlp-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Emnlp Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill emnlp-artifact-evaluation -a codex`. Or copy the skill folder (EMNLP-Skills/skills/emnlp-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/emnlp-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Emnlp Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill emnlp-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/emnlp-artifact-evaluation, .gemini/skills/emnlp-artifact-evaluation, .github/skills/emnlp-artifact-evaluation and .opencode/skills/emnlp-artifact-evaluation in your project.

What does Emnlp Artifact Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Emnlp Artifact Evaluation is instructions for the agent only.

Does Emnlp Artifact Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Emnlp Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Emnlp Artifact Evaluation use?

Emnlp Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Emnlp Artifact Evaluation use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Emnlp Artifact Evaluation?

Skills that share tags, products or a category with Emnlp Artifact Evaluation: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and Andrej Karpathy (K-Dense-AI/mimeo, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Emnlp Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.