Agent skill

Acl Artifact Evaluation

by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills

A skill your agent uses when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives…

MITAuto-check passedAI & LLM Engineering

Install Acl Artifact Evaluation

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills acl-artifact-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ACL-Skills/skills/acl-artifact-evaluation .claude/skills/acl-artifact-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
acl-artifact-evaluation
GitHub stars
1.2k
Token cost
~1.5k tokens
SKILL.md length
617 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives…

  • Works in 4 steps: The README — it has roughly one minute… → Prompt files and evaluation scripts, for… → Annotation guidelines, for any dataset… → …
  • Annotation materials for an ACL submission under ACL Rolling Review
  • SKILL.md covers What counts as an artifact here, Submission-time packaging rules, Checklist items your artifact… and What an ACL reviewer opens first, plus 5 more sections
  • Calls jupyter

What it does

Acl Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data statements, and post-acceptance public release.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Annotation materials for an ACL submission under ACL Rolling Review
  • Covering anonymized supplement archives
  • Scientific-artifact items of the Responsible NLP checklist
  • Licensing and intended-use documentation

Example prompts

  • “/acl-artifact-evaluation”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. The README — it has roughly one minute to orient them.
  2. Prompt files and evaluation scripts, for any LLM claim: exact prompts,
  3. Annotation guidelines, for any dataset or human-eval claim: reviewers judge
  4. A sample of the data itself — quality problems visible in twenty random

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jupyter

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Acl Artifact Evaluation loads about 1.5k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 617 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 617 words, ~1,452 tokens.

Download SKILL.mdSave it as .claude/skills/acl-artifact-evaluation/SKILL.md (or your agent's skills folder).
name
acl-artifact-evaluation
description
Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data statements, and post-acceptance public release.

ACL Artifact Evaluation

Use this to plan the evidence package around an ACL paper. ACL has no separate artifact-badge track; instead, artifact scrutiny is folded into review through the supplement archive and Section B ("scientific artifacts") of the Responsible NLP checklist, which reviewers cross-check against the PDF.

What counts as an artifact here

  • Code: training/inference scripts, evaluation harnesses, prompt templates.
  • Data: new corpora, annotations, filtered subsets of existing corpora, test suites, adversarial sets.
  • Model outputs: generations, ranked lists, logits used in analysis — often the cheapest way to make an LLM paper checkable without GPUs.
  • Human-subject materials: annotation guidelines, interface screenshots, consent text, compensation description.

Submission-time packaging rules

  • Supplements upload as .tgz/.zip through the OpenReview form; links to tracked cloud storage are not acceptable, and any linked page must be anonymous.
  • Scrub identity everywhere reviewers can look: file paths, git metadata, notebook author fields, license headers, dataset hosting pages, README contact lines.
  • Reviewers are not required to open supplements. The paper plus checklist must stand alone; the archive is for verification, not for essential content.

Checklist items your artifact must satisfy

Responsible NLP item (Section B)Artifact implication
Cited creators + versions of used artifactsPin dataset/model versions in the README and bibliography
License / terms of use statedInclude the license you release under and those you consumed under
Use consistent with intended useJustify research use of scraped or user-generated data
PII and offensive content handledDescribe scanning/anonymization steps actually performed
Documentation of domains, languages, demographicsShip a data statement or datasheet, not just row counts
Statistics on splits reportedTrain/dev/test sizes in both paper and README

Checklist answers contradicted by the archive read as misleading information — grounds for desk rejection under ARR policy, and a credibility wound even when not enforced.

What an ACL reviewer opens first

  1. The README — it has roughly one minute to orient them.
  2. Prompt files and evaluation scripts, for any LLM claim: exact prompts, decoding parameters, and scoring code are the reproduction spine.
  3. Annotation guidelines, for any dataset or human-eval claim: reviewers judge whether the labels could possibly mean what the paper says.
  4. A sample of the data itself — quality problems visible in twenty random examples have sunk otherwise strong resource papers.
Show full SKILL.md (247 more words)Show less

Vignette: packaging a multilingual benchmark submission

A hypothetical paper releases a 7-language reading-comprehension test suite built from news text plus a baseline evaluation of five LLMs.

  • Ship per-language provenance: source, license, collection window, and the filtering pipeline as runnable code, since "web text" alone fails checklist item B on documentation.
  • Include annotator guidelines, pay, recruitment channel, and agreement statistics; multilingual annotation quality is the first attack surface.
  • Provide the exact prompts and outputs for all five models so reviewers can re-score without API keys.
  • Keep a versioned, hash-stamped test file so post-publication contamination can be audited later.

Release ladder after acceptance

text
anonymous supplement  ->  public repo + dataset page  ->  archived, versioned release
   (review-time)          (camera-ready links)            (DOI/hub artifact, cited version)

Post-acceptance, register the artifact where your community actually looks (model/dataset hubs, a maintained repo), state the license explicitly, and put the citation-of-record (the Anthology entry) in the README.

Anonymization sweep, concretely

Run these before zipping, on a copy:

bash
# authorship trails in code and docs
grep -ri "yourname\|yourlab\|university" . --include="*.py" --include="*.md"
# git history and remotes leak owners
rm -rf .git; # or re-init a fresh repo for the archive copy
# notebook metadata carries usernames and kernel paths
jupyter nbconvert --clear-output --inplace *.ipynb
# absolute paths in configs and logs
grep -r "/home/\|/Users/" . | head

Then check the parts tools miss: license headers naming the lab, dataset hosting pages with institutional branding, model cards listing maintainers, and README badges pointing at owner-named CI.

Sizing and format sanity

  • Keep the archive lean: strip checkpoints reviewers cannot load anyway, cached datasets, and virtualenvs; describe big assets and provide them at camera-ready instead.
  • One top-level README, one environment file, one entry point per claimed result — reviewers grant roughly a minute before giving up.
  • Verify the .zip/.tgz opens on a machine that has never seen the project; OpenReview upload limits and accepted fields vary by cycle, so check the live form rather than last cycle's.

Output format

text
[Artifact role] anonymous supplement / camera-ready release / public benchmark
[Contents] <code/data/prompts/outputs/guidelines>
[Checklist alignment] <Section B items satisfied vs missing>
[Anonymity findings] <paths/metadata/hosting leaks>
[Release plan] <post-acceptance registry, license, versioning>

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in ACL-Skills/skills/acl-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Acl Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Acl Artifact Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Acl Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
OpenMed Model Card Writermaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 13 days ago
    Auto-check passed

Questions about Acl Artifact Evaluation

What does Acl Artifact Evaluation do?

A skill your agent uses when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives…. Acl Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging code, datasets, prompts, model outputs, or annotation materials for an ACL submission under ACL Rolling Review, covering anonymized supplement archives, scientific-artifact items of the Responsible NLP checklist, licensing and intended-use documentation, data statements, and post-acceptance public release.

When should I use Acl Artifact Evaluation?

Acl Artifact Evaluation fits situations like: annotation materials for an ACL submission under ACL Rolling Review; covering anonymized supplement archives; scientific-artifact items of the Responsible NLP checklist; licensing and intended-use documentation.

How do I install Acl Artifact Evaluation in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluation -a claude-code`. Or copy the skill folder (ACL-Skills/skills/acl-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/acl-artifact-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Acl Artifact Evaluation in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluation -a codex`. Or copy the skill folder (ACL-Skills/skills/acl-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/acl-artifact-evaluation in your project. Codex loads it when a task matches its description.

Can I use Acl Artifact Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill acl-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/acl-artifact-evaluation, .gemini/skills/acl-artifact-evaluation, .github/skills/acl-artifact-evaluation and .opencode/skills/acl-artifact-evaluation in your project.

What does Acl Artifact Evaluation need to run?

Going by SKILL.md and its folder, Acl Artifact Evaluation needs the command-line tools its instructions call (jupyter).

Does Acl Artifact Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Acl Artifact Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Acl Artifact Evaluation use?

Acl Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Acl Artifact Evaluation use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Acl Artifact Evaluation?

Skills that share tags, products or a category with Acl Artifact Evaluation: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and Andrej Karpathy (K-Dense-AI/mimeo, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Acl Artifact Evaluation?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.