Hugging Face Tokenizers
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
Agent skill
by brycewang-stanford in brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact…
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .claude/skills/naacl-artifact-evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "naacl-artifact-evaluation" agent skill from https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluation into .claude/skills/naacl-artifact-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "naacl-artifact-evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .agents/skills/naacl-artifact-evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "naacl-artifact-evaluation" agent skill from https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluation into .agents/skills/naacl-artifact-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "naacl-artifact-evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .cursor/skills/naacl-artifact-evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "naacl-artifact-evaluation" agent skill from https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluation into .cursor/skills/naacl-artifact-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "naacl-artifact-evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/brycewang-stanford/Awesome-Journal-Skills.git --path NAACL-Skills/skills/naacl-artifact-evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .gemini/skills/naacl-artifact-evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "naacl-artifact-evaluation" agent skill from https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluation into .gemini/skills/naacl-artifact-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "naacl-artifact-evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .github/skills/naacl-artifact-evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "naacl-artifact-evaluation" agent skill from https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluation into .github/skills/naacl-artifact-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "naacl-artifact-evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install brycewang-stanford/Awesome-Journal-Skills naacl-artifact-evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/NAACL-Skills/skills/naacl-artifact-evaluation .opencode/skills/naacl-artifact-evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "naacl-artifact-evaluation" agent skill from https://github.com/brycewang-stanford/Awesome-Journal-Skills/tree/main/NAACL-Skills/skills/naacl-artifact-evaluation into .opencode/skills/naacl-artifact-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "naacl-artifact-evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
naacl-artifact-evaluationA skill your agent uses when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact…
Naacl Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact questions, documenting provenance and licensing for language data, and handling community-owned or Indigenous-language resources correctly.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.
Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Naacl Artifact Evaluation loads about 1.5k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 622 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 622 words, ~1,462 tokens.
.claude/skills/naacl-artifact-evaluation/SKILL.md (or your agent's skills folder).NAACL has no separate artifact-badging track; artifacts are judged inside the ARR review itself, through the supplement upload and section B of the Responsible NLP checklist ("scientific artifacts"). That placement matters: your artifact documentation is not an optional extra but a set of sworn answers reviewers cross-examine against the PDF.
| Data situation | Minimum documentation for a NAACL reviewer | Extra step |
|---|---|---|
| Standard public benchmark | Version, split, license, citation | Note any known contamination reports |
| Web-scraped text | Collection dates, filtering rules, deduplication, license basis | PII handling statement |
| New annotated corpus | Guidelines, annotator recruitment and pay, agreement scores | Release the guidelines verbatim in the supplement |
| Dialectal / code-switched data | Variety labels and how they were assigned | Native-speaker validation description |
| Indigenous or community-owned language data | Consent and partnership terms, community approval for release | Verify whether public release is permitted at all |
The last row is a NAACL signature concern. Work on languages of the Americas increasingly follows community-controlled data norms: some corpora may be used but not redistributed, some require named attribution (which conflicts with anonymous review — use a placeholder and restore at camera-ready), and some communities set conditions on derived models. "We release everything" is not automatically the ethical high ground here; the checklist rewards accuracy about constraints, not maximal openness.
artifact.zip
├── README.md # one-screen orientation: what, how, how long
├── data/
│ ├── data_card.md # provenance, license, varieties, gaps
│ └── samples/ # enough rows to judge quality, not the corpus
├── prompts/ # exact strings, all variants tried
├── eval/
│ ├── score.py # runs on outputs/ with no network access
│ └── outputs/ # raw model outputs backing the main tables
└── annotation/
└── guidelines.pdf # the instrument, scrubbed of institution marksScrub before zipping: repository history, notebook execution metadata, absolute paths with usernames, license headers naming the lab, and any consent form carrying institutional letterhead (replace with a redacted copy; note that the original exists).
A submission introduces a 40k-pair Quechua-Spanish parallel corpus built with two community organizations, plus MT baselines. The packaging calls that follow from this skill:
The result is an artifact that scores as honest and inspectable rather than incomplete, which is the realistic best outcome for community-governed data.
Archive size caps, accepted formats, and whether supplements upload as one file or several are OpenReview-form details that shift between ARR cycles; read the live submission form before building the final zip, and never reverse-engineer the limits from a previous cycle's folklore.
At camera-ready, the anonymous bundle becomes the public record: move it to a persistent host with a DOI or a tagged release, apply the real license, restore attribution the community partnership requires, and update the checklist-facing statements in the paper if the release scope changed between review and publication.
[Artifact inventory] <data / code / prompts / guidelines / outputs>
[Checklist B alignment] <each B answer -> where the artifact proves it>
[Provenance gaps] <unlicensed, undocumented, or unclear-consent items>
[Community constraints] <redistribution / attribution / approval terms>
[Anonymity sweep] clean / issues found
[Release plan] <anonymous now -> public form at camera-ready>© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in NAACL-Skills/skills/naacl-artifact-evaluation of brycewang-stanford/Awesome-Journal-Skills.
Open the folder on GitHubat commit 932eb23
Naacl Artifact Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Naacl Artifact Evaluation this skillbrycewang-stanford/Awesome-Journal-Skills | 1.2k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.4k | Automated safety check: Pass | MIT | |
| OpenMed Model Card Writermaziyarpanahi/openmed | 5.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel | 1.3k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Andrej KarpathyK-Dense-AI/mimeo | 282 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Comparetaishi-i/awesome-japanese-nlp-resources | 1k | — | ~4.1k | Automated safety check: Notes | CC0-1.0 |
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
maziyarpanahi/openmed
Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
K-Dense-AI/mimeo
Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
taishi-i/awesome-japanese-nlp-resources
Analyze current trends and challenges in Japanese NLP for a topic.
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…
Categories
A skill your agent uses when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact…. Naacl Artifact Evaluation is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact questions, documenting provenance and licensing for language data, and handling community-owned or Indigenous-language resources correctly.
Naacl Artifact Evaluation fits situations like: packaging datasets; annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklists artifact questions; documenting provenance and licensing for language data; handling community-owned.
Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a claude-code`. Or copy the skill folder (NAACL-Skills/skills/naacl-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/naacl-artifact-evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a codex`. Or copy the skill folder (NAACL-Skills/skills/naacl-artifact-evaluation in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/naacl-artifact-evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill naacl-artifact-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/naacl-artifact-evaluation, .gemini/skills/naacl-artifact-evaluation, .github/skills/naacl-artifact-evaluation and .opencode/skills/naacl-artifact-evaluation in your project.
SKILL.md names no scripts, command-line tools or credentials: Naacl Artifact Evaluation is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Naacl Artifact Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Naacl Artifact Evaluation: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and Andrej Karpathy (K-Dense-AI/mimeo, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,231 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.
Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.