Compare
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
A skill your agent uses when choosing how to tokenize text or which transformer type fits an NLP task, when a tokenizer over-fragments non-English text or inflates token cost, when picking a…
$ npx skills add ericrisco/rsc-harness --skill nlp -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness nlp --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nlp .claude/skills/nlp && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nlp" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/nlp into .claude/skills/nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/nlpType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill nlp -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness nlp --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nlp .agents/skills/nlp && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nlp" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/nlp into .agents/skills/nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill nlp -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness nlp --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nlp .cursor/skills/nlp && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nlp" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/nlp into .cursor/skills/nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/nlp--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill nlp -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness nlp --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nlp .gemini/skills/nlp && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nlp" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/nlp into .gemini/skills/nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness nlpInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill nlp -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nlp .github/skills/nlp && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nlp" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/nlp into .github/skills/nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill nlp -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness nlp --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nlp .opencode/skills/nlp && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nlp" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/nlp into .opencode/skills/nlp/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nlpA skill your agent uses when choosing how to tokenize text or which transformer type fits an NLP task, when a tokenizer over-fragments non-English text or inflates token cost, when picking a…
NLP is an agent skill from ericrisco/rsc-harness. Use when choosing how to tokenize text or which transformer type fits an NLP task, when a tokenizer over-fragments non-English text or inflates token cost, when picking a language metric, or when classification, NER or summarization output looks wrong and it is unclear whether the tokenizer, the architecture or the metric is at fault. Covers subword tokenizers, encoder versus decoder versus encoder-decoder choice, sentence embeddings, and the metric families. NOT retrieval or vector search (that is…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/evaluation.md`).
It sits in AI & LLM Engineering, covering Natural language processing, Embeddings and Summarization. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
NLP loads about 3.5k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 166 tokens; SKILL.md has 1,486 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,486 words, ~3,541 tokens.
.claude/skills/nlp/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.You own the language-modeling discipline: how raw text becomes tokens, which transformer architecture fits a task, and which metric actually tells you whether it worked. When the question is "which tokenizer," "BERT or GPT or T5 for this," "why does my Catalan text cost 3× the tokens," or "is this BLEU score meaningful," this is the skill. You stop at retrieval, the RAG loop, prompt wording, and the training step itself — those route out (below).
../embeddings-search/SKILL.md. Sentence embeddings live
here as a task; using them to retrieve is theirs.../rag/SKILL.md.../prompt-engineering/SKILL.md.../finetuning/SKILL.md
../training-data/SKILL.md.Pick the architecture from the task's shape, not from what is trendy. A decoder LLM can technically classify, but a fine-tuned encoder is smaller, faster, cheaper, and usually more accurate on a fixed-label task.
| Task shape | Architecture | Why | Example families* |
|---|---|---|---|
| Understand / label a whole input (classification, NER, extractive QA, similarity) | Encoder (bidirectional) | Attends to the full sentence both directions; cheap to fine-tune and to serve | BERT, RoBERTa, DistilBERT, ModernBERT |
| Free-form generation, chat, few-shot | Decoder (autoregressive) | Attends only to prior tokens; predicts the next token | GPT-style, Llama, Gemma, Qwen |
| Transform input → new text (summarize, translate, generative QA) | Encoder-decoder / seq2seq | Encoder reads all of the source, decoder writes conditioned on it | T5 / FLAN-T5, BART, mT5 |
* Architecture families are stable; specific checkpoints and their licenses are not — check the
HF model card before you commit (licenses change; see open-weights). ModernBERT (2024) is a
current long-context encoder; verify the latest at author time.
The two most common own-goals: reaching for a 7B decoder to do sentiment on 5 classes (an encoder does it for a fraction of the cost), and forcing an encoder to generate (it cannot — it has no decoder).
Every downstream number depends on this step, and its failures are silent. The single load-bearing rule:
Load the tokenizer that shipped with the checkpoint, and use the same one at train and inference.
AutoTokenizer.from_pretrained(same_checkpoint). A train/inference tokenizer mismatch — different vocab, different special tokens, different casing/normalization — maps text to token ids the model never saw and corrupts everything downstream with no error.
from transformers import AutoTokenizer # transformers current major ~v5 (verify at author time)
tok = AutoTokenizer.from_pretrained("bert-base-cased")
enc = tok("Tokenizers matter.", return_offsets_mapping=True)
tok.convert_ids_to_tokens(enc["input_ids"])
# ['[CLS]', 'Token', '##izers', 'matter', '.', '[SEP]'] — note WordPiece '##' continuation + added specialsThe four algorithms (full mechanics in references/tokenization.md):
| Algorithm | Builds vocab by… | Applies by… | Used by |
|---|---|---|---|
| BPE | merging the most frequent adjacent pair, repeatedly | split to chars, replay learned merges | GPT-2 (byte-level), many |
| WordPiece | merging pairs that maximize a likelihood score | longest-match subword from the front (## continuations) | BERT family |
| Unigram (SentencePiece) | start large, remove tokens that least hurt corpus likelihood | most-probable segmentation | T5, ALBERT, mT5 |
| Byte-level BPE | BPE over the 256 raw bytes, not Unicode chars | same as BPE on bytes | GPT-2, RoBERTa |
[UNK]. Base vocab is exactly 256 (all byte values), so every
emoji, accent, and script maps to some byte sequence — nothing falls out as unknown
(verified: HF NLP course ch.6). WordPiece/word-level tokenizers do have [UNK] and lose OOV
content.▁ meta-symbol), so
decode(encode(x)) == x without language-specific detokenization rules. That is why it
dominates multilingual models.Special tokens are not decoration. [CLS]/<s> carries the pooled sentence
representation for classification; [SEP]/</s> marks segment/end; [PAD] fills a batch (and
must be masked out via attention_mask); [MASK] is the MLM target; [UNK] is the fallback.
Names differ by model ([CLS] in BERT vs <s> in RoBERTa) — another reason to never hand-roll
the tokenizer.
Why it matters — three concrete costs:
[UNK] throws away content it can't
represent; byte-level/SentencePiece degrade gracefully instead.from transformers import pipeline
clf = pipeline("text-classification", model="distilbert-base-uncased-finetuned-sst-2-english")
clf("The service was slow but the food was incredible.")
# [{'label': 'POSITIVE', 'score': 0.99...}]Metric: accuracy on balanced data; macro-F1 the moment classes are imbalanced (accuracy lies when 95% of rows are one class).
Labels are B-/I-/O spans aligned to subword tokens: the first subword of a word gets the
label, continuation subwords and special tokens get -100 (ignored by the loss). Evaluate with
seqeval at the entity level, never per-token accuracy (per-token accuracy is inflated by
the flood of O tokens).
from transformers import pipeline
ner = pipeline("token-classification", aggregation_strategy="simple")
ner("Ada Lovelace worked in London.")
# groups subwords back into entities: PER 'Ada Lovelace', LOC 'London'summ = pipeline("summarization", model="facebook/bart-large-cnn")
summ(long_article, max_length=130, min_length=30)Metric: ROUGE for summarization, BLEU/chrF for translation — with the heavy caveat in section 4.
from sentence_transformers import SentenceTransformer # sentence-transformers ~v5 (verify)
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
emb = model.encode(["The weather is lovely today.", "It's so sunny outside!"])
model.similarity(emb, emb) # semantic textual similarity / clustering / paraphrase miningProducing/judging embeddings for retrieval (model choice, chunking, recall@k, rerankers) is
embeddings-search, not here.
| Task | Primary metric | Catches | Trap |
|---|---|---|---|
| Classification | accuracy + macro-F1 | wrong labels | accuracy hides minority-class failure |
| NER / token | entity-level F1 (seqeval) | missed/partial spans | per-token accuracy is inflated by O |
| Translation | BLEU / chrF | n-gram overlap w/ reference | weak on meaning; chrF better for morphology |
| Summarization | ROUGE (1/2/L) | recall of reference n-grams | rewards copying; blind to faithfulness |
| Generation (LM) | perplexity | how well the model predicts held-out text | tokenizer-dependent — not comparable across tokenizers |
| Open-ended / chat | LLM-as-judge + human | quality overlap metrics miss | judge bias (position, verbosity, self-preference) |
The caveat that governs this whole section: BLEU, ROUGE, and chrF are n-gram/character overlap metrics and correlate weakly with human judgment on open-ended and creative generation — they reward matching the reference's exact phrasing, so a correct paraphrase scores low and a fluent-but-wrong copy scores high (well documented; e.g. the summarization and MT-evaluation literature). Use them for regression tracking on a fixed reference set, never as the final verdict on quality. For open-ended output, use an LLM-as-judge rubric plus a human spot-check — and know the judge has its own biases (position, verbosity, self-preference), so pin the rubric and randomize order.
Perplexity = exp(mean token NLL): lower means the model predicts held-out text better. It is
tokenizer-dependent, so two models with different tokenizers have non-comparable perplexities
— only compare within the same tokenizer/vocab. Runnable snippets for seqeval, sacrebleu, ROUGE,
perplexity, and an LLM-judge harness are in references/evaluation.md.
The English-centric trap: a tokenizer whose vocab was learned mostly on English over-fragments other scripts. The same sentence in Ukrainian, Arabic, Hindi, or even accented Catalan can take 2–15× more tokens than its English equivalent (Petrov et al., Language Model Tokenizers Introduce Unfairness Between Languages, NeurIPS 2023). That "fertility" (tokens per word) inflation is a triple tax:
Mitigations: prefer a multilingual tokenizer/model (mT5, XLM-R, a SentencePiece-based model) whose vocab actually covers your languages; measure fertility on your own corpus (tokens per word, per language) before you commit; and don't benchmark cost or latency only on English.
O tokens inflate accuracy toward 1.0.embeddings-search — retrieval embeddings, chunking, recall@k,
reranking. NLP owns making/judging sentence embeddings as a task; using them to search is theirs.rag — the full retrieve→generate answer loop and groundedness.prompt-engineering — the wording of the prompt.finetuning + deep-learning — actually
training/adapting the network (this skill picks the type and metric; those move the weights).training-data — building the labeled corpus you train on.attention_mask handled (padding masked, -100 on ignored labels).references/tokenization.md — BPE/WordPiece/Unigram/byte-level training mechanics, special
tokens per family, offset mapping, and a fertility-measuring snippet.references/evaluation.md — runnable seqeval, sacrebleu (BLEU/chrF), ROUGE, perplexity, and an
LLM-as-judge rubric, with when each lies.© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/nlp of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
NLP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| NLP this skillericrisco/rsc-harness | 156 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Comparetaishi-i/awesome-japanese-nlp-resources | 1k | — | ~4.1k | Automated safety check: Notes | CC0-1.0 | |
| Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Researchtaishi-i/awesome-japanese-nlp-resources | 1k | — | ~3.5k | Automated safety check: Notes | CC0-1.0 | |
| Searchtaishi-i/awesome-japanese-nlp-resources | 1k | — | ~4.3k | Automated safety check: Notes | CC0-1.0 | |
| Transformersynulihao/AgentSkillOS | 617 | — | ~2.9k | Automated safety check: Pass | None |
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
Orchestra-Research/AI-Research-SKILLs
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
taishi-i/awesome-japanese-nlp-resources
Analyze current trends and challenges in Japanese NLP for a topic.
taishi-i/awesome-japanese-nlp-resources
Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).
ynulihao/AgentSkillOS
Work with state-of-the-art machine learning models for NLP, computer vision, audio, and multimodal tasks using HuggingFace Transformers.
joshzyj/open-scholar-skill
Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Categories
A skill your agent uses when choosing how to tokenize text or which transformer type fits an NLP task, when a tokenizer over-fragments non-English text or inflates token cost, when picking a…. NLP is an agent skill from ericrisco/rsc-harness. Use when choosing how to tokenize text or which transformer type fits an NLP task, when a tokenizer over-fragments non-English text or inflates token cost, when picking a language metric, or when classification, NER or summarization output looks wrong and it is unclear whether the tokenizer, the architecture or the metric is at fault.
NLP fits situations like: choosing how to tokenize text; which transformer type fits an NLP task; A tokenizer over-fragments non-English text; inflates token cost.
Run `npx skills add ericrisco/rsc-harness --skill nlp -a claude-code`. Or copy the skill folder (skills/nlp in ericrisco/rsc-harness) into .claude/skills/nlp in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill nlp -a codex`. Or copy the skill folder (skills/nlp in ericrisco/rsc-harness) into .agents/skills/nlp in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill nlp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nlp, .gemini/skills/nlp, .github/skills/nlp and .opencode/skills/nlp in your project.
SKILL.md names no scripts, command-line tools or credentials: NLP is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
NLP is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with NLP: Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars), Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Research (taishi-i/awesome-japanese-nlp-resources, 1k stars) and Search (taishi-i/awesome-japanese-nlp-resources, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.