Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).

CC0-1.0Auto-check: notesAI & LLM Engineering

Install Search

skills CLI
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install taishi-i/awesome-japanese-nlp-resources search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/search .claude/skills/search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
search
GitHub stars
1k
Token cost
~4.3k tokens
SKILL.md length
1,447 words
Files
2
Skills in repo
4
Repo updated
First seen
Licence
CC0-1.0

At a glance

Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).

  • Works in 7 steps: Validate input → Interpret the query → Locate the data file → …
  • The user asks which Japanese NLP resource to use
  • SKILL.md covers Claude Code and Codex and Instructions
  • Calls python

What it does

Search is an agent skill from taishi-i/awesome-japanese-nlp-resources. Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face). Accepts keywords or natural language questions in any language. Use whenever the user asks which Japanese NLP resource to use, or wants to find one: tokenizers / morphological analyzers, BERT or LLM models, embeddings, NER, text classification, datasets / corpora, dictionaries, tutorials, or Hugging Face models. Trigger phrases include '日本語の形態素解析ライブラリ', 'おすすめの日本語tokenizer', '日本語BERTモデル', '日本語の感情分析データセット'…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering Natural language processing, Embeddings and Model hubs and datasets. It works with Hugging Face. The repository describes itself as: A curated list of resources for Japanese natural language processing (NLP): Python libraries, LLMs, dictionaries, corpora, and datasets. Includes Claude Code and Codex skills to… The licence is CC0-1.0.

When your agent uses it

  • The user asks which Japanese NLP resource to use
  • Wants to find one: tokenizers / morphological analyzers
  • Text classification
  • Datasets / corpora

Example prompts

  • “日本語の形態素解析ライブラリ”
  • “おすすめの日本語tokenizer”
  • “日本語BERTモデル”
  • “/search”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Validate input
  2. Interpret the query
  3. Locate the data file
  4. Search and score with Python
  5. Re-rank with your judgment
  6. Format the output
  7. Output use-case selection guide table

What it can do on your machine

Read from SKILL.md and the folder at commit 451770b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Search loads about 4.3k tokens when it runs. Until then it costs about 148 tokens; SKILL.md has 1,447 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~148
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from taishi-i/awesome-japanese-nlp-resources at commit 451770b, republished under its CC0-1.0 licence (© taishi-i). 1,447 words, ~4,301 tokens.

Download SKILL.mdSave it as .claude/skills/search/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
search
description
Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face). Accepts keywords or natural language questions in any language. Use whenever the user asks which Japanese NLP resource to use, or wants to find one: tokenizers / morphological analyzers, BERT or LLM models, embeddings, NER, text classification, datasets / corpora, dictionaries, tutorials, or Hugging Face models. Trigger phrases include '日本語の形態素解析ライブラリ', 'おすすめの日本語tokenizer', '日本語BERTモデル', '日本語の感情分析データセット', '日本語LLM 一覧', 'which Japanese embedding model', 'Japanese NER library'.
allowed-tools
Bash
argument-hint
query

Search the awesome-japanese-nlp-resources database for the user's query.

Claude Code and Codex

This skill is shared by the Claude Code and Codex versions of the plugin. The steps are the same in both tools; only these details differ:

  • Query — Claude Code: the arguments of /awesome-japanese-nlp-resources:search, appended at the end of this skill as ARGUMENTS: …. Codex: the user's message that invoked $awesome-japanese-nlp-resources:search, minus that $… mention. If the skill was picked automatically rather than invoked by name, use the user's request as the query.
  • Plugin root — Claude Code: ${CLAUDE_PLUGIN_ROOT}. Codex: the directory two levels above this SKILL.md (use its absolute path).
  • Shell — run the commands below with Claude Code's Bash tool or Codex's shell tool. Copy each Python script in full and run it as written, changing only its placeholders (RESOURCES_PATH, the keyword lists) — don't shorten it, drop passes, or alter its scores and thresholds.
  • Commands — write any command you show the user in the current tool's form: /awesome-japanese-nlp-resources:<skill> in Claude Code, $awesome-japanese-nlp-resources:<skill> in Codex.

Results must come from the bundled data. If the data file can't be read (for example, shell commands are blocked or fail to start), say so and link https://github.com/taishi-i/awesome-japanese-nlp-resources instead of answering from memory or web search.

Instructions

Step 0 — Validate input

If the query is empty or blank, stop immediately and output (in Codex, write the commands with $ instead of /):

Usage: /awesome-japanese-nlp-resources:search <query>

Examples:
  /awesome-japanese-nlp-resources:search morphological analysis
  /awesome-japanese-nlp-resources:search BERT
  /awesome-japanese-nlp-resources:search named entity recognition
  /awesome-japanese-nlp-resources:search text classification dataset
  /awesome-japanese-nlp-resources:search sentence embedding

Please pass the keyword(s) you want to search for as the argument.

---

使い方: /awesome-japanese-nlp-resources:search <query>

クエリ例:
  /awesome-japanese-nlp-resources:search 形態素解析
  /awesome-japanese-nlp-resources:search BERT
  /awesome-japanese-nlp-resources:search 固有表現認識
  /awesome-japanese-nlp-resources:search テキスト分類 データセット
  /awesome-japanese-nlp-resources:search 文埋め込み

検索したいキーワードを引数に指定してください。

Do not proceed to Step 1 if the query is empty.

Step 1 — Interpret the query

The data descriptions are in English, so always convert the query intent to English keywords before searching.

Keyword rules — read before choosing keywords:

  1. Use stems, not full words. Substring match is used, so morpholog catches "morphology", "morphological", "morphological analyzer". Other examples: embed → embedding/embeddings, classif → classification/classifier, translat → translation/translate, generat → generation/generative, segment → segmentation/segmenter, recogni → recognition/recognizer, extract → extraction/extractor, retriev → retrieval/retrieve.
  2. Add domain-specific tool names. When the query maps to a known NLP domain, include the well-known tool names present in the database:
Domain (Japanese query hint)Stem keywordsTool names to add
形態素解析 / morphological analysismorpholog, segmentmecab, janome, sudachi, kytea, kuromoji, jumanpp, nagisa
固有表現認識 / NERnamed entit, NER, recogniginza, spacy, knp
係り受け解析 / dependency parsingdepend, parse, syntaxcabocha, knp, ginza, spacy
文章分類 / text classificationclassif, sentiment, categorbert, fasttext
感情分析 / sentiment analysissentiment, emotion, opinionoseti, wrime
埋め込み / word vectors / embeddingsembed, vector, representword2vec, fasttext, bert, sbert
事前学習モデル / pretrained modelpretrain, language model, bert, gptbert, gpt, llama, rinna, elyza, calm, swallow
テキスト生成 / text generationgenerat, language modelgpt, llm, llama, rinna, elyza
機械翻訳 / machine translationtranslat, machine translationopus, marian, fairseq
音声認識 / speech recognitionspeech, recogni, audio, asrwhisper, julius, espnet
音声合成 / text-to-speechspeech, synthesis, ttsvoicevox, espnet
質問応答 / QAquestion, answer, qabert, t5
要約 / summarizationsummari, abstractbart, t5, pegasus
辞書・IME / dictionarydict, lexicon, imemecab, sudachi, mozc
コーパス・データセット / corpuscorpus, dataset, annot(rely on stems)
チュートリアル / learningtutorial, introduc, learn(rely on stems)
OCR / 光学文字認識ocr, optical character, recognimanga-ocr, donut, tesseract
RAG / 検索拡張生成retriev, rag, embedruri, glucose, faiss
ファインチューニング / fine-tuningfine-tun, finetun, lora, peftlora, peft, qlora
ベンチマーク・評価 / benchmarkbenchmark, evaluat, jgluellm-jp-eval, jglue, nejumi
  1. When the query contains Japanese text, also keep 2–4 raw Japanese terms/phrases lifted directly from the query (not translated) as a separate ja_keywords list. Aliases and some descriptions (al, d_ja — see Step 3) are Japanese-only, so a literal Japanese substring catches entries an English-only translation would miss entirely — nicknames like ボイボ (VOICEVOX), めかぶ (mecab), or a Japanese technical term that never got glossed into the English description. Leave ja_keywords empty for English queries.
  2. Aim for 4–6 keywords. Fewer miss items; more than 6 inflates low-quality partial matches.
  3. If none of the above domains fit, translate the query intent literally to English stems.
Step 2 — Locate the data file

The data file ships with the plugin at data/resources.json under the plugin root (see "Claude Code and Codex" above). Resolve its absolute path, falling back to a scoped search only if the install is unusual:

bash
PLUGIN_ROOT="${CLAUDE_PLUGIN_ROOT}"  # Codex: replace with the plugin root, two levels above this SKILL.md
RESOURCES_PATH="$PLUGIN_ROOT/data/resources.json"
[ -f "$RESOURCES_PATH" ] || RESOURCES_PATH="$(find "${CODEX_HOME:-$HOME/.codex}/plugins" "${HOME}/.claude/plugins" -type f -name resources.json 2>/dev/null | grep "awesome-japanese-nlp-resources/" | head -1)"
echo "RESOURCES_PATH=$RESOURCES_PATH"

Use the resulting absolute RESOURCES_PATH wherever Step 3 opens the data file — write the path itself into the script, since shell variables may not persist between commands.

The plugin also ships data/multilingual_resources.json (same item format) listing multilingual libraries, models, and datasets (GitHub repositories) that also support Japanese, from docs/multilingual.md. The scripts below load it automatically when it exists; its items have categories like Multilingual (Speech recognition).

Step 3 — Search and score with Python

Do not read the data file directly (no Read tool, cat, head, or similar) — it is about 660 KB and would flood the context. Instead, run the scoring in a single shell command using Python.

Each item in the JSON array has:

  • u: GitHub or Hugging Face URL
  • n: repository/model name
  • d: English description
  • d_ja: Japanese description (GitHub-origin items only; match your ja_keywords against this)
  • al: curated alternate names / kana nicknames, e.g. ["VOICEVOX", "ボイスボックス", "ボイボ"] (array of strings, only ~40 items have this — treat a hit here as strong as a name match)
  • c: category (e.g. Python library, HuggingFace Model (Text Generation), Corpus, Tutorial, Multilingual (Speech recognition), ...)
  • s: subcategory / semantic labels (array of strings)
  • st: GitHub star count (GitHub items only; absent or 0 otherwise)
  • ns: normalized star score 0–10 (log-scaled, GitHub items only)
  • dl: Hugging Face download count (HF items only; absent or 0 otherwise)
  • nd: normalized download score 0–10 (log-scaled, HF items only)
  • sc: pre-computed quality score (higher = more popular/active)
  • status: "ok" or "not_found" — items whose repo 404s (~8 of ~1200) are filtered out below; never recommend one

Run the following, substituting RESOURCES_PATH with the absolute path from Step 2, keywords with your English keywords and ja_keywords with your raw Japanese terms, both from Step 1 (ja_keywords may be []):

python
python3 << 'EOF'
import json, os

with open("RESOURCES_PATH") as f:    # absolute path from Step 2
    data = json.load(f)
multilingual_path = os.path.join(os.path.dirname("RESOURCES_PATH"), "multilingual_resources.json")
if os.path.exists(multilingual_path):
    with open(multilingual_path) as f:
        data += json.load(f)

keywords = ["keyword1", "keyword2", "keyword3"]  # English stems, from Step 1
ja_keywords = []  # raw Japanese terms from Step 1 -- [] for English queries

results = []
for item in data:
    if item.get("status") == "not_found":
        continue  # dead repo -- never recommend it

    n = item.get("n", "").lower()
    d = item.get("d", "").lower()
    d_ja = item.get("d_ja") or ""
    s = " ".join(item.get("s") or []).lower()
    c = item.get("c", "").lower()
    al = " ".join(item.get("al") or []).lower()

    text_score = 0
    for kw in keywords:
        kw = kw.lower()
        if n == kw:       text_score += 20
        elif kw in n:     text_score += 10
        if kw in d:       text_score += 5
        if kw in s:       text_score += 3
        if kw in c:       text_score += 2
        if kw in al:      text_score += 10  # alias hit is name-equivalent

    for kw in ja_keywords:
        if kw in n:       text_score += 10
        if kw in d_ja:    text_score += 5
        if kw in al:      text_score += 10

    if text_score < 8:
        continue

    ns = item.get("ns") or 0
    nd = item.get("nd") or 0
    sc = item.get("sc") or 0
    pop = (ns if ns else nd) * 2.5
    qual = min(5, sc * 5 / 21)
    combined = text_score + pop + qual

    results.append((combined, text_score, item))

results.sort(key=lambda x: -x[0])
seen = {item['n'] for _, _, item in results}

# Supplemental pass: surface high-popularity items from matching categories
# that may have been missed because their descriptions are in Japanese.
# Keys are stems to match against user keywords; values are category prefixes
# (prefix match covers "HuggingFace Model (Text Generation)" etc.).
CATEGORY_KEYWORDS = {
    "tutorial": "Tutorial", "introduc": "Tutorial", "learn": "Tutorial",
    "morpholog": "Python library", "segment": "Python library",
    "mecab": "Python library", "janome": "Python library", "sudachi": "Python library",
    "spacy": "Python library", "ginza": "Python library",
    "corpus": "Corpus", "dataset": "Corpus",
    "bert": "HuggingFace Model", "gpt": "HuggingFace Model",
    "llm": "HuggingFace Model", "llama": "HuggingFace Model",
    "pretrain": "HuggingFace Model", "embed": "HuggingFace Model",
    "model": "Pretrained model",
}
supplement_cats = set()
for kw in keywords:
    for ck, cat in CATEGORY_KEYWORDS.items():
        if ck in kw.lower():
            supplement_cats.add(cat)

if supplement_cats:
    def cat_match(c):
        return any(c == cat or c.startswith(cat + " ") for cat in supplement_cats)
    extras = [
        item for item in data
        if cat_match(item.get("c", ""))
        and (item.get("st", 0) or item.get("dl", 0))
        and item["n"] not in seen
        and item.get("status") != "not_found"
    ]
    extras.sort(key=lambda x: -max(x.get("ns") or 0, x.get("nd") or 0))
    for item in extras[:5]:
        ns = item.get("ns") or 0
        nd = item.get("nd") or 0
        sc = item.get("sc") or 0
        # base 8 = category-match credit (same as the text_score threshold)
        combined = 8 + max(ns, nd) * 2.5 + min(5, sc * 5 / 21)
        results.append((combined, 0, item))
        seen.add(item["n"])

results.sort(key=lambda x: -x[0])
for combined, text_score, item in results[:20]:
    st = item.get("st", 0) or 0
    dl = item.get("dl", 0) or 0
    flag = " [supplemental]" if text_score == 0 else ""
    print(f"score={combined:.1f} text={text_score} st={st} dl={dl}{flag}")
    print(f"  n={item['n']}")
    print(f"  u={item['u']}")
    print(f"  c={item['c']}")
    print(f"  s={item.get('s','')}")
    if item.get('al'):
        print(f"  al={item['al']}")
    print(f"  d={item.get('d','')[:120]}")
    if item.get('d_ja'):
        print(f"  d_ja={item['d_ja'][:120]}")
    print()
EOF

This returns up to 20 candidates. Items marked [supplemental] were added by the category-based pass to recover high-star resources whose descriptions are in Japanese. In Step 4, evaluate supplemental items on semantic fit before including them in the final list.

Show full SKILL.md (464 more words)Show less
Step 4 — Re-rank with your judgment

You now have up to 20 candidates. Apply your semantic judgment to produce the final ordered list of up to 10 results.

Re-rank by evaluating each candidate on:

  1. Semantic centrality — how directly does this resource address the query's core intent? A BERT model is more central to "BERT fine-tuning" than a generic transformer library.
  2. Popularity as a proxy for quality — high stars/downloads generally signal battle-tested, well-documented tools. Prefer them when candidates are otherwise equivalent.
  3. Category fit — match the resource type to the implied need:
    • "how to learn / 勉強" → prefer Tutorial, Research summary
    • "I need a model" → prefer Pretrained model, HuggingFace Model
    • "find a dataset / コーパス" → prefer Corpus, HuggingFace Dataset
    • "build an app / ライブラリ" → prefer Python library, language-specific libs
    • "multilingual / 多言語 / other languages too" → include Multilingual (...) items; otherwise prefer Japanese-specific resources when they fit equally well, and use Multilingual (...) items to fill gaps such as speech, OCR, language detection or search engines
  4. Specificity — a resource specialized for the exact task beats a general one.
  5. Recency signal — when sc is significantly higher among otherwise-similar items, it usually reflects more recent activity; prefer those.

Do not mechanically follow the combined score from Step 3 — use it as a starting point, then move items up or down based on the criteria above.

Step 5 — Format the output

Language detection rule (apply before writing any output):

  • The query contains Japanese characters (hiragana / katakana / kanji) → Japanese
  • Otherwise → English (default)

Apply the detected language to all headings and prose.

Present the final re-ranked results:

## Search results for "<query>"

*(Searched for: keyword1, keyword2, ...)*

Found N result(s).

### 1. [repository-name](url)
**Category:** category > subcategory
**Popularity:** ⭐ {st} stars  (or  📥 {dl} downloads for HF)
Description text here.

### 2. ...

If no results, suggest alternate keywords and link to: https://github.com/taishi-i/awesome-japanese-nlp-resources

Step 6 — Output use-case selection guide table

After the search results list, append a guide table that helps the user pick the right resource for their specific situation.

Match the section heading and table language to the query language — translate the heading and column headers into the query language (e.g. Japanese query → Japanese heading and headers).

## Use-case Selection Guide

| Use case | Recommended | Popularity | Why |
|---|---|---|---|
| ... | [name](url) | ⭐N or 📥N | short reason |

Rules:

  • List 3–6 distinct use cases derived from the top 10 results. Each row should represent a meaningfully different scenario (e.g., "fine-tune an LLM" vs "evaluate an LLM"), not just a restatement of the search query.
  • For each row, select the single best resource from the top 10 results.
  • Popularity column: use ⭐{st} for GitHub stars, 📥{dl} for HuggingFace downloads. If both are 0, omit.
  • Why: write a 10–15 word reason in the query language explaining why this resource is the best fit for that use case. Do not copy the description verbatim. Focus on the practical benefit.
  • If two use cases would map to the same resource, merge them into one row or drop the weaker one.
  • If there are fewer than 3 meaningfully distinct use cases in the results, output as many rows as make sense (minimum 1).

© taishi-i, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/awesome-japanese-nlp-resources/skills/search of taishi-i/awesome-japanese-nlp-resources.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 451770b

Compare with similar skills

Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Search this skilltaishi-i/awesome-japanese-nlp-resources1k—~4.3kAutomated safety check: NotesCC0-1.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.0
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k2 repos~1.6kAutomated safety check: PassMIT
Hugging Face Transformers Usagedavila7/claude-code-templates33k11 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    33k GitHub starsUsed in 11 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface Inference Guide

    wentorai/research-plugins

    Run NLP and CV model inference via Hugging Face free-tier API

    298 GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed

More from taishi-i/awesome-japanese-nlp-resources

  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 4 days ago
    Auto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check: notes
  • Discover

    taishi-i/awesome-japanese-nlp-resources

    Given a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed…

    1k GitHub stars~6.5k tokensUpdated 4 days ago
    Auto-check: notes

Works with

Questions about Search

What does Search do?

Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face). Search is an agent skill from taishi-i/awesome-japanese-nlp-resources. Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).

When should I use Search?

Search fits situations like: the user asks which Japanese NLP resource to use; wants to find one: tokenizers / morphological analyzers; text classification; datasets / corpora.

How do I install Search in Claude Code?

Run `npx skills add taishi-i/awesome-japanese-nlp-resources --skill search -a claude-code`. Or copy the skill folder (plugins/awesome-japanese-nlp-resources/skills/search in taishi-i/awesome-japanese-nlp-resources) into .claude/skills/search in your project. Claude Code loads it when a task matches its description.

How do I install Search in Codex?

Run `npx skills add taishi-i/awesome-japanese-nlp-resources --skill search -a codex`. Or copy the skill folder (plugins/awesome-japanese-nlp-resources/skills/search in taishi-i/awesome-japanese-nlp-resources) into .agents/skills/search in your project. Codex loads it when a task matches its description.

Can I use Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add taishi-i/awesome-japanese-nlp-resources --skill search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/search, .gemini/skills/search, .github/skills/search and .opencode/skills/search in your project.

What does Search need to run?

Going by SKILL.md and its folder, Search needs the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash.

Does Search access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Search safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Search use?

Search is published under the CC0-1.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Search use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Search?

Skills that share tags, products or a category with Search: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Esmfold2 (JimLiu/science-skills, 228 stars), Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Search?

taishi-i (a GitHub user) maintains it in taishi-i/awesome-japanese-nlp-resources, which has 1,021 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 6, 2026.

Source: taishi-i/awesome-japanese-nlp-resources on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.