Esmfold2
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
Given a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed…
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install taishi-i/awesome-japanese-nlp-resources discover --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/discover .claude/skills/discover && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "discover" agent skill from https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discover into .claude/skills/discover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "discover", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discoverType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install taishi-i/awesome-japanese-nlp-resources discover --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/discover .agents/skills/discover && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "discover" agent skill from https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discover into .agents/skills/discover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "discover", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install taishi-i/awesome-japanese-nlp-resources discover --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/discover .cursor/skills/discover && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "discover" agent skill from https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discover into .cursor/skills/discover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "discover", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/taishi-i/awesome-japanese-nlp-resources.git --path plugins/awesome-japanese-nlp-resources/skills/discover--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install taishi-i/awesome-japanese-nlp-resources discover --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/discover .gemini/skills/discover && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "discover" agent skill from https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discover into .gemini/skills/discover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "discover", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install taishi-i/awesome-japanese-nlp-resources discoverInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/discover .github/skills/discover && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "discover" agent skill from https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discover into .github/skills/discover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "discover", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install taishi-i/awesome-japanese-nlp-resources discover --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/taishi-i/awesome-japanese-nlp-resources.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/awesome-japanese-nlp-resources/skills/discover .opencode/skills/discover && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "discover" agent skill from https://github.com/taishi-i/awesome-japanese-nlp-resources/tree/main/plugins/awesome-japanese-nlp-resources/skills/discover into .opencode/skills/discover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "discover", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
discoverGiven a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed…
Discover is an agent skill from taishi-i/awesome-japanese-nlp-resources. Given a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed (contribution candidates). Mines the bundled dataset, then expands via web research across GitHub and Hugging Face. Use when the user names a SPECIFIC repository, model, or tool and wants alternatives/equivalents, OR wants to discover Japanese NLP resources for a topic that are NOT yet in the list, OR wants to prepare a…
Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).
It sits in AI & LLM Engineering, covering Natural language processing, Model hubs and datasets and Comics and storyboards. It works with Hugging Face and GitHub. The repository describes itself as: A curated list of resources for Japanese natural language processing (NLP): Python libraries, LLMs, dictionaries, corpora, and datasets. Includes Claude Code and Codex skills to… The licence is CC0-1.0.
11 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 451770b. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashWebSearchWebFetchFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.cogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Discover loads about 6.5k tokens when it runs. Until then it costs about 229 tokens; SKILL.md has 1,951 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, WebSearch, WebFetchAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from taishi-i/awesome-japanese-nlp-resources at commit 451770b, republished under its CC0-1.0 licence (© taishi-i). 1,951 words, ~6,513 tokens.
.claude/skills/discover/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Discover Japanese NLP resources related to the user's query — both what's already in awesome-japanese-nlp-resources and what's not yet listed.
This skill is shared by the Claude Code and Codex versions of the plugin. The steps are the same in both tools; only these details differ:
/awesome-japanese-nlp-resources:discover, appended at the end of this skill as ARGUMENTS: …. Codex: the user's message that invoked $awesome-japanese-nlp-resources:discover, minus that $… mention. If the skill was picked automatically rather than invoked by name, use the user's request as the query.${CLAUDE_PLUGIN_ROOT}. Codex: the directory two levels above this SKILL.md (use its absolute path).Bash tool or Codex's shell tool. Copy each Python script in full and run it as written, changing only its placeholders (RESOURCES_PATH, SEED, the keyword list, the candidate list) — don't shorten it, drop passes, or alter its scores and thresholds.WebSearch to search and WebFetch to read a page. Codex: the built-in web search tool (search, then open the page). Do not use the gh CLI, curl, or other network commands from the shell./awesome-japanese-nlp-resources:<skill> in Claude Code, $awesome-japanese-nlp-resources:<skill> in Codex.The "Already in the list" results and the "not yet in the list" check both come from the bundled data. If the data file can't be read (for example, shell commands are blocked or fail to start), say so and link https://github.com/taishi-i/awesome-japanese-nlp-resources instead of guessing which resources are listed.
Before anything else, run this once and remember the values — every step that mentions a year refers to them:
echo "YEAR_NOW=$(date +%Y)"
echo "YEAR_PREV=$(($(date +%Y) - 1))"Substitute these wherever this skill writes ${YEAR_NOW} or ${YEAR_PREV}. Do not hardcode years.
If the query is empty or blank, treat it as a topic-mode request for a general search for the latest Japanese NLP resources. Use these defaults for the rest of the steps:
llm, bert, embed, speech, morpholog — short stems for local substring matching (a phrase like japanese language processing almost never occurs verbatim in a description and would silently match nothing; see Step 3's substring-matching note)${YEAR_PREV} ${YEAR_NOW} to every query, and include:japanese NLP new library github ${YEAR_NOW}日本語 NLP 新しい ライブラリ github ${YEAR_NOW}japanese NLP new model huggingface ${YEAR_NOW}huggingface japanese nlp ${YEAR_PREV} ${YEAR_NOW} newDiscover: "<query>"Then skip Step 1 (mode is topic mode) and continue from Step 2.
Seed mode — the query names ONE specific existing repository, model, or tool:
(a) A direct identifier — a full GitHub URL (https://github.com/owner/repo), a full Hugging Face URL (https://huggingface.co/owner/name or https://huggingface.co/datasets/owner/name), an owner/repo pair, or a bare tool/model name (e.g. mecab). Pass it straight through as SEED.
(b) A natural-language phrase wrapping an identifier (e.g. mecabに似たツール, alternatives to fugashi) — extract the core identifier (the URL, owner/repo, or tool/model name) and pass only that as SEED.
If (a) or (b) applies, proceed to Step 2 in seed mode.
Topic mode — the query is a descriptive/topical phrase with no single specific name in it (e.g. 日本語の要約データセット, japanese sentiment analysis dataset, 形態素解析). There is no identifier to extract and no reason to force one — proceed to Step 2 in topic mode directly. Do not invent a seed; topic mode already covers this case via keyword search plus web discovery.
Keep the full query for the language-detection rule in Step 9.
The data ships with the plugin at data/resources.json under the plugin root (see "Claude Code and Codex" above). Resolve its absolute path, falling back to a scoped search only if the install is unusual:
PLUGIN_ROOT="${CLAUDE_PLUGIN_ROOT}" # Codex: replace with the plugin root, two levels above this SKILL.md
RESOURCES_PATH="$PLUGIN_ROOT/data/resources.json"
[ -f "$RESOURCES_PATH" ] || RESOURCES_PATH="$(find "${CODEX_HOME:-$HOME/.codex}/plugins" "${HOME}/.claude/plugins" -type f -name resources.json 2>/dev/null | grep "awesome-japanese-nlp-resources/" | head -1)"
echo "RESOURCES_PATH=$RESOURCES_PATH"Write the resulting absolute path into every script below — shell variables may not persist between commands.
The plugin also ships data/multilingual_resources.json (same item format) listing multilingual libraries, models, and datasets (GitHub repositories) that also support Japanese, from docs/multilingual.md. The scripts below load it automatically when it exists; its items have categories like Multilingual (Speech recognition).
Seed mode: locate the seed and score every other item by shared category, shared semantic labels, and shared description tokens (IDF-weighted so rare, discriminative labels count more than ubiquitous ones). Substitute RESOURCES_PATH (Step 2) and SEED (from Step 1):
python3 << 'EOF'
import json, re, math, os
from collections import Counter
RESOURCES_PATH = "RESOURCES_PATH" # from Step 2
SEED_RAW = "SEED" # from Step 1
with open(RESOURCES_PATH) as f:
data = json.load(f)
multilingual_path = os.path.join(os.path.dirname(RESOURCES_PATH), "multilingual_resources.json")
if os.path.exists(multilingual_path):
with open(multilingual_path) as f:
data += json.load(f)
N = len(data)
STOP = {
"the","and","for","with","that","this","from","into","your","you","are","was",
"japanese","nlp","python","library","tool","tools","text","based","using","use",
"used","language","data","model","models","repository","repo","support","simple",
"fast","easy","also","can","via","etc","https","http","github","com","www","org",
"not","but","all","any","other","such","more","most","than","its","each","which",
}
def norm_url(u): return (u or "").lower().rstrip("/")
def subs_of(it): return set(t.strip().lower() for t in (it.get("s") or []))
def family(c): return (c or "").split("(")[0].strip().lower()
def toks(text):
return set(w for w in re.findall(r"[a-z0-9]+", (text or "").lower())
if len(w) >= 4 and w not in STOP)
seed = SEED_RAW.strip().lower().rstrip("/")
gh_m = re.search(r"github\.com/([^/#?]+/[^/#?]+)", seed)
hf_m = re.search(r"huggingface\.co/(?:datasets/)?([^/#?]+/[^/#?]+)", seed)
slug = (gh_m or hf_m).group(1) if (gh_m or hf_m) else (seed if seed.count("/") == 1 else None)
basename = seed.split("/")[-1]
def find_matches():
exact = [x for x in data if norm_url(x["u"]) == seed]
if exact: return exact, "exact URL"
if slug:
sm = [x for x in data if norm_url(x["u"]).endswith("/" + slug)]
if sm: return sm, "owner/repo"
return [], None
nm = [x for x in data if x["n"].lower() == basename]
if nm: return nm, "name"
loose = [x for x in data if basename and (basename in x["n"].lower() or ("/" + basename) in norm_url(x["u"]))]
if loose: return loose, "loose substring"
return [], None
matches, how = find_matches()
matches.sort(key=lambda x: -(max(x.get("ns") or 0, x.get("nd") or 0)))
seed_item = matches[0] if matches else None
if not seed_item:
print("SEED_NOT_FOUND")
near = [x for x in data if basename and basename[:4] and basename[:4] in x["n"].lower()]
near.sort(key=lambda x: -(max(x.get("ns") or 0, x.get("nd") or 0)))
for x in near[:8]:
print(f" near? {x['n']} | {x['c']} | {x['u']}")
raise SystemExit
print(f"SEED_FOUND via {how}: {seed_item['n']}")
print(f" url={seed_item['u']}")
print(f" c={seed_item['c']}")
print(f" s={seed_item.get('s','')}")
print(f" d={seed_item.get('d','')[:200]}")
if len(matches) > 1:
print(f" NOTE: {len(matches)} entries matched '{seed}' — chose the most popular. Others:")
for x in matches[1:6]:
print(f" - {x['n']} ({x['u']})")
print()
sub_df, tok_df = Counter(), Counter()
for x in data:
for s in subs_of(x): sub_df[s] += 1
for t in toks((x.get("d") or "") + " " + x["n"] + " " + " ".join(x.get("s") or [])): tok_df[t] += 1
SUB_FLOOR, TOK_FLOOR = 1.5, 1.5
def idf_sub(l): return max(0.0, math.log(N / sub_df.get(l, 1)) - SUB_FLOOR)
def idf_tok(t): return max(0.0, math.log(N / tok_df.get(t, 1)) - TOK_FLOOR)
SUB_W, TOK_W, THRESHOLD = 3.0, 1.5, 8.0
seed_cat = seed_item["c"]; seed_fam = family(seed_cat)
seed_subs = subs_of(seed_item)
seed_tok = toks((seed_item.get("d") or "") + " " + seed_item["n"] + " " + " ".join(seed_item.get("s") or []))
seed_name = seed_item["n"].lower()
seed_urls = {norm_url(seed_item["u"])} | {norm_url(x["u"]) for x in matches if x["n"].lower() == seed_name}
results = []
for x in data:
if x.get("status") == "not_found":
continue
if norm_url(x["u"]) in seed_urls:
continue
score = 0.0
if x["c"] == seed_cat: score += 10
elif family(x["c"]) == seed_fam: score += 5
sh_subs = seed_subs & subs_of(x)
score += SUB_W * sum(idf_sub(l) for l in sh_subs)
sh_tok = seed_tok & toks((x.get("d") or "") + " " + x["n"] + " " + " ".join(x.get("s") or []))
score += TOK_W * sum(idf_tok(t) for t in sh_tok)
if score < THRESHOLD:
continue
pop = max(x.get("ns") or 0, x.get("nd") or 0)
results.append((score + 0.5 * pop, score, sorted(sh_subs), sorted(sh_tok), x))
results.sort(key=lambda r: -r[0])
print(f"=== LOCAL SIMILAR ({len(results)} candidates; showing top 15) ===")
for combined, sc, sh_subs, sh_tok, x in results[:15]:
st = x.get("st", 0) or 0; dl = x.get("dl", 0) or 0
print(f"score={combined:.1f}")
print(f" n={x['n']}")
print(f" u={x['u']}")
print(f" c={x['c']}")
print(f" st={st} dl={dl}")
print(f" shared_labels={', '.join(sh_subs) if sh_subs else '-'}")
print(f" shared_terms={', '.join(sh_tok[:6]) if sh_tok else '-'}")
print(f" d={x.get('d','')[:120]}")
print()
EOFTopic mode: translate the topic to 3–5 English keywords (same stem + tool-name conventions as the search skill — see its domain table if unsure), then score by keyword match instead of seed fingerprint. Matching is by literal substring, so prefer short stems (llm, embed, morpholog) over multi-word phrases, which rarely occur verbatim in a description:
python3 << 'EOF'
import json, os
with open("RESOURCES_PATH") as f: # from Step 2
data = json.load(f)
multilingual_path = os.path.join(os.path.dirname("RESOURCES_PATH"), "multilingual_resources.json")
if os.path.exists(multilingual_path):
with open(multilingual_path) as f:
data += json.load(f)
keywords = ["keyword1", "keyword2", "keyword3"] # from Step 1 (topic mode)
results = []
for item in data:
if item.get("status") == "not_found":
continue
n = item.get("n", "").lower()
d = item.get("d", "").lower()
s = " ".join(item.get("s") or []).lower()
c = item.get("c", "").lower()
al = " ".join(item.get("al") or []).lower()
text_score = 0
for kw in keywords:
kw = kw.lower()
if n == kw: text_score += 20
elif kw in n: text_score += 10
if kw in d: text_score += 5
if kw in s: text_score += 3
if kw in c: text_score += 2
if kw in al: text_score += 10
if text_score < 8:
continue
ns = item.get("ns") or 0
nd = item.get("nd") or 0
results.append((text_score + max(ns, nd) * 2.5, item))
results.sort(key=lambda x: -x[0])
print(f"=== LOCAL MATCHES ({len(results)} candidates; showing top 15) ===")
for combined, item in results[:15]:
st = item.get("st", 0) or 0; dl = item.get("dl", 0) or 0
print(f"score={combined:.1f} st={st} dl={dl}")
print(f" n={item['n']}")
print(f" u={item['u']}")
print(f" c={item['c']}")
print(f" s={item.get('s','')}")
print(f" d={item.get('d','')[:120]}")
print()
EOFThis prints the resolved seed (seed mode) or keyword matches (topic mode), plus up to 15 local candidates. Use the shared labels/terms (seed mode) or matched keywords (topic mode) to justify the "Why" column later.
If Step 3 printed SEED_NOT_FOUND:
owner/repo, fetch its page (Claude Code: WebFetch as below; Codex: open it with the web search tool and extract the same fields):WebFetch url="https://github.com/<owner>/<repo>" prompt="Extract as JSON: name, one-line description, primary language, the NLP task it performs, star count, and whether it targets Japanese. If a field is unavailable, set it to null."WebFetch url="https://huggingface.co/<owner>/<name>" prompt="Extract as JSON: name, one-line description / model card summary, pipeline or task tag, whether it is a model or a dataset, downloads count, likes count, and whether it targets Japanese. If a field is unavailable, set it to null."<name> japanese nlp) to identify what it does and where it lives.If the seed cannot be found locally or resolved on the web, this is the only case where you stop and report back: state what you tried, and suggest running the search skill with the same query.
Seed mode — using the seed's task and category, build 6–8 queries mixing GitHub and Hugging Face, English and Japanese:
GitHub-oriented: <seed-name> alternative github, <seed-name> vs, japanese <task> library github, 日本語 <task> ライブラリ github, <seed-name> 類似 ツール
Hugging Face-oriented: <seed-name> alternative huggingface, japanese <task> huggingface, huggingface japanese <task>
Topic mode — reuse the same keywords derived in Step 3 (the search skill's domain-table conventions), build 6–9 queries. Unlike Step 3's local substring matching, these are search-engine queries, so natural multi-word phrasing (e.g. japanese sentence embedding) works fine here even where it wouldn't as a literal substring:
GitHub-oriented: <keyword> japanese site:github.com, japanese <keyword> ${YEAR_NOW} site:github.com, <topic> 日本語 github, awesome japanese <keyword>
Hugging Face-oriented (always include at least 2–3): <keyword> japanese site:huggingface.co, huggingface japanese <keyword> ${YEAR_NOW}, <topic> 日本語 huggingface
Both modes: add ${YEAR_NOW} recency variants when the topic is fast-moving (LLMs, embeddings, speech).
Use the web search and page-fetch tools only (see "Claude Code and Codex" above) — do not use the gh CLI in this project.
Run the queries from Step 5 (6–9 of them). From each result, extract every URL matching:
https://github.com/<owner>/<repo> (ignore deeper paths like /issues, /pull/, /blob/, /tree/)https://huggingface.co/<owner>/<name> for models, https://huggingface.co/datasets/<owner>/<name> for datasets (ignore deeper paths)Collect them, lowercased, trailing slashes stripped, tagged by kind (github / hf_model / hf_dataset), de-duplicated by owner/repo (or owner/name) within each kind, and drop the seed itself (seed mode).
Keep only resources NOT already in the dataset — this section reports unlisted items only; anything already catalogued belongs in the "Already in the list" table from Step 3. Check all collected candidates in one run (both modes), substituting RESOURCES_PATH (Step 2) and the candidate URLs. Besides resources.json, the script also reads the multilingual list, a pre-built data/existing_urls.txt when the plugin ships one, and the list's own README.md / docs/huggingface.md when run from a clone of the repository — resources.json can lag behind the live README. It reads files only, so it also works in read-only sandboxes:
python3 << 'EOF'
import json, re, os
RESOURCES_PATH = "RESOURCES_PATH" # from Step 2
CANDIDATES = """
https://github.com/owner/repo
https://huggingface.co/owner/name
""".split() # every candidate URL collected above, one per line
def canon(u):
u = (u or "").strip().lower().rstrip("/")
m = re.search(r"(github\.com|huggingface\.co)/((?:datasets/|spaces/)?[^/#?\s]+/[^/#?\s]+)", u)
if not m:
return u
path = m.group(2)
if path.endswith(".git"):
path = path[:-4]
return f"https://{m.group(1)}/{path}"
data_dir = os.path.dirname(os.path.abspath(RESOURCES_PATH))
existing = set()
for name in ("resources.json", "multilingual_resources.json"):
path = os.path.join(data_dir, name)
if os.path.exists(path):
with open(path) as f:
existing.update(canon(item.get("u")) for item in json.load(f))
scan_files = [os.path.join(data_dir, "existing_urls.txt")]
p = data_dir
for _ in range(6):
p = os.path.dirname(p)
if os.path.exists(os.path.join(p, "awesome-japanese-nlp-resources.json")):
scan_files += [os.path.join(p, "README.md"), os.path.join(p, "docs", "huggingface.md")]
break
url_pattern = re.compile(r"https://(?:github\.com|huggingface\.co)/(?:datasets/)?[A-Za-z0-9_.-]+/[A-Za-z0-9_.-]+")
for scan_file in scan_files:
if os.path.exists(scan_file):
with open(scan_file) as f:
existing.update(canon(u) for u in url_pattern.findall(f.read()))
print(f"{len(existing)} URLs already in the list")
for c in CANDIDATES:
print(("in list - drop " if canon(c) in existing else "unlisted - keep ") + c)
EOFCap the survivors at 10–15 total (seed mode: 5, spread across both kinds), prioritizing candidates that appear in multiple result sets and keeping a mix of both platforms.
For each surviving candidate, fetch its page — in Claude Code with WebFetch (up to 5 calls in parallel — single message, multiple tool calls); in Codex, open it with the web search tool and extract the same fields:
For github candidates:
WebFetch url="https://github.com/<owner>/<repo>" prompt="Extract as JSON: name, one-line description, primary language, star count, last-updated date (YYYY-MM), whether archived, whether a fork, and what NLP task it performs. Note if it targets Japanese."For hf_model / hf_dataset candidates:
WebFetch url="https://huggingface.co/<owner>/<name>" prompt="Extract as JSON: name, one-line description / model card summary, pipeline or task tag, whether model or dataset, downloads count, likes count, last-updated date (YYYY-MM), and what NLP task it performs. Note if it targets Japanese."Drop any candidate that: is archived (GitHub); is an inactive/throwaway fork (keep only if independently maintained and published — e.g. on PyPI/crates.io/npm — with a distinct improvement); clearly does not perform the same/related task; is not Japanese-NLP-related (unless a widely-known direct equivalent); or (topic mode, general-search default) has stars < 3 / likes < 3 AND is >24 months stale. Keep stars/downloads/likes and last-updated for the survivors.
Group survivors under awesome-japanese-nlp-resources section headings, contribution-ready:
Hugging Face candidates go to "Hugging Face model" or "Hugging Face dataset", inferred from the fetched page (hf_model/hf_dataset kind, or whether the URL contains /datasets/).
GitHub candidates are grouped by primary language and description:
| Heading | Cues |
|---|---|
| Python library / C++ / Rust crate / JavaScript / Go / Java | primary language |
| Pretrained model | "model", "weights", "fine-tuned", links to HuggingFace |
| ChatGPT / LLM application | LLM app, RAG, agent |
| Dictionary and IME | dictionary, lexicon, IME |
| Corpus | corpus, dataset, annotation |
| Tutorial | tutorial, course, lecture, "学習", "入門" |
| Research summary | survey, paper list, "サーベイ" |
Language detection rule (apply before writing any output):
Apply the detected language to all headings and prose. Repository/model descriptions stay in English (the awesome list standard).
English output template:
## Discover: "<query>"
**Seed:** [<seed-name>](<seed-url>) — <category> — <one-line what it does> ← seed mode only
*(Similarity from the bundled dataset + web research across GitHub and Hugging Face)*
### Already in awesome-japanese-nlp-resources
| # | Repository | Category | Popularity | Why |
|---|---|---|---|---|
| 1 | [name](url) | category | ⭐N / 📥N | shares <task/labels> (seed mode) or matched <keywords> (topic mode) |
*(Omit this table, with a one-line note, if Step 3 found nothing.)*
### Not yet in the list — candidates
Found **N** resources not yet in awesome-japanese-nlp-resources (**G** GitHub, **H** Hugging Face).
#### Python library
* [repo-name](https://github.com/owner/repo) - One-line English description. (⭐ 123, last updated: YYYY-MM)
#### Hugging Face model
* [model-name](https://huggingface.co/owner/name) - 📥 {downloads} / ⭐ {likes} / One-line English description.
*(Only the section headings that actually have candidates. Omit the whole block, with a one-line note, if Step 6/7 found nothing unlisted.)*
### Recommendation
- **Closest alternative** (seed mode) or **Highlight** (topic mode): [name](url) — one-line reason.
- **Next step:** `/awesome-japanese-nlp-resources:search <query>` for the full catalog, or `/awesome-japanese-nlp-resources:research <topic>` for a trend/challenge report.
Sources:
- [Title 1](https://...)Japanese output template (when the query is in Japanese): mirror the structure with ## "<query>" の発見結果, ### awesome-japanese-nlp-resources 内の既存リソース, ### 未収録の候補, ### おすすめ, and per-bullet suffixes (⭐ N, 最終更新: YYYY-MM).
Rules:
* [name](url) - description. (⭐ N, last updated: YYYY-MM); Hugging Face * [name](url) - 📥 {downloads} / ⭐ {likes} / description. (abbreviate with k/M, e.g. 📥 367k).Sources: is mandatory (web research requirement) — list the result URLs actually used.$awesome-japanese-nlp-resources:… in Codex.© taishi-i, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in plugins/awesome-japanese-nlp-resources/skills/discover of taishi-i/awesome-japanese-nlp-resources.
Open the folder on GitHubat commit 451770b
Discover next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Discover this skilltaishi-i/awesome-japanese-nlp-resources | 1k | — | ~6.5k | Automated safety check: Notes | CC0-1.0 | |
| Esmfold2JimLiu/science-skills | 228 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Publish Tracelab Huggingfaceuw-syfi/TraceLab | 142 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Transformers Usagedavila7/claude-code-templates | 33k | 11 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Transformers.jshuggingface/skills | 11k | 1 repos | ~6.2k | Automated safety check: Pass | Apache-2.0 |
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
uw-syfi/TraceLab
Prepare, publish, refresh, or validate the TraceLab public dataset on Hugging Face under UW-SyFI/TraceLab.
davila7/claude-code-templates
Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.
huggingface/skills
Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks.
dracohu2025-cloud/draco-skills-collection
A skill your agent uses when collecting public news candidates from Hacker News, GitHub Trending, Hugging Face papers, and other web sources for briefings and daily reports.
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
taishi-i/awesome-japanese-nlp-resources
Analyze current trends and challenges in Japanese NLP for a topic.
taishi-i/awesome-japanese-nlp-resources
Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).
Works with
Categories
Given a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed…. Discover is an agent skill from taishi-i/awesome-japanese-nlp-resources. Given a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed (contribution candidates).
Discover fits situations like: the user names a SPECIFIC repository; tool and wants alternatives/equivalents; wants to discover Japanese NLP resources for a topic that are NOT yet in the list; wants to prepare a contribution.
Run `npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a claude-code`. Or copy the skill folder (plugins/awesome-japanese-nlp-resources/skills/discover in taishi-i/awesome-japanese-nlp-resources) into .claude/skills/discover in your project. Claude Code loads it when a task matches its description.
Run `npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a codex`. Or copy the skill folder (plugins/awesome-japanese-nlp-resources/skills/discover in taishi-i/awesome-japanese-nlp-resources) into .agents/skills/discover in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add taishi-i/awesome-japanese-nlp-resources --skill discover -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/discover, .gemini/skills/discover, .github/skills/discover and .opencode/skills/discover in your project.
SKILL.md names no scripts, command-line tools or credentials: Discover is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, WebSearch, WebFetch.
SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Discover is published under the CC0-1.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Discover: Esmfold2 (JimLiu/science-skills, 228 stars), Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Publish Tracelab Huggingface (uw-syfi/TraceLab, 142 stars) and Hugging Face Transformers Usage (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
taishi-i (a GitHub user) maintains it in taishi-i/awesome-japanese-nlp-resources, which has 1,021 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 6, 2026.
Source: taishi-i/awesome-japanese-nlp-resources on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.