Compare
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug…
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithms --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .claude/skills/linguistic-semantic-algorithms && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "linguistic-semantic-algorithms" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithms into .claude/skills/linguistic-semantic-algorithms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linguistic-semantic-algorithms", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithmsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithms --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .agents/skills/linguistic-semantic-algorithms && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "linguistic-semantic-algorithms" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithms into .agents/skills/linguistic-semantic-algorithms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linguistic-semantic-algorithms", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithms --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .cursor/skills/linguistic-semantic-algorithms && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "linguistic-semantic-algorithms" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithms into .cursor/skills/linguistic-semantic-algorithms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linguistic-semantic-algorithms", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pproenca/dot-skills.git --path skills/.experimental/linguistic-semantic-algorithms--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithms --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .gemini/skills/linguistic-semantic-algorithms && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "linguistic-semantic-algorithms" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithms into .gemini/skills/linguistic-semantic-algorithms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linguistic-semantic-algorithms", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithmsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .github/skills/linguistic-semantic-algorithms && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "linguistic-semantic-algorithms" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithms into .github/skills/linguistic-semantic-algorithms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linguistic-semantic-algorithms", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithms --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .opencode/skills/linguistic-semantic-algorithms && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "linguistic-semantic-algorithms" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/linguistic-semantic-algorithms into .opencode/skills/linguistic-semantic-algorithms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linguistic-semantic-algorithms", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
linguistic-semantic-algorithmsMapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug…
Linguistic Semantic Algorithms is an agent skill from pproenca/dot-skills. Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug localization, identifier linguistics, and complexity metrics. Trigger when hunting bugs across many files, scoping a new feature, identifying domain entities, or analyzing commit history — even if the user doesn't explicitly mention algorithms — apply when they ask "where does X live in this codebase?", "what is this codebase…
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 47 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).
It sits in AI & LLM Engineering, covering Internationalization, Embeddings and Natural language processing. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
tree-sitter.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Linguistic Semantic Algorithms loads about 2.7k tokens when it runs, and up to ~49k if it reads all its reference files. Until then it costs about 165 tokens; SKILL.md has 930 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 930 words, ~2,733 tokens.
.claude/skills/linguistic-semantic-algorithms/SKILL.md (or your agent's skills folder). This skill also uses 44 other files; get the full folder from GitHub.Reference of 40 algorithms an agent should reach for when extracting structure, meaning, history, or risk signals from source code and commit data. Categories are ordered by insight-per-effort — how much non-obvious truth the technique exposes relative to how easy it is to apply. The first two categories target the highest-leverage questions: what business entities live in this code? and where else does this concept already exist? — questions that grep and intuition cannot answer.
Reach for these algorithms when:
| Priority | Category | Impact | Prefix | Question answered |
|---|---|---|---|---|
| 1 | Concept & Domain Extraction | CRITICAL | concept- | What business entities live in this code? |
| 2 | Semantic Similarity & Feature Mapping | CRITICAL | sim- | Where else does this concept already exist? |
| 3 | Architectural Topology | HIGH | graph- | What is the shape of this codebase? |
| 4 | Co-Change & Temporal Mining | HIGH | mine- | What hidden couplings does history reveal? |
| 5 | Clone & Duplication Detection | MEDIUM-HIGH | clone- | Where are we repeating ourselves? |
| 6 | Bug & Feature Localization | MEDIUM-HIGH | local- | Given a description, where in code? |
| 7 | Identifier Linguistics | MEDIUM | ling- | How to prepare tokens so the other algorithms work? |
| 8 | Complexity & Risk Metrics | MEDIUM | risk- | Where is the danger concentrated? |
concept-lda-topic-modeling — LDA over identifier tokens surfaces latent business themesconcept-noun-phrase-mining — POS-tag + chunk identifiers to extract entity candidatesconcept-tfidf-rare-terms — IDF against a generic corpus isolates domain vocabulary from framework noiseconcept-identifier-cooccurrence-network — PMI-weighted co-occurrence graph reveals conceptual neighborhoodsconcept-entity-name-resolution — Cluster name variants (user/usr/u/userAccount) via embedding + edit distanceconcept-bounded-context-detection — Louvain + Jensen-Shannon divergence detects DDD bounded contextssim-codebert-embeddings — CodeBERT + cosine for semantic code search across renamessim-pdg-semantic-clones — Program Dependence Graph isomorphism finds Type-4 clonessim-cross-pr-feature-mapping — Embed merged PRs once, retrieve precedent at feature-design timesim-cosine-vsm-files — TF-IDF VSM file similarity when no GPU is availablesim-call-pattern-similarity — N-grams on call-sequence find behavioral twinssim-doc-code-alignment — Joint code-doc embedding flags drift between docs and codegraph-pagerank-core — PageRank the import graph to find the codebase coregraph-betweenness-bottlenecks — Betweenness centrality surfaces bottleneck modulesgraph-louvain-modules — Louvain community detection reveals natural module boundariesgraph-scc-cycle-tangles — Tarjan's SCC algorithm exposes circular-dependency tanglesgraph-feedback-arcs — Eades-Lin-Smyth FAS chooses the smallest cycle-breaking cutmine-change-coupling — Conditional probability over commit history exposes hidden couplingmine-hotspots-churn-complexity — Churn × complexity = canonical hotspot score (Tornhill)mine-bus-factor — Per-file authorship Gini coefficient surfaces knowledge concentrationmine-commit-topic-modeling — LDA on commit messages reveals quarterly themesmine-bug-fix-density — Classify commits, rank files by fix-density to find defect magnetsmine-codebase-aging — Last-modified age + reachability splits stable code from dead codeclone-minhash-lsh — MinHash + LSH for sub-linear near-duplicate retrievalclone-simhash — SimHash 64-bit fingerprints for O(1) Hamming-distance lookupsclone-suffix-array-cpd — Token-level suffix array (PMD CPD) for precise clone boundariesclone-ast-gumtree — GumTree algorithm for fine-grained AST differencingclone-zhang-shasha-ted — Zhang-Shasha tree edit distance for exact subtree similaritylocal-tfidf-bug-reports — TF-IDF rank source files against bug report tokenslocal-bm25-saturation — BM25 handles length normalization and TF saturationlocal-history-prior-localization — Bayesian fusion of IR score with bug-history priorlocal-embedding-bug-text — Two-stage BM25 + embedding re-rank for semantic localizationling-camel-snake-split — Split camelCase, snake_case, digit-boundaries before any analysisling-abbreviation-expansion — Expand idx→index, mgr→manager via dictionary + miningling-porter-stemming — Apply Porter stemmer to unify singular/plural formsling-pos-tagging-identifiers — POS-tag identifier heads to flag misnamed functions/classesrisk-cyclomatic-mccabe — McCabe cyclomatic complexity for branch-test surfacerisk-cognitive-complexity — SonarSource Cognitive Complexity for readability gatesrisk-halstead-volume — Halstead volume for language-agnostic size and effortrisk-shannon-entropy-naming — Per-token entropy flags overloaded namesPick the category that matches the user's question, then read one or two specific rules from that category. Most rules cite combinable partners ("Combine with mine-change-coupling...") that compound the signal — read the partner rule when you need higher precision.
For unfamiliar repos, the highest-ROI starting sequence is:
graph-pagerank-core → read the top-20 most central filesconcept-lda-topic-modeling + concept-tfidf-rare-terms → identify the business themesmine-hotspots-churn-complexity → find where the bugs concentratemine-change-coupling → uncover hidden architectural couplingsFor a single-task bug or feature, the pipeline is:
local-bm25-saturation (broad candidates) → local-embedding-bug-text (semantic re-rank) → local-history-prior-localization (fix-history boost)sim-cross-pr-feature-mapping for prior precedent on new featuresmine-change-coupling to surface partner files that historically move togetherAlways preprocess identifier tokens via ling-camel-snake-split → ling-abbreviation-expansion → ling-porter-stemming before any vocabulary-based algorithm. Skipping this step silently degrades every downstream signal.
Cross-language parsing. Most rule code examples use Python's built-in ast module for brevity. For real cross-language work (Go, Rust, Java, TS, C++ in the same repo), use tree-sitter — it provides robust parsers for 40+ languages with a uniform API. Every AST-based rule in this skill (PDG clones, GumTree, Zhang-Shasha, POS-tag heads, identifier co-occurrence) maps cleanly onto tree-sitter ASTs.
| File | Description |
|---|---|
| references/_sections.md | Category definitions and impact ordering |
| assets/templates/_template.md | Template for adding new algorithm rules |
| metadata.json | Version and reference information |
© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 44 other files (references, assets) in skills/.experimental/linguistic-semantic-algorithms of pproenca/dot-skills.
Open the folder on GitHubat commit cf93c57
Linguistic Semantic Algorithms next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Linguistic Semantic Algorithms this skillpproenca/dot-skills | 214 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Comparetaishi-i/awesome-japanese-nlp-resources | 1k | — | ~4.1k | Automated safety check: Notes | CC0-1.0 | |
| Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Researchtaishi-i/awesome-japanese-nlp-resources | 1k | — | ~3.5k | Automated safety check: Notes | CC0-1.0 | |
| Searchtaishi-i/awesome-japanese-nlp-resources | 1k | — | ~4.3k | Automated safety check: Notes | CC0-1.0 | |
| Scholar Computejoshzyj/open-scholar-skill | 167 | — | ~15k | Automated safety check: Pass | Custom licence |
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
Orchestra-Research/AI-Research-SKILLs
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
taishi-i/awesome-japanese-nlp-resources
Analyze current trends and challenges in Japanese NLP for a topic.
taishi-i/awesome-japanese-nlp-resources
Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).
joshzyj/open-scholar-skill
Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…
majiayu000/claude-skill-registry
A skill your agent uses when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization.
pproenca/dot-skills
Audio forensics and voice recovery guidelines for CSI-level audio analysis.
pproenca/dot-skills
Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.
pproenca/dot-skills
Create well-structured RFCs and technical proposals for software projects.
pproenca/dot-skills
Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.
pproenca/dot-skills
Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…
pproenca/dot-skills
Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…
Categories
Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug…. Linguistic Semantic Algorithms is an agent skill from pproenca/dot-skills. Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug localization, identifier linguistics, and complexity metrics.
Linguistic Semantic Algorithms fits situations like: hunting bugs across many files; scoping a new feature; identifying domain entities; analyzing commit history — even if the user doesnt explicitly mention algorithms — apply when they ask where does X live in this codebase?.
Run `npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a claude-code`. Or copy the skill folder (skills/.experimental/linguistic-semantic-algorithms in pproenca/dot-skills) into .claude/skills/linguistic-semantic-algorithms in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a codex`. Or copy the skill folder (skills/.experimental/linguistic-semantic-algorithms in pproenca/dot-skills) into .agents/skills/linguistic-semantic-algorithms in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/linguistic-semantic-algorithms, .gemini/skills/linguistic-semantic-algorithms, .github/skills/linguistic-semantic-algorithms and .opencode/skills/linguistic-semantic-algorithms in your project.
SKILL.md names no scripts, command-line tools or credentials: Linguistic Semantic Algorithms is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: tree-sitter.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Linguistic Semantic Algorithms is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 46k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Linguistic Semantic Algorithms: Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars), Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Research (taishi-i/awesome-japanese-nlp-resources, 1k stars) and Search (taishi-i/awesome-japanese-nlp-resources, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.
Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.