Agent skill

Linguistic Semantic Algorithms

by pproenca in pproenca/dot-skills

Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug…

MITAuto-check passedAI & LLM Engineering

Install Linguistic Semantic Algorithms

skills CLI
$ npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills linguistic-semantic-algorithms --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/linguistic-semantic-algorithms .claude/skills/linguistic-semantic-algorithms && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
linguistic-semantic-algorithms
GitHub stars
214
Token cost
~2.7k tokens
SKILL.md length
930 words
Files
45 (incl. references, assets)
Skills in repo
182
Repo updated
First seen
Licence
MIT

At a glance

Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug…

  • Works in 8 steps: Concept & Domain Extraction (CRITICAL) → Semantic Similarity & Feature Mapping… → Architectural Topology (HIGH) → …
  • Hunting bugs across many files
  • SKILL.md covers When to Apply, Rule Categories by Priority, Quick Reference and How to Use, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Linguistic Semantic Algorithms is an agent skill from pproenca/dot-skills. Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug localization, identifier linguistics, and complexity metrics. Trigger when hunting bugs across many files, scoping a new feature, identifying domain entities, or analyzing commit history — even if the user doesn't explicitly mention algorithms — apply when they ask "where does X live in this codebase?", "what is this codebase…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 47 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).

It sits in AI & LLM Engineering, covering Internationalization, Embeddings and Natural language processing. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Hunting bugs across many files
  • Scoping a new feature
  • Identifying domain entities
  • Analyzing commit history — even if the user doesnt explicitly mention algorithms — apply when they ask where does X live in this codebase?

Example prompts

  • “t explicitly mention algorithms — apply when they ask”
  • “what is this codebase about?”
  • “find duplicated logic”
  • “/linguistic-semantic-algorithms”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Concept & Domain Extraction (CRITICAL)
  2. Semantic Similarity & Feature Mapping (CRITICAL)
  3. Architectural Topology (HIGH)
  4. Co-Change & Temporal Mining (HIGH)
  5. Clone & Duplication Detection (MEDIUM-HIGH)
  6. Bug & Feature Localization (MEDIUM-HIGH)
  7. Identifier Linguistics (MEDIUM)
  8. Complexity & Risk Metrics (MEDIUM)

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • tree-sitter.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Linguistic Semantic Algorithms loads about 2.7k tokens when it runs, and up to ~49k if it reads all its reference files. Until then it costs about 165 tokens; SKILL.md has 930 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~49k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 930 words, ~2,733 tokens.

Download SKILL.mdSave it as .claude/skills/linguistic-semantic-algorithms/SKILL.md (or your agent's skills folder). This skill also uses 44 other files; get the full folder from GitHub.
name
linguistic-semantic-algorithms
description
Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug localization, identifier linguistics, and complexity metrics. Trigger when hunting bugs across many files, scoping a new feature, identifying domain entities, or analyzing commit history — even if the user doesn't explicitly mention algorithms — apply when they ask "where does X live in this codebase?", "what is this codebase about?", "find duplicated logic", "what changed recently?", "who owns this code?", or "is this function risky?".

pproenca Linguistic and Semantic Algorithms Best Practices

Reference of 40 algorithms an agent should reach for when extracting structure, meaning, history, or risk signals from source code and commit data. Categories are ordered by insight-per-effort — how much non-obvious truth the technique exposes relative to how easy it is to apply. The first two categories target the highest-leverage questions: what business entities live in this code? and where else does this concept already exist? — questions that grep and intuition cannot answer.

When to Apply

Reach for these algorithms when:

  • Orienting in an unfamiliar codebase: PageRank the import graph to find the core, run LDA over identifier tokens to discover business themes, mine change coupling to surface hidden architectural couplings.
  • Hunting a bug from a description: BM25 + history prior + embedding re-rank produces a ranked file shortlist far better than grep.
  • Scoping a feature: find prior PRs that did similar work via embedding similarity; map the feature's vocabulary against the codebase's domain via TF-IDF and noun-phrase mining.
  • Reviewing a refactor: AST-level GumTree diff reveals semantic impact text diff hides; PDG isomorphism finds the "same logic, different code" twin you should also update.
  • Auditing risk: hotspots (churn × complexity), bus factor, defect-magnet density, dead-code candidates — together they direct attention to the parts of the codebase that pay back attention.
  • Identifying domain entities and bounded contexts: noun-phrase mining + TF-IDF rare-term extraction + Louvain communities + Jensen-Shannon divergence on per-cluster vocabulary.

Rule Categories by Priority

PriorityCategoryImpactPrefixQuestion answered
1Concept & Domain ExtractionCRITICALconcept-What business entities live in this code?
2Semantic Similarity & Feature MappingCRITICALsim-Where else does this concept already exist?
3Architectural TopologyHIGHgraph-What is the shape of this codebase?
4Co-Change & Temporal MiningHIGHmine-What hidden couplings does history reveal?
5Clone & Duplication DetectionMEDIUM-HIGHclone-Where are we repeating ourselves?
6Bug & Feature LocalizationMEDIUM-HIGHlocal-Given a description, where in code?
7Identifier LinguisticsMEDIUMling-How to prepare tokens so the other algorithms work?
8Complexity & Risk MetricsMEDIUMrisk-Where is the danger concentrated?

Quick Reference

1. Concept & Domain Extraction (CRITICAL)
2. Semantic Similarity & Feature Mapping (CRITICAL)
3. Architectural Topology (HIGH)
4. Co-Change & Temporal Mining (HIGH)
Show full SKILL.md (368 more words)Show less
5. Clone & Duplication Detection (MEDIUM-HIGH)
6. Bug & Feature Localization (MEDIUM-HIGH)
7. Identifier Linguistics (MEDIUM)
8. Complexity & Risk Metrics (MEDIUM)

How to Use

Pick the category that matches the user's question, then read one or two specific rules from that category. Most rules cite combinable partners ("Combine with mine-change-coupling...") that compound the signal — read the partner rule when you need higher precision.

For unfamiliar repos, the highest-ROI starting sequence is:

  1. graph-pagerank-core → read the top-20 most central files
  2. concept-lda-topic-modeling + concept-tfidf-rare-terms → identify the business themes
  3. mine-hotspots-churn-complexity → find where the bugs concentrate
  4. mine-change-coupling → uncover hidden architectural couplings

For a single-task bug or feature, the pipeline is:

  1. local-bm25-saturation (broad candidates) → local-embedding-bug-text (semantic re-rank) → local-history-prior-localization (fix-history boost)
  2. sim-cross-pr-feature-mapping for prior precedent on new features
  3. mine-change-coupling to surface partner files that historically move together

Always preprocess identifier tokens via ling-camel-snake-split → ling-abbreviation-expansion → ling-porter-stemming before any vocabulary-based algorithm. Skipping this step silently degrades every downstream signal.

Cross-language parsing. Most rule code examples use Python's built-in ast module for brevity. For real cross-language work (Go, Rust, Java, TS, C++ in the same repo), use tree-sitter — it provides robust parsers for 40+ languages with a uniform API. Every AST-based rule in this skill (PDG clones, GumTree, Zhang-Shasha, POS-tag heads, identifier co-occurrence) maps cleanly onto tree-sitter ASTs.

Reference Files

FileDescription
references/_sections.mdCategory definitions and impact ordering
assets/templates/_template.mdTemplate for adding new algorithm rules
metadata.jsonVersion and reference information

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 44 other files (references, assets) in skills/.experimental/linguistic-semantic-algorithms of pproenca/dot-skills.

  • SKILL.md
  • AGENTS.md
  • assets/templates/_template.md
  • metadata.json
  • references/_sections.md
  • references/clone-ast-gumtree.md
  • references/clone-minhash-lsh.md
  • references/clone-simhash.md
  • references/clone-suffix-array-cpd.md
  • references/clone-zhang-shasha-ted.md
  • references/concept-bounded-context-detection.md
  • references/concept-entity-name-resolution.md
  • references/concept-identifier-cooccurrence-network.md
  • references/concept-lda-topic-modeling.md
  • references/concept-noun-phrase-mining.md
  • references/concept-tfidf-rare-terms.md
  • references/graph-betweenness-bottlenecks.md
  • references/graph-feedback-arcs.md
  • … and 27 more

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Linguistic Semantic Algorithms next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Linguistic Semantic Algorithms compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Linguistic Semantic Algorithms this skillpproenca/dot-skills214—~2.7kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.6kAutomated safety check: PassMIT
Researchtaishi-i/awesome-japanese-nlp-resources1k—~3.5kAutomated safety check: NotesCC0-1.0
Searchtaishi-i/awesome-japanese-nlp-resources1k—~4.3kAutomated safety check: NotesCC0-1.0
Scholar Computejoshzyj/open-scholar-skill167—~15kAutomated safety check: PassCustom licence

Similar skills

  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 3 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Search

    taishi-i/awesome-japanese-nlp-resources

    Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).

    1k GitHub stars~4.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Scholar Compute

    joshzyj/open-scholar-skill

    Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…

    167 GitHub stars~15k tokensUpdated 19 days ago
    AI & LLM EngineeringAuto-check passed
  • NLP Engineering

    majiayu000/claude-skill-registry

    A skill your agent uses when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization.

    666 GitHub starsUsed in 1 repo~4.4k tokens
    AI & LLM EngineeringAuto-check passed

More from pproenca/dot-skills

All 182 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    214 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    214 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    214 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    214 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    214 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    214 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Linguistic Semantic Algorithms

What does Linguistic Semantic Algorithms do?

Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug…. Linguistic Semantic Algorithms is an agent skill from pproenca/dot-skills. Mapping out an unfamiliar codebase via NLP and graph algorithms — 40 algorithms across topic modelling, semantic embeddings, code graphs, repository mining, clone detection, IR-based bug localization, identifier linguistics, and complexity metrics.

When should I use Linguistic Semantic Algorithms?

Linguistic Semantic Algorithms fits situations like: hunting bugs across many files; scoping a new feature; identifying domain entities; analyzing commit history — even if the user doesnt explicitly mention algorithms — apply when they ask where does X live in this codebase?.

How do I install Linguistic Semantic Algorithms in Claude Code?

Run `npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a claude-code`. Or copy the skill folder (skills/.experimental/linguistic-semantic-algorithms in pproenca/dot-skills) into .claude/skills/linguistic-semantic-algorithms in your project. Claude Code loads it when a task matches its description.

How do I install Linguistic Semantic Algorithms in Codex?

Run `npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a codex`. Or copy the skill folder (skills/.experimental/linguistic-semantic-algorithms in pproenca/dot-skills) into .agents/skills/linguistic-semantic-algorithms in your project. Codex loads it when a task matches its description.

Can I use Linguistic Semantic Algorithms in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill linguistic-semantic-algorithms -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/linguistic-semantic-algorithms, .gemini/skills/linguistic-semantic-algorithms, .github/skills/linguistic-semantic-algorithms and .opencode/skills/linguistic-semantic-algorithms in your project.

What does Linguistic Semantic Algorithms need to run?

SKILL.md names no scripts, command-line tools or credentials: Linguistic Semantic Algorithms is instructions for the agent only.

Does Linguistic Semantic Algorithms access the network?

SKILL.md names 1 domain. As links in the text: tree-sitter.github.io. This is read from the text; nothing was executed.

Is Linguistic Semantic Algorithms safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Linguistic Semantic Algorithms use?

Linguistic Semantic Algorithms is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Linguistic Semantic Algorithms use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 46k tokens, read only when the agent opens those files.

What are the alternatives to Linguistic Semantic Algorithms?

Skills that share tags, products or a category with Linguistic Semantic Algorithms: Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars), Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Research (taishi-i/awesome-japanese-nlp-resources, 1k stars) and Search (taishi-i/awesome-japanese-nlp-resources, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Linguistic Semantic Algorithms?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.