Agent skill

Algo NLP Similarity

by asgard-ai-platform in asgard-ai-platform/skills

Calculate text similarity using lexical and semantic methods for matching and deduplication.

MITAuto-check passedAI & LLM Engineering

Install Algo NLP Similarity

skills CLI
$ npx skills add asgard-ai-platform/skills --skill algo-nlp-similarity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install asgard-ai-platform/skills algo-nlp-similarity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/asgard-ai-platform/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/algo-nlp-similarity .claude/skills/algo-nlp-similarity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
algo-nlp-similarity
GitHub stars
242
Token cost
~1k tokens
SKILL.md length
386 words
Files
4 (incl. references)
Skills in repo
207
Repo updated
First seen
Licence
MIT

At a glance

Calculate text similarity using lexical and semantic methods for matching and deduplication.

  • Works in 4 steps: Input Validation → Core Algorithm → Verification → …
  • The user needs to find similar documents
  • SKILL.md covers Overview, When to Use, Algorithm and Output Format, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Algo NLP Similarity is an agent skill from asgard-ai-platform/skills. Calculate text similarity using lexical and semantic methods for matching and deduplication. Use this skill when the user needs to find similar documents, detect near-duplicates, or measure semantic closeness between texts — even if they say 'how similar are these texts', 'find duplicates', or 'semantic matching'.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `examples/sample_scenario.md`, `references/ann-search.md` and `references/model-benchmarks.md`).

It sits in AI & LLM Engineering, covering Natural language processing and Data cleaning. The repository describes itself as: 301 open-source coding agent skills across 22 domains — methodology, judgment & gotchas packaged as Claude Agent Skills for the Asgard AI Platform. The licence is MIT.

When your agent uses it

  • The user needs to find similar documents
  • Detect near-duplicates
  • Measure semantic closeness between texts — even if they say how similar are these texts
  • Find duplicates

Example prompts

  • “how similar are these texts”
  • “find duplicates”
  • “semantic matching”
  • “/algo-nlp-similarity”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Input Validation
  2. Core Algorithm
  3. Verification
  4. Output

What it can do on your machine

Read from SKILL.md and the folder at commit 4e7f4f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Algo NLP Similarity loads about 1k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 386 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from asgard-ai-platform/skills at commit 4e7f4f8, republished under its MIT licence (© asgard-ai-platform). 386 words, ~1,038 tokens.

Download SKILL.mdSave it as .claude/skills/algo-nlp-similarity/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
algo-nlp-similarity
description
Calculate text similarity using lexical and semantic methods for matching and deduplication. Use this skill when the user needs to find similar documents, detect near-duplicates, or measure semantic closeness between texts — even if they say 'how similar are these texts', 'find duplicates', or 'semantic matching'.
metadata.category
WP-45 NLP 演算法
metadata.tags
nlp, text-similarity, embeddings, deduplication

Text Similarity

Overview

Text similarity measures how close two texts are in meaning or surface form. Lexical methods (Jaccard, cosine on TF-IDF) compare word overlap. Semantic methods (sentence embeddings) capture meaning even with different words. Choice depends on whether you need exact matching or meaning matching.

When to Use

Trigger conditions:

  • Finding similar or duplicate documents in a collection
  • Matching queries to FAQ answers or knowledge base entries
  • Detecting plagiarism or content reuse

When NOT to use:

  • For topic-level grouping (use topic modeling / LDA)
  • For entity extraction from text (use NER)

Algorithm

IRON LAW: Lexical Similarity ≠ Semantic Similarity
"The car is fast" and "The automobile is speedy" have LOW lexical
similarity (different words) but HIGH semantic similarity (same meaning).
"Bank of the river" and "Bank account" have HIGH lexical similarity
but LOW semantic similarity. Choose the method that matches your
definition of "similar."
Phase 1: Input Validation

Determine: similarity type needed (lexical or semantic), text preprocessing requirements, scale (pairwise vs all-pairs vs query-to-corpus). Gate: Texts preprocessed, method selected.

Phase 2: Core Algorithm

Lexical methods:

  • Jaccard: |A∩B| / |A∪B| on word sets
  • Cosine on TF-IDF vectors: cos(θ) = (A·B) / (|A|×|B|)

Semantic methods:

  • Sentence embeddings: encode texts with sentence-transformers (all-MiniLM-L6-v2)
  • Cosine similarity on embedding vectors
  • For large-scale: use FAISS or Annoy for approximate nearest neighbor search
Phase 3: Verification

Spot-check: highly similar pairs should be genuinely similar. Low-similarity pairs should be genuinely different. Check threshold calibration. Gate: Similarity scores align with human judgment on sample pairs.

Phase 4: Output

Return similarity scores or nearest neighbors.

Output Format

json
{
  "similarities": [{"text_a": "doc1", "text_b": "doc5", "score": 0.92, "method": "semantic_cosine"}],
  "metadata": {"method": "sentence-transformers", "model": "all-MiniLM-L6-v2", "pairs_computed": 500}
}

Examples

Sample I/O

Input: Text A: "How to reset my password", Text B: "I forgot my login credentials" Expected: Lexical (Jaccard) ≈ 0.07 (almost no word overlap). Semantic ≈ 0.82 (same intent).

Show full SKILL.md (154 more words)Show less
Edge Cases
InputExpectedWhy
Identical textsScore = 1.0Exact match
Empty textUndefined or 0Handle gracefully
Different languagesLexical=0, semantic depends on modelMultilingual models can match cross-language

Gotchas

  • Threshold is use-case specific: 0.8 similarity might mean "duplicate" for deduplication but "somewhat related" for recommendation. Calibrate threshold on labeled examples.
  • Text length effects: Cosine on TF-IDF is sensitive to document length. Very short texts have sparse vectors with unreliable similarity. Use embeddings for short texts.
  • Embedding model choice: Different models have different strengths. all-MiniLM-L6-v2 is fast but less accurate than larger models. Match model to performance needs.
  • Computational scaling: All-pairs similarity on N documents is O(N²). For large corpora, use approximate methods (locality-sensitive hashing, FAISS).
  • Domain adaptation: General-purpose embedding models may not capture domain-specific similarity (legal, medical). Fine-tune on domain data for best results.

References

  • For embedding model comparison and benchmarks, see references/model-benchmarks.md
  • For approximate nearest neighbor search at scale, see references/ann-search.md

© asgard-ai-platform, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in algo-nlp-similarity of asgard-ai-platform/skills.

  • SKILL.md
  • examples/sample_scenario.md
  • references/ann-search.md
  • references/model-benchmarks.md

Open the folder on GitHubat commit 4e7f4f8

Compare with similar skills

Algo NLP Similarity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Algo NLP Similarity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Algo NLP Similarity this skillasgard-ai-platform/skills242—~1kAutomated safety check: PassMIT
Big Data Labeling Variable ConstructionDrchronx/ai-agent-research-starter-kit139—~1.1kAutomated safety check: PassCustom licence
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
OpenMed Model Card Writermaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Big Data Labeling Variable Construction

    Drchronx/ai-agent-research-starter-kit

    Route text-labeling requests across LDA topic modeling, sklearn baselines, pretrained transformer models, and OpenAI-compatible LLM labeling.

    139 GitHub stars~1.1k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from asgard-ai-platform/skills

All 207 skills in this repo
  • Algo Ecom Bm25

    asgard-ai-platform/skills

    Implement BM25 ranking function for e-commerce product search relevance scoring.

    242 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Mfg Cpk

    asgard-ai-platform/skills

    Calculate Cpk process capability index to assess whether a process meets specification requirements.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Price Elasticity

    asgard-ai-platform/skills

    Calculate price elasticity of demand to quantify how price changes affect sales volume.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Bayesian

    asgard-ai-platform/skills

    Apply Bayesian averaging to rank items by combining observed ratings with prior expectations.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Elo

    asgard-ai-platform/skills

    Implement Elo rating system to rank items or players from pairwise comparison outcomes.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Wilson

    asgard-ai-platform/skills

    Calculate Wilson Score confidence intervals for ranking items by positive proportion with sample size correction.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Algo NLP Similarity

What does Algo NLP Similarity do?

Calculate text similarity using lexical and semantic methods for matching and deduplication. Algo NLP Similarity is an agent skill from asgard-ai-platform/skills. Calculate text similarity using lexical and semantic methods for matching and deduplication.

When should I use Algo NLP Similarity?

Algo NLP Similarity fits situations like: the user needs to find similar documents; detect near-duplicates; measure semantic closeness between texts — even if they say how similar are these texts; find duplicates.

How do I install Algo NLP Similarity in Claude Code?

Run `npx skills add asgard-ai-platform/skills --skill algo-nlp-similarity -a claude-code`. Or copy the skill folder (algo-nlp-similarity in asgard-ai-platform/skills) into .claude/skills/algo-nlp-similarity in your project. Claude Code loads it when a task matches its description.

How do I install Algo NLP Similarity in Codex?

Run `npx skills add asgard-ai-platform/skills --skill algo-nlp-similarity -a codex`. Or copy the skill folder (algo-nlp-similarity in asgard-ai-platform/skills) into .agents/skills/algo-nlp-similarity in your project. Codex loads it when a task matches its description.

Can I use Algo NLP Similarity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add asgard-ai-platform/skills --skill algo-nlp-similarity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/algo-nlp-similarity, .gemini/skills/algo-nlp-similarity, .github/skills/algo-nlp-similarity and .opencode/skills/algo-nlp-similarity in your project.

What does Algo NLP Similarity need to run?

SKILL.md names no scripts, command-line tools or credentials: Algo NLP Similarity is instructions for the agent only.

Does Algo NLP Similarity access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Algo NLP Similarity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Algo NLP Similarity use?

Algo NLP Similarity is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Algo NLP Similarity use?

About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.

What are the alternatives to Algo NLP Similarity?

Skills that share tags, products or a category with Algo NLP Similarity: Big Data Labeling Variable Construction (Drchronx/ai-agent-research-starter-kit, 139 stars), Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Dingo Verify (MigoXLab/dingo, 757 stars) and OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Algo NLP Similarity?

asgard-ai-platform (a GitHub organization) maintains it in asgard-ai-platform/skills, which has 242 GitHub stars. The repository holds 207 skills in this directory. The repository was last updated on June 6, 2026.

Source: asgard-ai-platform/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.