Agent skill

Extracting Keywords

by oaustegard in oaustegard/claude-skills

Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese).

MITAuto-check passedWriting & Content

Install Extracting Keywords

skills CLI
$ npx skills add oaustegard/claude-skills --skill extracting-keywords -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills extracting-keywords --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/extracting-keywords .claude/skills/extracting-keywords && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extracting-keywords
GitHub stars
150
Token cost
~2.9k tokens
SKILL.md length
632 words
Files
4 (incl. assets)
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese).

  • Users request keyword extraction
  • SKILL.md covers Installation, Stopwords Configuration, Basic Usage and Domain-Specific Extraction, plus 6 more sections
  • Calls uv and python
  • Topic identification

What it does

Extracting Keywords is an agent skill from oaustegard/claude-skills. Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including assets (for example `README.md`).

It sits in Writing & Content, covering Summarization. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Users request keyword extraction
  • Topic identification
  • Content summarization
  • Document analysis

Example prompts

  • “/extracting-keywords”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90b0f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extracting Keywords loads about 2.9k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 632 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 90b0f1b, republished under its MIT licence (© oaustegard). 632 words, ~2,913 tokens.

Download SKILL.mdSave it as .claude/skills/extracting-keywords/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
extracting-keywords
description
Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Use when users request keyword extraction, key terms, topic identification, content summarization, or document analysis. Includes domain-specific stopwords for AI/ML and life sciences. Optional deeper extraction mode (n=2+n=3 combined) for comprehensive coverage.
metadata.version
0.2.1

Extracting Keywords

Extract keywords from text using YAKE (Yet Another Keyword Extractor), an unsupervised statistical keyword extraction algorithm.

Installation

First time only: Install YAKE with optimized dependencies to avoid unnecessary downloads.

bash
cd /home/claude
uv venv yake-venv --system-site-packages
uv pip install yake --python yake-venv/bin/python --no-deps
uv pip install jellyfish segtok regex --python yake-venv/bin/python

This reuses system packages (numpy, networkx) instead of downloading them (~0.08s vs ~5s).

Stopwords Configuration

Built-in YAKE stopwords (34 languages): Use lan="<code>" parameter

  • See Parameters section below for all 34 supported language codes
  • English (lan="en") is the default

Custom domain stopwords (bundled in assets/):

AI/ML: stopwords_ai.txt

  • English stopwords + 783 AI/ML domain-specific terms (1357 total)
  • Filters AI/ML methodology noise (model, training, network, algorithm, parameter)
  • Filters ML boilerplate (dataset, baseline, benchmark, experiment, evaluation)
  • Filters technical terms (transformer, embedding, attention, optimization, inference)
  • Includes full lemmatization (train/trains/trained/training/trainer)
  • Use for AI/ML papers, technical reports, machine learning literature
  • Performance impact: +4-5% runtime vs English stopwords

Life Sciences: stopwords_ls.txt

  • English stopwords + 719 life sciences domain-specific terms (1293 total)
  • Filters research methodology noise (study, results, analysis, significant, observed)
  • Filters academic boilerplate (paper, manuscript, publication, review, editing)
  • Filters statistical terms (correlation, distribution, deviation, variance)
  • Filters clinical terms (patient, treatment, diagnosis, symptom, therapy)
  • Filters biology/medicine (cell, tissue, protein, gene, organism)
  • Includes full lemmatization (analyze/analyzes/analyzed/analyzing/analysis)
  • Use for biomedical papers, clinical studies, research articles, scientific literature
  • Performance impact: +4-5% runtime vs English stopwords

Basic Usage

python
import yake

# Read text
with open('document.txt', 'r') as f:
    text = f.read()

# Extract with English stopwords (default)
kw_extractor = yake.KeywordExtractor(
    lan="en",           # Language code
    n=3,                # Max n-gram size (1-3 word phrases)
    dedupLim=0.9,       # Deduplication threshold (0-1)
    top=20              # Number of keywords to return
)

keywords = kw_extractor.extract_keywords(text)

# Display results (lower score = more important)
for kw, score in keywords:
    print(f"{score:.4f}  {kw}")

Domain-Specific Extraction

Using Life Sciences Stopwords

Option 1: Install custom stopwords file

bash
# Copy life sciences stopwords to YAKE package
cp assets/stopwords_ls.txt /home/claude/yake-venv/lib/python3.12/site-packages/yake/core/StopwordsList/stopwords_ls.txt

# Use with lan="ls"
kw_extractor = yake.KeywordExtractor(lan="ls", n=3, top=20)

Option 2: Load custom stopwords directly

python
# Load stopwords from file
with open('assets/stopwords_ls.txt', 'r') as f:
    custom_stops = set(line.strip().lower() for line in f)

# Pass to extractor
kw_extractor = yake.KeywordExtractor(
    stopwords=custom_stops,
    n=3,
    top=20
)
Using AI/ML Stopwords
python
# Load AI/ML stopwords
with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
    ai_stops = set(line.strip().lower() for line in f)

# Extract with AI stopwords
kw_extractor = yake.KeywordExtractor(
    stopwords=ai_stops,
    n=3,
    top=20
)
keywords = kw_extractor.extract_keywords(text)

Deeper Extraction (n=2 + n=3 Combined)

For more comprehensive extraction, run both n=2 and n=3 and consolidate results. This captures both focused phrases and broader context with ~100% time overhead (still <2s for large documents).

python
import yake

# Load domain stopwords
with open('/mnt/skills/user/extracting-keywords/assets/stopwords_ai.txt', 'r') as f:
    stops = set(line.strip().lower() for line in f)

# Extract with n=2 (captures focused phrases)
kw_n2 = yake.KeywordExtractor(stopwords=stops, n=2, dedupLim=0.9, top=50)
results_n2 = kw_n2.extract_keywords(text)

# Extract with n=3 (captures broader context)
kw_n3 = yake.KeywordExtractor(stopwords=stops, n=3, dedupLim=0.9, top=50)
results_n3 = kw_n3.extract_keywords(text)

# Consolidate: union with score averaging for overlaps
combined = {}
for kw, score in results_n2:
    combined[kw] = score
for kw, score in results_n3:
    if kw in combined:
        combined[kw] = (combined[kw] + score) / 2
    else:
        combined[kw] = score

# Sort by score (lower = more important)
consolidated = sorted(combined.items(), key=lambda x: x[1])

# Display top 30
for kw, score in consolidated[:30]:
    print(f"{score:.4f}  {kw}")

Benefits:

  • n=2 extracts cleaner domain-specific phrases ("disk move", "error rate")
  • n=3 captures contextual combinations ("Move disk 1", "per-step error rate")
  • Consolidation provides richer keyword set for topic modeling or search indexing

Performance:

  • Combined approach: ~2x runtime of single extraction
  • Typical timing: 0.4s (small doc) to 1.0s (large doc)
  • Use when quality matters more than speed
Show full SKILL.md (318 more words)Show less

Parameters

lan (str): Language code for built-in stopwords

  • "en" - English (default)
  • "ai" - AI/ML (if stopwords_ai.txt installed in YAKE)
  • "ls" - Life sciences (if stopwords_ls.txt installed in YAKE)

Built-in YAKE languages (34 total):

  • "ar" - Arabic
  • "bg" - Bulgarian
  • "br" - Breton
  • "cz" - Czech
  • "da" - Danish
  • "de" - German
  • "el" - Greek
  • "es" - Spanish
  • "et" - Estonian
  • "fa" - Farsi/Persian
  • "fi" - Finnish
  • "fr" - French
  • "hi" - Hindi
  • "hr" - Croatian
  • "hu" - Hungarian
  • "hy" - Armenian
  • "id" - Indonesian
  • "it" - Italian
  • "ja" - Japanese
  • "lt" - Lithuanian
  • "lv" - Latvian
  • "nl" - Dutch
  • "no" - Norwegian
  • "pl" - Polish
  • "pt" - Portuguese
  • "ro" - Romanian
  • "ru" - Russian
  • "sk" - Slovak
  • "sl" - Slovenian
  • "sv" - Swedish
  • "tr" - Turkish
  • "uk" - Ukrainian
  • "zh" - Chinese

n (int): Maximum n-gram size (default: 3)

  • 1 - Single words only
  • 2 - Up to 2-word phrases
  • 3 - Up to 3-word phrases (recommended)
  • 4-5 - May produce suboptimal results with YAKE's algorithm

dedupLim (float): Deduplication threshold (default: 0.9)

  • Range: 0.0 to 1.0
  • Higher values = more aggressive deduplication
  • Controls handling of similar terms (e.g., "cancer cell" vs "cancer cells")

top (int): Number of keywords to return (default: 20)

stopwords (set): Custom stopwords set (overrides lan parameter)

Workflow Patterns

Single Document Analysis
python
import yake

# Read document
with open('/mnt/user-data/uploads/article.txt', 'r') as f:
    text = f.read()

# Extract keywords
kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=30)
keywords = kw_extractor.extract_keywords(text)

# Format results
results = []
for kw, score in keywords:
    results.append(f"{score:.4f}  {kw}")

print("\n".join(results))
Comparing Stopwords Strategies
python
import yake

# Load life sciences stopwords
with open('assets/stopwords_ls.txt', 'r') as f:
    ls_stops = set(line.strip().lower() for line in f)

# Extract with English stopwords
kw_en = yake.KeywordExtractor(lan="en", n=3, top=20)
keywords_en = kw_en.extract_keywords(text)

# Extract with life sciences stopwords
kw_ls = yake.KeywordExtractor(stopwords=ls_stops, n=3, top=20)
keywords_ls = kw_ls.extract_keywords(text)

# Compare results
print("English stopwords:")
for kw, score in keywords_en:
    print(f"  {score:.4f}  {kw}")

print("\nLife sciences stopwords:")
for kw, score in keywords_ls:
    print(f"  {score:.4f}  {kw}")
Batch Processing
python
import yake
import os

# Initialize extractor
kw_extractor = yake.KeywordExtractor(lan="en", n=3, top=15)

# Process multiple files
results = {}
for filename in os.listdir('/mnt/user-data/uploads'):
    if filename.endswith('.txt'):
        with open(f'/mnt/user-data/uploads/{filename}', 'r') as f:
            text = f.read()
        
        keywords = kw_extractor.extract_keywords(text)
        results[filename] = keywords

# Output results
for filename, keywords in results.items():
    print(f"\n{filename}:")
    for kw, score in keywords[:10]:  # Top 10
        print(f"  {score:.4f}  {kw}")
Multilingual Extraction
python
import yake

# French document
with open('/mnt/user-data/uploads/article_fr.txt', 'r') as f:
    french_text = f.read()

# Extract with French stopwords
kw_fr = yake.KeywordExtractor(lan="fr", n=3, top=20)
keywords_fr = kw_fr.extract_keywords(french_text)

print("Mots-clés (French):")
for kw, score in keywords_fr:
    print(f"  {score:.4f}  {kw}")

# German document
with open('/mnt/user-data/uploads/artikel_de.txt', 'r') as f:
    german_text = f.read()

# Extract with German stopwords
kw_de = yake.KeywordExtractor(lan="de", n=3, top=20)
keywords_de = kw_de.extract_keywords(german_text)

print("\nSchlüsselwörter (German):")
for kw, score in keywords_de:
    print(f"  {score:.4f}  {kw}")

Output Formats

Plain Text
python
for kw, score in keywords:
    print(f"{kw}: {score:.4f}")
CSV
python
import csv

with open('/mnt/user-data/outputs/keywords.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['Keyword', 'Score'])
    writer.writerows(keywords)
JSON
python
import json

output = [{"keyword": kw, "score": score} for kw, score in keywords]
with open('/mnt/user-data/outputs/keywords.json', 'w') as f:
    json.dump(output, f, indent=2)

Notes

  • Lower scores indicate more important keywords
  • YAKE is unsupervised - no training data required
  • Supports 34 languages - built-in stopwords for Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Estonian, Farsi, Finnish, French, German, Greek, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Lithuanian, Latvian, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, and more
  • Optimal n-gram size is 2 or 3 for most use cases
  • For longer technical phrases (4+ words), consider post-processing or ontology matching
  • Always specify full venv path: /home/claude/yake-venv/bin/python

Troubleshooting

Import errors: Verify venv installation

bash
/home/claude/yake-venv/bin/python -c "import yake; print(yake.__version__)"

Empty results: Check text length (YAKE needs sufficient content, typically 100+ words)

Poor quality keywords: Adjust parameters:

  • Increase dedupLim for more aggressive deduplication
  • Try domain-specific stopwords
  • Increase top to see more candidates

Generic terms appearing: Add custom stopwords for your domain:

python
with open('assets/stopwords_ls.txt', 'r') as f:
    stops = set(line.strip().lower() for line in f)

# Add domain-specific terms
stops.update(['term1', 'term2', 'term3'])

kw_extractor = yake.KeywordExtractor(stopwords=stops, n=3, top=20)

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (assets) in extracting-keywords of oaustegard/claude-skills.

  • SKILL.md
  • README.md
  • assets/stopwords_ai.txt
  • assets/stopwords_ls.txt

Open the folder on GitHubat commit 90b0f1b

Compare with similar skills

Extracting Keywords next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extracting Keywords compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extracting Keywords this skilloaustegard/claude-skills150—~2.9kAutomated safety check: PassMIT
News Aggregator Skillcclank/news-aggregator-skill1.3k—~2.1kAutomated safety check: PassNone
AI Daily Newsgeekjourneyx/ai-daily-skill235—~2.3kAutomated safety check: PassNone
Vss Search ArchiveNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~3.3kAutomated safety check: PassApache-2.0
AnalyzeriBigQiang/feedgrab614—~1kAutomated safety check: PassMIT
Reportmicrosoft/data-formulator18k—~1.5kAutomated safety check: PassMIT

Similar skills

  • News Aggregator Skill

    cclank/news-aggregator-skill

    Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…

    1.3k GitHub stars~2.1k tokensUpdated 4 mo ago
    Writing & ContentAuto-check passed
  • AI Daily News

    geekjourneyx/ai-daily-skill

    Fetches AI news from smol.ai RSS and generates structured markdown with intelligent summarization and categorization.

    235 GitHub stars~2.3k tokensUpdated today
    Writing & ContentAuto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…

    1.9k GitHub stars~3.3k tokensUpdated today
    Writing & ContentAuto-check passed
  • Analyzer

    iBigQiang/feedgrab

    Content Analyzer — any content (URL, text, transcript) into structured analysis report with actionable insights.

    614 GitHub stars~1k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed
  • Report

    microsoft/data-formulator

    Official

    Turn an exploration (threads, findings, charts) into a single Markdown report — note, blog post, executive summary, KPI dashboard, slide brief, or multi-section analytical report, with embedded…

    18k GitHub stars~1.5k tokensUpdated 2 days ago
    Writing & ContentAuto-check passed
  • Tldr

    earlyaidopters/claudeclaw

    Summarize the current conversation into a TLDR note and save it to your notes folder.

    173 GitHub stars~953 tokensUpdated 5 mo ago
    Writing & ContentAuto-check passed

More from oaustegard/claude-skills

All 67 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Deciding With Confidence

    oaustegard/claude-skills

    Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

    150 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Extracting Keywords

What does Extracting Keywords do?

Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese). Extracting Keywords is an agent skill from oaustegard/claude-skills. Extract keywords from documents using YAKE algorithm with support for 34 languages (Arabic to Chinese).

When should I use Extracting Keywords?

Extracting Keywords fits situations like: users request keyword extraction; topic identification; content summarization; document analysis.

How do I install Extracting Keywords in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill extracting-keywords -a claude-code`. Or copy the skill folder (extracting-keywords in oaustegard/claude-skills) into .claude/skills/extracting-keywords in your project. Claude Code loads it when a task matches its description.

How do I install Extracting Keywords in Codex?

Run `npx skills add oaustegard/claude-skills --skill extracting-keywords -a codex`. Or copy the skill folder (extracting-keywords in oaustegard/claude-skills) into .agents/skills/extracting-keywords in your project. Codex loads it when a task matches its description.

Can I use Extracting Keywords in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill extracting-keywords -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extracting-keywords, .gemini/skills/extracting-keywords, .github/skills/extracting-keywords and .opencode/skills/extracting-keywords in your project.

What does Extracting Keywords need to run?

Going by SKILL.md and its folder, Extracting Keywords needs the command-line tools its instructions call (uv and python). Our summary lists: Python 3.

Does Extracting Keywords access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Extracting Keywords safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extracting Keywords use?

Extracting Keywords is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extracting Keywords use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extracting Keywords?

Skills that share tags, products or a category with Extracting Keywords: News Aggregator Skill (cclank/news-aggregator-skill, 1.3k stars), AI Daily News (geekjourneyx/ai-daily-skill, 235 stars), Vss Search Archive (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars) and Analyzer (iBigQiang/feedgrab, 614 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extracting Keywords?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 9, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.