Agent skill

Extracting Keywords

by hashgraph-online in hashgraph-online/awesome-codex-plugins

A skill your agent uses when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Extracting Keywords

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-keywords -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins extracting-keywords --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-keywords .claude/skills/extracting-keywords && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extracting-keywords
GitHub stars
1.3k
Token cost
~1.5k tokens
SKILL.md length
475 words
Files
1
Skills in repo
714
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search.

  • Extracting keywords (YAKE/RAKE) from documents — and
  • SKILL.md covers Keywords (YAKE / RAKE), Language detection, Embeddings (embed command) and Programmatic access, plus 1 more section
  • Calls jq; needs OPENAI_API_KEY
  • Detecting document language

What it does

Extracting Keywords is an agent skill from hashgraph-online/awesome-codex-plugins. Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), --detect-language, and the standalone embed command with real flags.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • Extracting keywords (YAKE/RAKE) from documents — and
  • Detecting document language
  • Generating embeddings for RAG and search

Example prompts

  • “/extracting-keywords”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 9e7b281. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extracting Keywords loads about 1.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 475 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 9e7b281, republished under its Apache-2.0 licence (© hashgraph-online). 475 words, ~1,520 tokens.

Download SKILL.mdSave it as .claude/skills/extracting-keywords/SKILL.md (or your agent's skills folder).
name
extracting-keywords
description
Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Covers the keyword config (and its feature gating), `--detect-language`, and the standalone `embed` command with real flags.

Extracting keywords, language, and embeddings

Use this for the enrichment surface around extraction: statistical keyword extraction, language detection, and vector embeddings. Keywords and language detection ride along with extraction and land on the result; embeddings are produced by a dedicated embed command.

Keywords (YAKE / RAKE)

Keyword extraction is configured via the [keywords] config block (or inline JSON) — there is no single --keywords CLI flag. When enabled, extracted keywords appear on result.keywords. Two algorithms are available:

  • YAKE ("yake") — statistical, unsupervised single-document extraction. Good general default.
  • RAKE ("rake") — co-occurrence / phrase-based. Favors multi-word key phrases.

Feature-gated: keyword extraction requires the CLI to be built with the keywords-yake and/or keywords-rake Cargo features (both are in the default/full build). If the CLI was built without them, the [keywords] config block is silently ignored — result.keywords simply stays empty rather than erroring. The "yake" algorithm needs keywords-yake; "rake" needs keywords-rake.

Enable via inline JSON on the CLI:

bash
kreuzberg extract paper.pdf --format json \
  --config-json '{"keywords":{"algorithm":"yake","max_keywords":15,"language":"en"}}' \
  | jq '.keywords'

Or in a config file:

toml
[keywords]
algorithm = "rake"       # "yake" or "rake"
max_keywords = 10        # default 10
min_score = 0.0          # filter below this score (ranges differ per algorithm)
ngram_range = [1, 3]     # unigrams..trigrams (default)
language = "en"          # stopword language; omit to skip stopword filtering
bash
kreuzberg extract report.pdf --config kreuzberg.toml --format json | jq '.keywords'

Field notes:

  • max_keywords caps how many keywords are returned (default 10).
  • min_score filters low-scoring keywords; note YAKE scores are lower-is-better while RAKE scores are higher-is-better, so a single threshold behaves differently per algorithm.
  • ngram_range is [min, max]: [1,1] unigrams only, [1,2] adds bigrams, [1,3] (default) adds trigrams.
  • language enables stopword filtering for that language; omit it to disable stopword filtering entirely.

Language detection

Language detection is a real CLI flag: --detect-language. Detected languages appear on result.detected_languages:

bash
kreuzberg extract multilingual.pdf --detect-language true --format json \
  | jq '.detected_languages'

In a config file it lives under [language_detection]:

toml
[language_detection]
enabled = true
min_confidence = 0.8
detect_multiple = false

The CLI flag enables detection with min_confidence = 0.8 and single-language mode; use the config block to detect multiple languages or tune confidence.

Show full SKILL.md (214 more words)Show less

Embeddings (embed command)

The standalone embed command produces vector embeddings for text from --text (repeatable) or stdin. It does not run extraction — pipe extracted content in if you want document embeddings.

bash
# Local ONNX preset model (default provider)
kreuzberg embed --text "first passage" --text "second passage" --preset balanced

# Embed extracted document text
kreuzberg extract report.pdf | kreuzberg embed --preset quality

Presets for the local provider: fast, balanced (default), quality, multilingual. Output defaults to JSON (--format json).

--provider selects the embedding source:

ProviderFlagNotes
local--preset <fast|balanced|quality|multilingual>Default. ONNX model, no API key.
llm--model <id> --api-key <key>liter-llm routing, e.g. openai/text-embedding-3-small.
plugin--plugin <name>A backend pre-registered in-process via the plugin API.
bash
# Provider-hosted embeddings via an LLM
kreuzberg embed --text "query text" \
  --provider llm --model openai/text-embedding-3-small --api-key "$OPENAI_API_KEY"

Local embedding presets must be downloaded first if not cached. Pre-warm them with the cache command:

bash
kreuzberg cache warm --embedding-model balanced   # one preset
kreuzberg cache warm --all-embeddings             # all four presets

Programmatic access

Keywords and detected languages live on the extraction result:

python
from kreuzberg import extract_file_sync, ExtractionConfig

result = extract_file_sync(
    "paper.pdf",
    config=ExtractionConfig(),  # configure keywords/language_detection on the config
)
print(result.keywords)             # extracted keywords (when enabled)
print(result.detected_languages)   # detected languages (when enabled)

See references/python-api.md and references/configuration.md in the sibling kreuzberg skill for the keyword / language-detection config classes and the embedding presets.

Common pitfalls

  • No --keywords flag — keyword extraction is config-only. Use --config-json '{"keywords":{...}}' or a [keywords] config block.
  • min_score direction — lower is better for YAKE, higher is better for RAKE; pick the threshold to match the algorithm.
  • Embeddings ≠ extraction — embed only takes raw text. Pipe kreuzberg extract output into it for document vectors.
  • Cold embedding models — first local run downloads the preset; run kreuzberg cache warm --all-embeddings to pre-populate.

See references/advanced-features.md for the embeddings pipeline and references/cli-reference.md for the embed and cache warm flag sets.

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-keywords of hashgraph-online/awesome-codex-plugins.

Open the folder on GitHubat commit 9e7b281

Compare with similar skills

Extracting Keywords next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extracting Keywords compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extracting Keywords this skillhashgraph-online/awesome-codex-plugins1.3k—~1.5kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Mashup Mods

    rehan-remade/universal-modder

    Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.

    5.8k GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 714 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.3k GitHub stars~922 tokensUpdated today
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.3k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.3k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.3k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Calle

    hashgraph-online/awesome-codex-plugins

    Use CALL-E from Codex through the calle CLI. An agent skill from hashgraph-online/awesome-codex-plugins.

    1.3k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.3k GitHub stars~618 tokensUpdated today
    Auto-check passed

Questions about Extracting Keywords

What does Extracting Keywords do?

A skill your agent uses when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search. Extracting Keywords is an agent skill from hashgraph-online/awesome-codex-plugins. Use when extracting keywords (YAKE/RAKE) from documents — and, secondarily, when detecting document language or generating embeddings for RAG and search.

When should I use Extracting Keywords?

Extracting Keywords fits situations like: extracting keywords (YAKE/RAKE) from documents — and; detecting document language; generating embeddings for RAG and search.

How do I install Extracting Keywords in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-keywords -a claude-code`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-keywords in hashgraph-online/awesome-codex-plugins) into .claude/skills/extracting-keywords in your project. Claude Code loads it when a task matches its description.

How do I install Extracting Keywords in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-keywords -a codex`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/extracting-keywords in hashgraph-online/awesome-codex-plugins) into .agents/skills/extracting-keywords in your project. Codex loads it when a task matches its description.

Can I use Extracting Keywords in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill extracting-keywords -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extracting-keywords, .gemini/skills/extracting-keywords, .github/skills/extracting-keywords and .opencode/skills/extracting-keywords in your project.

What does Extracting Keywords need to run?

Going by SKILL.md and its folder, Extracting Keywords needs the command-line tools its instructions call (jq) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Extracting Keywords access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Extracting Keywords safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extracting Keywords use?

Extracting Keywords is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extracting Keywords use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extracting Keywords?

Skills that share tags, products or a category with Extracting Keywords: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extracting Keywords?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,255 GitHub stars. The repository holds 714 skills in this directory. The repository was last updated on October 9, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.