Agent skill

Keyword Clustering

by Infrasity-Labs in Infrasity-Labs/dev-gtm-claude-skills

Cluster a list of keywords into topical groups with search intent labels, validated search volume, KD, and CPC data.

MITAuto-check passedDocuments & Office

Install Keyword Clustering

skills CLI
$ npx skills add Infrasity-Labs/dev-gtm-claude-skills --skill keyword-clustering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Infrasity-Labs/dev-gtm-claude-skills keyword-clustering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Infrasity-Labs/dev-gtm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/seo-skills/keyword-clustering .claude/skills/keyword-clustering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
keyword-clustering
GitHub stars
136
Token cost
~2k tokens
SKILL.md length
929 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Cluster a list of keywords into topical groups with search intent labels, validated search volume, KD, and CPC data.

  • Works in 5 steps: Input Handling → DataForSEO Validation → Clustering → …
  • A user provides a list of keywords (pasted
  • SKILL.md covers What This Skill Does, Step 1: Input Handling, Step 2: DataForSEO Validation and Step 3: Clustering, plus 3 more sections
  • Calls python

What it does

Keyword Clustering is an agent skill from Infrasity-Labs/dev-gtm-claude-skills. Cluster a list of keywords into topical groups with search intent labels, validated search volume, KD, and CPC data. Use this skill whenever a user provides a list of keywords (pasted, uploaded as CSV/Excel, or via Google Sheet URL) and asks to cluster, group, map, organize, or categorize them. Also trigger when a user says "keyword cluster", "cluster my keywords", "group these keywords", "keyword map", "keyword strategy from this list", "organize keywords by topic", or pastes or uploads a list of keywords and…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Keyword research and Excel spreadsheets. It works with Microsoft Excel and Google Sheets. The repository describes itself as: Open-source Claude skills for GEO, AI discoverability, and developer GTM workflows. Built for developer-focused companies that want their documentation to be found, parsed, and… The licence is MIT.

When your agent uses it

  • A user provides a list of keywords (pasted
  • Uploaded as CSV/Excel
  • Via Google Sheet URL) and asks to cluster
  • Categorize them

Example prompts

  • “keyword cluster”
  • “cluster my keywords”
  • “group these keywords”
  • “/keyword-clustering”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Input Handling
  2. DataForSEO Validation
  3. Clustering
  4. Build the Excel Output
  5. Save and Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit 02cfefb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Keyword Clustering loads about 2k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 929 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Infrasity-Labs/dev-gtm-claude-skills at commit 02cfefb, republished under its MIT licence (© Infrasity-Labs). 929 words, ~2,008 tokens.

Download SKILL.mdSave it as .claude/skills/keyword-clustering/SKILL.md (or your agent's skills folder).
name
keyword-clustering
description
Cluster a list of keywords into topical groups with search intent labels, validated search volume, KD, and CPC data. Use this skill whenever a user provides a list of keywords (pasted, uploaded as CSV/Excel, or via Google Sheet URL) and asks to cluster, group, map, organize, or categorize them. Also trigger when a user says "keyword cluster", "cluster my keywords", "group these keywords", "keyword map", "keyword strategy from this list", "organize keywords by topic", or pastes or uploads a list of keywords and wants structure from it. Always use this skill — do not attempt keyword clustering manually without following this workflow.

Keyword Clustering Skill

Turn a raw keyword list into a structured, validated cluster map delivered as a downloadable Excel file.


What This Skill Does

  1. Accepts keywords from any input format
  2. Validates each keyword via DataForSEO — fetches search volume, KD, CPC
  3. Filters out keywords with 0 or null search volume (dead keywords)
  4. Clusters the validated keywords by topic + search intent
  5. Outputs a formatted, downloadable Excel file

Step 1: Input Handling

Detect which input format the user provided and extract the keyword list.

Pasted list

If the user pasted keywords directly in chat (one per line, comma-separated, or numbered), extract each keyword into a clean array. Strip numbers, bullets, extra whitespace.

CSV or Excel upload

The file will be at /mnt/user-data/uploads/. Read it with pandas:

python
import pandas as pd

# CSV
df = pd.read_csv('/mnt/user-data/uploads/filename.csv')

# Excel
df = pd.read_excel('/mnt/user-data/uploads/filename.xlsx')

# Identify keyword column: look for columns named 'keyword', 'keywords', 'query', 'term', 'search term'
# If ambiguous, pick the first text column or ask the user
keyword_candidates = ['keyword', 'keywords', 'query', 'term', 'search term']
keyword_col = next((col for col in df.columns if col.lower() in keyword_candidates), df.columns[0])
keywords = df[keyword_col].dropna().tolist()
Google Sheet URL

Use web_fetch to fetch the sheet as CSV (append /export?format=csv to the base URL). Parse with pandas.

Cap at 500 keywords per run. If input exceeds 500, tell the user and process the first 500, or ask which subset to use.


Step 2: DataForSEO Validation

Use the dataforseo_labs_google_keyword_overview tool to validate all keywords and fetch metrics. Batch in groups of 100 to stay within API limits.

Required fields to extract per keyword:

  • search_volume — monthly searches (Google)
  • keyword_difficulty — KD score (0–100)

Filter rule: Drop any keyword where search_volume is 0, null, or missing. These are dead keywords.

Tell the user upfront: "Validating [N] keywords via DataForSEO. This may take a moment..."

After validation, report:

  • Total keywords submitted
  • Keywords that passed (have search volume)
  • Keywords that were dropped (zero/null volume) — list them so the user can see what was removed

If DataForSEO is unavailable or returns an error: Tell the user "DataForSEO validation failed — I'll cluster the full list but cannot validate search volume or provide KD/CPC data." Then proceed with the full unvalidated list and skip the volume/KD/CPC columns in the output.


Step 3: Clustering

Cluster the validated keywords only using a two-axis approach:

Axis 1: Topical Cluster

Group keywords by shared topic/theme. Use the root concept to name the cluster.

Rules:

  • Each cluster should have a 2–5 word descriptive name (e.g., "Email Marketing Tools", "Python Data Analysis")
  • Aim for clusters of 3–8 keywords. Prefer more clusters over fewer — if a topic can reasonably be split into two distinct subtopics, split it
  • Do NOT over-merge: keywords that target different angles, audiences, or content types should be in separate clusters even if they share a root word
  • If a topic is very broad, always create sub-clusters (use a parent + child naming convention, e.g., "SEO – On-Page", "SEO – Technical")
  • Assign each cluster a numeric Cluster ID (1, 2, 3…)
Axis 2: Search Intent

For each keyword, assign one of the four standard intent labels:

LabelMeaningSignal words
InformationalUser wants to learnwhat is, how to, guide, tutorial, definition, examples
NavigationalUser wants a specific site/brandbrand name + login/sign in/pricing
CommercialUser is comparing optionsbest, top, vs, review, alternative, comparison
TransactionalUser wants to act/buybuy, download, get, free trial, sign up, hire

When intent is ambiguous, pick the most likely based on the full keyword phrase. Do not leave intent blank.

Show full SKILL.md (408 more words)Show less
Clustering Logic

Use your understanding of keyword semantics. Group keywords that:

  • Share the same core topic or product area
  • Would logically appear in the same content piece or site section
  • Target the same audience stage

Do NOT group purely by shared word (e.g., don't put "best email marketing software" and "email marketing statistics" in the same cluster just because they share "email marketing" — one is Commercial, one is Informational, and they serve different pages).


Step 4: Build the Excel Output

Read the xlsx SKILL first if available. Use openpyxl for formatting.

Sheet 1: "Keyword Clusters" (main output)

Columns in order:

ColumnDescription
Cluster IDNumeric cluster number
Cluster NameDescriptive topic name
KeywordThe validated keyword
Search VolumeMonthly search volume from DataForSEO
KDKeyword difficulty (0–100)
IntentInformational / Navigational / Commercial / Transactional

Sorting: Sort by Cluster ID ascending, then by Search Volume descending within each cluster.

Formatting:

  • Row 1: Bold header, dark background (#1F2D40), white text, center-aligned
  • Alternate cluster groups with light row shading to visually separate clusters (#F2F2F2 every other cluster group)
  • Freeze top row
  • Auto-fit column widths (min 15, max 40)
  • Number format for Search Volume: #,##0
  • KD: plain integer
Sheet 2: "Cluster Summary"

One row per cluster:

ColumnDescription
Cluster IDNumber
Cluster NameName
# KeywordsCount of keywords in cluster
Avg Search VolumeAverage volume across cluster
Avg KDAverage KD
Dominant IntentMost common intent label in cluster
Top KeywordHighest-volume keyword in cluster

Sort by Avg Search Volume descending — highest-opportunity clusters first.

Sheet 3: "Dropped Keywords"

List all keywords removed at the validation step:

ColumnDescription
KeywordThe dropped keyword
Reason"Zero search volume" or "No data returned"

If no keywords were dropped, add a single row: "No keywords were dropped."


Step 5: Save and Deliver

python
# Save the workbook
output_path = '/mnt/user-data/outputs/keyword_clusters.xlsx'
wb.save(output_path)

Then run recalc if formulas are used:

bash
python scripts/recalc.py /mnt/user-data/outputs/keyword_clusters.xlsx

Use present_files to deliver the file to the user.

After presenting the file, give a short summary in chat:

  • Total keywords clustered
  • Number of clusters created
  • Number of keywords dropped
  • Top 3 clusters by average search volume

Edge Cases

  • Duplicate keywords: Deduplicate before validation. If duplicates exist, mention it.
  • Non-English keywords: DataForSEO supports multi-language. Cluster by semantic meaning. Note the language.
  • All keywords dropped: Tell the user all keywords had zero volume. Offer to cluster them anyway without validation.
  • Single keyword input: Tell the user clustering requires at least 10 keywords to be meaningful.
  • Very large input (200+ keywords): Warn that DataForSEO batching may take 2–3 minutes. Proceed automatically.

© Infrasity-Labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in seo-skills/keyword-clustering of Infrasity-Labs/dev-gtm-claude-skills.

Open the folder on GitHubat commit 02cfefb

Compare with similar skills

Keyword Clustering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Keyword Clustering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Keyword Clustering this skillInfrasity-Labs/dev-gtm-claude-skills136—~2kAutomated safety check: PassMIT
XLSXzzhonglei/GeoCode-Release189—~3.1kAutomated safety check: PassMIT
Spreadsheet Agentmastra-ai/mastra29k—~2.1kAutomated safety check: PassCustom licence
Sheets Artifactasgeirtj/system_prompts_leaks69k—~1.2kAutomated safety check: PassCC0-1.0
XLSXflonat/flonat-research146—~2.7kAutomated safety check: PassProprietary
Spreadsheet Formula Helpercomposio-community/awesome-codex-skills17k—~328Automated safety check: PassNone

Similar skills

  • XLSX

    zzhonglei/GeoCode-Release

    Create, edit, analyze, or convert Excel spreadsheets (.xlsx, .xlsm) where the workbook file is the primary deliverable.

    189 GitHub stars~3.1k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed
  • Spreadsheet Agent

    mastra-ai/mastra

    Authoring playbook for building agents that read or write tabular data — Google Sheets, Microsoft Excel, CSV, Airtable, Notion databases, or any spreadsheet.

    29k GitHub stars~2.1k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Sheets Artifact

    asgeirtj/system_prompts_leaks

    A skill your agent uses when creating, editing, or inspecting a spreadsheet or workbook (Excel, Google Sheets, or CSV), or when the task calls for a reusable budget, model, tracker, or structured…

    69k GitHub stars~1.2k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • XLSX

    flonat/flonat-research

    Create, read, edit, clean, format, chart, or convert spreadsheet files while preserving spreadsheet-native deliverables.

    146 GitHub stars~2.7k tokensUpdated 12 days ago
    Documents & OfficeAuto-check passed
  • Spreadsheet Formula Helper

    composio-community/awesome-codex-skills

    Write and debug spreadsheet formulas (Excel/Google Sheets), pivot tables, and array formulas; translate between dialects; use when users need working formulas with examples and edge-case checks.

    17k GitHub stars~328 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • CSV Formula Injection

    yaklang/hack-skills

    CSV/spreadsheet formula injection (DDE, Excel/LibreOffice, Google Sheets IMPORT).

    2.4k GitHub stars~1.1k tokensUpdated 28 days ago
    Documents & OfficeAuto-check passed

More from Infrasity-Labs/dev-gtm-claude-skills

All 26 skills in this repo
  • Brief Outline Generator

    Infrasity-Labs/dev-gtm-claude-skills

    Generates a fully structured SEO content outline (not a finished brief) and exports it as a formatted .docx Word document.

    136 GitHub stars~4k tokensUpdated 3 mo ago
    Auto-check passed
  • Content Brief

    Infrasity-Labs/dev-gtm-claude-skills

    Generates a fully structured SEO content brief for a target keyword and optionally pushes it to a Notion database.

    136 GitHub stars~2.8k tokensUpdated 3 mo ago
    Auto-check passed
  • API Docs Quality Report

    Infrasity-Labs/dev-gtm-claude-skills

    Audits any API documentation site by crawling every endpoint page and scoring each one across 5 checks: description quality, OpenAPI spec presence, body param descriptions, response codes, and…

    136 GitHub stars~2.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Docs Auditor

    Infrasity-Labs/dev-gtm-claude-skills

    Audits any developer documentation site across 33 checks in 7 categories and produces a scored report (out of 100) with Pass / Warn / Fail status per check.

    136 GitHub stars~3k tokensUpdated 3 mo ago
    Auto-check passed
  • Growth Report

    Infrasity-Labs/dev-gtm-claude-skills

    Generates a 3-month SEO performance HTML report for any domain using DataForSEO data.

    136 GitHub stars~4k tokensUpdated 3 mo ago
    Auto-check passed
  • Inbox

    Infrasity-Labs/dev-gtm-claude-skills

    Email triage system that handles both one-time setup and recurring triage in a single skill.

    136 GitHub stars~4.8k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Keyword Clustering

What does Keyword Clustering do?

Cluster a list of keywords into topical groups with search intent labels, validated search volume, KD, and CPC data. Keyword Clustering is an agent skill from Infrasity-Labs/dev-gtm-claude-skills. Cluster a list of keywords into topical groups with search intent labels, validated search volume, KD, and CPC data.

When should I use Keyword Clustering?

Keyword Clustering fits situations like: A user provides a list of keywords (pasted; uploaded as CSV/Excel; via Google Sheet URL) and asks to cluster; categorize them.

How do I install Keyword Clustering in Claude Code?

Run `npx skills add Infrasity-Labs/dev-gtm-claude-skills --skill keyword-clustering -a claude-code`. Or copy the skill folder (seo-skills/keyword-clustering in Infrasity-Labs/dev-gtm-claude-skills) into .claude/skills/keyword-clustering in your project. Claude Code loads it when a task matches its description.

How do I install Keyword Clustering in Codex?

Run `npx skills add Infrasity-Labs/dev-gtm-claude-skills --skill keyword-clustering -a codex`. Or copy the skill folder (seo-skills/keyword-clustering in Infrasity-Labs/dev-gtm-claude-skills) into .agents/skills/keyword-clustering in your project. Codex loads it when a task matches its description.

Can I use Keyword Clustering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Infrasity-Labs/dev-gtm-claude-skills --skill keyword-clustering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/keyword-clustering, .gemini/skills/keyword-clustering, .github/skills/keyword-clustering and .opencode/skills/keyword-clustering in your project.

What does Keyword Clustering need to run?

Going by SKILL.md and its folder, Keyword Clustering needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Keyword Clustering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Keyword Clustering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Keyword Clustering use?

Keyword Clustering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Keyword Clustering use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Keyword Clustering?

Skills that share tags, products or a category with Keyword Clustering: XLSX (zzhonglei/GeoCode-Release, 189 stars), Spreadsheet Agent (mastra-ai/mastra, 29k stars), Sheets Artifact (asgeirtj/system_prompts_leaks, 69k stars) and XLSX (flonat/flonat-research, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Keyword Clustering?

Infrasity-Labs (a GitHub user) maintains it in Infrasity-Labs/dev-gtm-claude-skills, which has 136 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on June 28, 2026.

Source: Infrasity-Labs/dev-gtm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.