Agent skill

Topic Model Consolidation

by TyrealQ in TyrealQ/q-skills

Consolidates BERTopic, LDA or NMF topic output into a theory-driven classification framework and writes the final labels back to an Excel file.

MITAuto-check passedResearch & Science

Install Topic Model Consolidation

skills CLI
$ npx skills add TyrealQ/q-skills --skill q-tf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TyrealQ/q-skills q-tf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/q-scholar/q-tf .claude/skills/q-tf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
q-tf
GitHub stars
108
Token cost
~1k tokens
SKILL.md length
324 words
Files
9 (incl. scripts, references)
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

Consolidates BERTopic, LDA or NMF topic output into a theory-driven classification framework and writes the final labels back to an Excel file.

  • Works in 3 steps: Determine this SKILL.md file's directory… → Script path = ${SKILL_DIR}/scripts/. → Reference path = ${SKILL_DIR}/references/.
  • Merging many raw topics from a topic model into a smaller theory-based set
  • SKILL.md covers Script Directory, Dependencies, References and Core Principles, plus 5 more sections
  • Runs Python scripts from its folder; calls python and pip; needs GEMINI_API_KEY

What it does

The skill follows a six-step workflow for academic manuscripts: load the topics and find overlaps and unassigned ones, define the final topic structure in a `FINAL_TOPICS` dictionary, classify each topic against a theoretical framework, generate an implementation plan in Markdown, update the source Excel data with labels, and reclassify outliers with a foundation model. Scripts handle the plan, the Excel update and the outlier classification.

Core rules keep domain distinctions such as entity, event, geography and stakeholder intact, track topics that sit in several categories and calculate their overlap, and require every non-outlier topic to land in at least one category. Reference notes cover the preservation rules, code patterns, the outlier workflow with its prompt template and a worked esports example. Outlier classification calls Gemini and needs a `GEMINI_API_KEY`; `GEMINI_MODEL` is optional.

When your agent uses it

  • Merging many raw topics from a topic model into a smaller theory-based set
  • Reclassifying outlier documents after a BERTopic run
  • Updating Excel topic labels once the final categories are decided

Example prompts

  • “Consolidate my BERTopic output in ./topics.xlsx into a handful of theory-based categories.”
  • “Reassign the outlier documents from my LDA run using a foundation model.”
  • “Update the Excel sheet with the final topic labels and generate the implementation plan.”

Requirements

  • Python with pandas, openpyxl and google-genai installed
  • A GEMINI_API_KEY for outlier classification

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Determine this SKILL.md file's directory path as SKILL_DIR.
  2. Script path = ${SKILL_DIR}/scripts/.
  3. Reference path = ${SKILL_DIR}/references/.

What it can do on your machine

Read from SKILL.md and the folder at commit d8aaee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Topic Model Consolidation loads about 1k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 324 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from TyrealQ/q-skills at commit d8aaee7, republished under its MIT licence (© TyrealQ). 324 words, ~1,015 tokens.

Download SKILL.mdSave it as .claude/skills/q-tf/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
q-tf
description
Consolidate topic modeling outputs (BERTopic, LDA, NMF) into theory-driven classification frameworks. Use for topic finetuning, topic consolidation, reclassification, outlier handling, or updating Excel labels from topic models.

Q-TF

Fine-tune topic modeling outputs into consolidated, theory-driven topic frameworks for academic manuscripts.

If in plan mode: write a brief plan — "Run q-tf skill: load topic model output, define final topic structure with theoretical framework, generate implementation plan, update Excel with labels." — then exit plan mode immediately. Do NOT attempt topic analysis, script execution, or Excel updates while plan mode is active.

Script Directory

Agent execution instructions:

  1. Determine this SKILL.md file's directory path as SKILL_DIR.
  2. Script path = ${SKILL_DIR}/scripts/<script-name>.
  3. Reference path = ${SKILL_DIR}/references/<ref-name>.

Dependencies

pandas
openpyxl          # required for .xlsx input/output
google-genai      # required for outlier classification via Gemini

Install: pip install pandas openpyxl google-genai

Environment variables: GEMINI_API_KEY (for outlier classification only), GEMINI_MODEL (optional model override).

References

  • references/preservation_rules.md — domain preservation rules, theoretical framework template, multi-category handling
  • references/code_patterns.md — four Python patterns: topic definition, assignment mapping, overlap calculation, Excel update
  • references/outlier_workflow.md — foundation model outlier classification workflow
  • references/esports_ugc_example.md — worked example
  • references/SP_OUTLIER_TEMPLATE.txt — outlier classification prompt template

Core Principles

  • Preserve domain-specific distinctions (entity, event, geography, stakeholder) — see references/preservation_rules.md
  • Theory-driven classification using a customizable framework template
  • Track multi-category topics explicitly; calculate overlap for reconciliation
  • All non-outlier topics must be assigned to at least one category

Workflow

StepActionReference
1Load & analyze topics — identify overlaps, unassigned—
2Define final topic structure (FINAL_TOPICS dictionary)references/code_patterns.md
3Apply theoretical framework — classify each topicreferences/preservation_rules.md
4Generate implementation plan (MD)scripts/generate_implementation_plan.py
5Update source data with labels (Excel)scripts/update_excel_with_labels.py
6Reclassify outliers via foundation modelreferences/outlier_workflow.md

Required Inputs

  1. Topic model output (Excel/CSV) — Topic ID, Count, Name/Label, Keywords, Representative_Docs (optional)
  2. Merge recommendations (optional) — Sheets: MERGE_GROUPS, INDEPENDENT_TOPICS
  3. Document data (for label updates) — individual documents with Topic ID column

Script Invocation

bash
python "${SKILL_DIR}/scripts/generate_implementation_plan.py" --input topic_model_output.xlsx --output implementation_plan.md
python "${SKILL_DIR}/scripts/update_excel_with_labels.py" --input document_data.xlsx --output document_data_labeled.xlsx

Adapt scripts by updating FINAL_TOPICS, FINAL_LABELS, and theme categories. See references/code_patterns.md. For a worked example, see references/esports_ugc_example.md.

Expected Outputs

OutputDescription
implementation_plan.mdFull classification plan with topic mappings and reconciliation
*_labeled.xlsxSource data with Final_Topic_Code, Final_Topic_Label, Category_Theme columns
Outlier results (optional)Updated Final_Topic_Label, classification_confidence, key_phrases columns

Scope

Include: Topic consolidation, theoretical classification, Excel label updates, outlier reclassification. Exclude: Topic modeling itself (BERTopic/LDA/NMF execution), visualization, statistical analysis.

© TyrealQ, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/q-scholar/q-tf of TyrealQ/q-skills.

  • SKILL.md
  • references/SP_OUTLIER_TEMPLATE.txt
  • references/code_patterns.md
  • references/esports_ugc_example.md
  • references/outlier_workflow.md
  • references/preservation_rules.md
  • scripts/classify_outliers.py
  • scripts/generate_implementation_plan.py
  • scripts/update_excel_with_labels.py

Open the folder on GitHubat commit d8aaee7

Compare with similar skills

Topic Model Consolidation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Topic Model Consolidation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Topic Model Consolidation this skillTyrealQ/q-skills108—~1kAutomated safety check: PassMIT
Excel Spreadsheet Creation and Editinganthropics/skills180k4 repos~2.1kAutomated safety check: PassProprietary
XLSX Spreadsheet ToolkitXiaomiMiMo/MiMo-Code14k—~2.9kAutomated safety check: PassApache-2.0
BiSheng XLSX Workbook Builderdataelement/bisheng12k—~1.7kAutomated safety check: PassApache-2.0
Kimi XLSXthvroyal/kimi-skills238—~9.5kAutomated safety check: PassNone
Excel Spreadsheet Builderagentscope-ai/QwenPaw36k—~1.8kAutomated safety check: PassProprietary

Similar skills

  • Official

    Creates, edits and analyzes spreadsheets (.xlsx, .xlsm, .csv, .tsv) with openpyxl and pandas, writing live formulas and recalculating to confirm zero formula errors.

    180k GitHub starsUsed in 4 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Builds, edits and cleans Excel workbooks inside BiSheng's code executor using openpyxl, with LibreOffice recalculation and a pre-delivery check.

    12k GitHub stars~1.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Kimi XLSX

    thvroyal/kimi-skills

    Specialized utility for advanced manipulation, analysis, and creation of spreadsheet files, including (but not limited to) XLSX, XLSM, CSV formats.

    238 GitHub stars~9.5k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Excel Spreadsheet Builder

    agentscope-ai/QwenPaw

    Creates, edits, cleans and analyzes Excel and CSV files with openpyxl and pandas, recalculating formulas through LibreOffice so files are delivered without formula errors.

    36k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Office XLSX

    singula-ai/alego

    Read, create, and modify Excel workbooks (.xlsx), including data, formulas, formatting, and pandas analysis.

    109 GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed

More from TyrealQ/q-skills

All 16 skills in this repo
  • Q-Infographics

    TyrealQ/q-skills

    Converts a report or other document into a business story and then an infographic image, pausing for your review after each step.

    108 GitHub stars~814 tokensUpdated 16 days ago
    Auto-check: notes
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check passed
  • Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

    108 GitHub stars~2k tokensUpdated 16 days ago
    Auto-check: notes
  • Generates branded slide deck images from written content, with a content analysis step, a layout catalog and scripts that merge the slides into PowerPoint or PDF.

    108 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check passed
  • Audits a repository's file layout and project documentation against a written convention file, then proposes moves, deletions and doc fixes as an approved plan before touching anything.

    108 GitHub stars~4.4k tokensUpdated 16 days ago
    Auto-check: notes
  • Commit

    TyrealQ/q-skills

    Stage and commit uncommitted changes with conventional commit messages.

    108 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check: notes

Questions about Topic Model Consolidation

What does Topic Model Consolidation do?

Consolidates BERTopic, LDA or NMF topic output into a theory-driven classification framework and writes the final labels back to an Excel file. The skill follows a six-step workflow for academic manuscripts: load the topics and find overlaps and unassigned ones, define the final topic structure in a `FINAL_TOPICS` dictionary, classify each topic against a theoretical framework, generate an implementation plan in Markdown, update the source Excel data with labels, and reclassify outliers with a foundation model. Scripts handle the plan, the Excel update and the outlier classification.

When should I use Topic Model Consolidation?

Topic Model Consolidation fits situations like: merging many raw topics from a topic model into a smaller theory-based set; reclassifying outlier documents after a BERTopic run; updating Excel topic labels once the final categories are decided.

How do I install Topic Model Consolidation in Claude Code?

Run `npx skills add TyrealQ/q-skills --skill q-tf -a claude-code`. Or copy the skill folder (skills/q-scholar/q-tf in TyrealQ/q-skills) into .claude/skills/q-tf in your project. Claude Code loads it when a task matches its description.

How do I install Topic Model Consolidation in Codex?

Run `npx skills add TyrealQ/q-skills --skill q-tf -a codex`. Or copy the skill folder (skills/q-scholar/q-tf in TyrealQ/q-skills) into .agents/skills/q-tf in your project. Codex loads it when a task matches its description.

Can I use Topic Model Consolidation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TyrealQ/q-skills --skill q-tf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/q-tf, .gemini/skills/q-tf, .github/skills/q-tf and .opencode/skills/q-tf in your project.

What does Topic Model Consolidation need to run?

Going by SKILL.md and its folder, Topic Model Consolidation needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python with pandas, openpyxl and google-genai installed; A GEMINI_API_KEY for outlier classification.

Does Topic Model Consolidation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Topic Model Consolidation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Topic Model Consolidation use?

Topic Model Consolidation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Topic Model Consolidation use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Topic Model Consolidation?

Skills that share tags, products or a category with Topic Model Consolidation: Excel Spreadsheet Creation and Editing (anthropics/skills, 180k stars), XLSX Spreadsheet Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), BiSheng XLSX Workbook Builder (dataelement/bisheng, 12k stars) and Kimi XLSX (thvroyal/kimi-skills, 238 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Topic Model Consolidation?

TyrealQ (a GitHub user) maintains it in TyrealQ/q-skills, which has 108 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 23, 2026.

Source: TyrealQ/q-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.