Agent skill

Text Analysis Basic

by Drchronx in Drchronx/ai-agent-research-starter-kit

Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity.

Custom licenceAuto-check passedAI & LLM Engineering

Install Text Analysis Basic

skills CLI
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill text-analysis-basic -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Drchronx/ai-agent-research-starter-kit text-analysis-basic --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'综合学术部署包/05_文本挖掘与NLP Skills/text-analysis-basic' .claude/skills/text-analysis-basic && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
text-analysis-basic
GitHub stars
134
Token cost
~737 tokens
SKILL.md length
311 words
Files
17 (incl. scripts, references)
Skills in repo
53
Repo updated
First seen
Licence
Custom licence

At a glance

Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity.

  • Works in 5 steps: Decide whether the user's request is… → If it is not covered, state that the… → If it is covered, write a short plan… → …
  • A user asks Codex to extract text from PDFs
  • SKILL.md covers Operating Rule, Capability Map, Input Expectations and Course Context
  • Runs Python scripts from its folder

What it does

Text Analysis Basic is an agent skill from Drchronx/ai-agent-research-starter-kit. Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity. Use when a user asks Codex to extract text from PDFs, segment Chinese text, count terms or sentences, create word clouds, vectorize text, train Word2Vec, encode embeddings, or compute text similarity from files or short inputs.

Its SKILL.md is about 740 tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts and reference files (for example `agents/openai.yaml`, `readme.md` and `references/embedding.md`).

It sits in AI & LLM Engineering, covering Natural language processing, Embeddings and PDF. The repository describes itself as: AI Agent 科研全流程教学包,能教学生从零部署 Codex、Claude Code、OpenClaw、Hermes 等 Agent,学会使用 Skills、飞书、AMiner、AI4Scholar、Zotero、Obsidian…

When your agent uses it

  • A user asks Codex to extract text from PDFs
  • Segment Chinese text
  • Create word clouds
  • Encode embeddings

Example prompts

  • “/text-analysis-basic”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Decide whether the user's request is covered by the capabilities below.
  2. If it is not covered, state that the skill cannot solve it and identify the missing capability.
  3. If it is covered, write a short plan before running commands. The plan must list
  4. Ask only for missing data paths or parameters that cannot be inferred safely.
  5. Execute the relevant script with command-line arguments. Do not write one-off Python code for the operation.

What it can do on your machine

Read from SKILL.md and the folder at commit aab1133. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Text Analysis Basic loads about 737 tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 311 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~737
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 311 words (~737 tokens).

“Use this skill as a command-line workflow, not as a code-writing prompt.”

— opening of SKILL.md by Drchronx, Custom licence
name
text-analysis-basic

Read the full SKILL.md on GitHub

Files

SKILL.md and 16 other files (scripts, references) in 综合学术部署包/05_文本挖掘与NLP Skills/text-analysis-basic of Drchronx/ai-agent-research-starter-kit.

  • SKILL.md
  • agents/openai.yaml
  • readme.md
  • references/embedding.md
  • references/frequency.md
  • references/jieba-tokenization.md
  • references/pdf-extraction.md
  • references/similarity.md
  • references/tfidf.md
  • references/word2vec.md
  • references/wordcloud.md
  • scripts/frequencies.py
  • scripts/jieba_tokenize.py
  • scripts/pdf_extract.py
  • scripts/similarity.py
  • scripts/vectorize.py
  • scripts/wordcloud_generate.py

Open the folder on GitHubat commit aab1133

Compare with similar skills

Text Analysis Basic next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Text Analysis Basic compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Text Analysis Basic this skillDrchronx/ai-agent-research-starter-kit134—~737Automated safety check: PassCustom licence
Scholar Computejoshzyj/open-scholar-skill167—~15kAutomated safety check: PassCustom licence
I3brycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.8kAutomated safety check: PassCustom licence
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.6kAutomated safety check: PassMIT
Researchtaishi-i/awesome-japanese-nlp-resources1k—~3.5kAutomated safety check: NotesCC0-1.0

Similar skills

  • Scholar Compute

    joshzyj/open-scholar-skill

    Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…

    167 GitHub stars~15k tokensUpdated 19 days ago
    AI & LLM EngineeringAuto-check passed
  • I3

    brycewang-stanford/Auto-Empirical-Research-Skills

    RAG Builder with Parallel Document Processing Vector database construction with local embeddings (zero cost) Handles PDF download, text extraction, chunking, and vector database creation Absorbed B5…

    4.5k GitHub stars~1.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 3 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Search

    taishi-i/awesome-japanese-nlp-resources

    Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face).

    1k GitHub stars~4.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes

More from Drchronx/ai-agent-research-starter-kit

All 53 skills in this repo
  • Literature Reviewer Skill

    Drchronx/ai-agent-research-starter-kit

    Build high-quality literature reviews from a research topic using a 10-phase workflow.

    134 GitHub stars~2.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Scenario Experiment Analysis

    Drchronx/ai-agent-research-starter-kit

    Analyze scenario/vignette experiment datasets for behavioral research.

    134 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Scenario Experiment Benchmark Mining

    Drchronx/ai-agent-research-starter-kit

    Mine and synthesize real top-journal scenario/vignette experiment patterns for behavioral research.

    134 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Web Browsing

    Drchronx/ai-agent-research-starter-kit

    Browse and summarize websites, extract content from URLs, search the web for information.

    134 GitHub stars~593 tokensUpdated 4 mo ago
    Auto-check passed
  • Multi Source Data Integration Extraction

    Drchronx/ai-agent-research-starter-kit

    Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

    134 GitHub stars~671 tokensUpdated 4 mo ago
    Auto-check passed
  • Scientific Slides

    Drchronx/ai-agent-research-starter-kit

    Build slide decks and presentations for research talks using Nano Banana Pro AI.

    134 GitHub stars~9.8k tokensUpdated 4 mo ago
    Auto-check: notes

Questions about Text Analysis Basic

What does Text Analysis Basic do?

Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity. Text Analysis Basic is an agent skill from Drchronx/ai-agent-research-starter-kit. Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity.

When should I use Text Analysis Basic?

Text Analysis Basic fits situations like: A user asks Codex to extract text from PDFs; segment Chinese text; create word clouds; encode embeddings.

How do I install Text Analysis Basic in Claude Code?

Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill text-analysis-basic -a claude-code`. Or copy the skill folder (综合学术部署包/05_文本挖掘与NLP Skills/text-analysis-basic in Drchronx/ai-agent-research-starter-kit) into .claude/skills/text-analysis-basic in your project. Claude Code loads it when a task matches its description.

How do I install Text Analysis Basic in Codex?

Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill text-analysis-basic -a codex`. Or copy the skill folder (综合学术部署包/05_文本挖掘与NLP Skills/text-analysis-basic in Drchronx/ai-agent-research-starter-kit) into .agents/skills/text-analysis-basic in your project. Codex loads it when a task matches its description.

Can I use Text Analysis Basic in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Drchronx/ai-agent-research-starter-kit --skill text-analysis-basic -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/text-analysis-basic, .gemini/skills/text-analysis-basic, .github/skills/text-analysis-basic and .opencode/skills/text-analysis-basic in your project.

What does Text Analysis Basic need to run?

Going by SKILL.md and its folder, Text Analysis Basic needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Text Analysis Basic access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Text Analysis Basic safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Text Analysis Basic use?

Text Analysis Basic has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Text Analysis Basic use?

About 737 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Text Analysis Basic?

Skills that share tags, products or a category with Text Analysis Basic: Scholar Compute (joshzyj/open-scholar-skill, 167 stars), I3 (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars) and Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Text Analysis Basic?

Drchronx (a GitHub user) maintains it in Drchronx/ai-agent-research-starter-kit, which has 134 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on May 19, 2026.

Source: Drchronx/ai-agent-research-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.