Hugging Face Tokenizers
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
Design and validate LLM-based text classification. An agent skill from majiayu000/claude-skill-registry.
$ npx skills add majiayu000/claude-skill-registry --skill text-classification -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install majiayu000/claude-skill-registry text-classification --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-llm/text-classification .claude/skills/text-classification && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "text-classification" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classification into .claude/skills/text-classification/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-classification", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classificationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add majiayu000/claude-skill-registry --skill text-classification -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install majiayu000/claude-skill-registry text-classification --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ai-llm/text-classification .agents/skills/text-classification && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "text-classification" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classification into .agents/skills/text-classification/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-classification", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add majiayu000/claude-skill-registry --skill text-classification -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install majiayu000/claude-skill-registry text-classification --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ai-llm/text-classification .cursor/skills/text-classification && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "text-classification" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classification into .cursor/skills/text-classification/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-classification", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/majiayu000/claude-skill-registry.git --path skills/ai-llm/text-classification--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add majiayu000/claude-skill-registry --skill text-classification -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install majiayu000/claude-skill-registry text-classification --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ai-llm/text-classification .gemini/skills/text-classification && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "text-classification" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classification into .gemini/skills/text-classification/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-classification", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install majiayu000/claude-skill-registry text-classificationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add majiayu000/claude-skill-registry --skill text-classification -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ai-llm/text-classification .github/skills/text-classification && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "text-classification" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classification into .github/skills/text-classification/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-classification", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add majiayu000/claude-skill-registry --skill text-classification -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install majiayu000/claude-skill-registry text-classification --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ai-llm/text-classification .opencode/skills/text-classification && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "text-classification" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-llm/text-classification into .opencode/skills/text-classification/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-classification", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
text-classificationDesign and validate LLM-based text classification. An agent skill from majiayu000/claude-skill-registry.
Text Classification is an agent skill from majiayu000/claude-skill-registry. Design and validate LLM-based text classification. Use for codebooks, prompts, validation samples, agreement statistics, and reporting model-coded data.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).
It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Text Classification loads about 4.8k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 2,540 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 2,540 words, ~4,809 tokens.
.claude/skills/text-classification/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.none_of_above or uncodeable) for responses that are too vague, too short, or off-topic. Define this category as precisely as the substantive codes (Halterman & Keith 2025).reference/example-codebook-and-prompt.md.Follow the decision framework from Chae & Davidson (2025), which maps document characteristics and available resources to the appropriate approach:
Zero-shot prompting: Use when classifying short documents with a large decoder model (GPT-4o, Llama3-70B+) and no labeled training data. Best for rapid prototyping and tasks where constructs are well-defined. GPT-4o achieves the best zero-shot performance across tasks (Chae & Davidson 2025).
Few-shot prompting: Add labeled examples to the prompt. Results are inconsistent — adding examples helps some models but degrades others (Chae & Davidson 2025). Always compare few-shot against zero-shot on a held-out sample before committing. Select diverse examples covering edge cases, not just prototypical instances.
Fine-tuning: Train a model on labeled data. Effective with as few as 100 hand-coded examples for smaller models (Chae & Davidson 2025). Fine-tuned smaller models (Llama3-8B, GPT-3 Davinci) can match GPT-4o zero-shot performance. Prefer this when you have labeled data and need cost-effective classification at scale.
Instruction-tuning: Combine detailed prompting with fine-tuning on paired instruction-output examples. Most powerful regime for complex tasks — instruction-tuned Llama3-70B surpasses GPT-4o zero-shot on stance detection (Chae & Davidson 2025). Requires more technical infrastructure but yields the highest accuracy.
Encoder-only fine-tuning: A distinct fourth regime often omitted from generative-LLM discussions. Fine-tuning a smaller encoder-only model (BERT, DeBERTa, SBERT; ~86–110M parameters, personal-computer hardware) on modest labeled data can match or exceed zero-shot generative LLMs on many classification tasks at a fraction of the cost and with fully reproducible (deterministic) output (Chae & Davidson 2025, Table 1; Ziems et al. 2024 find fine-tuned RoBERTa rarely under-performs larger generative models across 20 tasks). Prefer encoder fine-tuning when the label set is fixed, labeled data exists, and reproducibility matters more than generative flexibility.
When resources permit, test multiple regimes on the same pilot sample and select based on empirical performance, not assumptions.
gpt-4o-2024-08-06), not the model family name. Commercial models are modified or deprecated without notice — GPT-3 was withdrawn from OpenAI's API entirely (Barrie, Palmer & Spirling 2025; Chae & Davidson 2025)."Code this response:\n\n{text}").pre-registration-writing.methods-reporting. When the underlying category set is not fixed in advance and discovery of categories is itself the goal, unsupervised approaches may be more appropriate — see topic-modeling.© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/ai-llm/text-classification of majiayu000/claude-skill-registry.
Open the folder on GitHubat commit 000116a
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.
Text Classification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Text Classification this skillmajiayu000/claude-skill-registry | 666 | 1 repos | ~4.8k | Automated safety check: Pass | MIT | |
| Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~3.4k | Automated safety check: Pass | MIT | |
| OpenMed Model Card Writermaziyarpanahi/openmed | 5.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel | 1.3k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Andrej KarpathyK-Dense-AI/mimeo | 282 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Comparetaishi-i/awesome-japanese-nlp-resources | 1k | — | ~4.1k | Automated safety check: Notes | CC0-1.0 |
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
maziyarpanahi/openmed
Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
K-Dense-AI/mimeo
Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).
taishi-i/awesome-japanese-nlp-resources
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
Orchestra-Research/AI-Research-SKILLs
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
majiayu000/claude-skill-registry
Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.
majiayu000/claude-skill-registry
Neural search via Exa MCP for web, code, and company research.
majiayu000/claude-skill-registry
Unified media generation via fal.ai MCP — image, video, and audio.
majiayu000/claude-skill-registry
Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.
majiayu000/claude-skill-registry
Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.
majiayu000/claude-skill-registry
Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.
Categories
Design and validate LLM-based text classification. An agent skill from majiayu000/claude-skill-registry. Text Classification is an agent skill from majiayu000/claude-skill-registry. Design and validate LLM-based text classification.
Text Classification fits situations like: validation samples; agreement statistics; reporting model-coded data.
Run `npx skills add majiayu000/claude-skill-registry --skill text-classification -a claude-code`. Or copy the skill folder (skills/ai-llm/text-classification in majiayu000/claude-skill-registry) into .claude/skills/text-classification in your project. Claude Code loads it when a task matches its description.
Run `npx skills add majiayu000/claude-skill-registry --skill text-classification -a codex`. Or copy the skill folder (skills/ai-llm/text-classification in majiayu000/claude-skill-registry) into .agents/skills/text-classification in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill text-classification -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/text-classification, .gemini/skills/text-classification, .github/skills/text-classification and .opencode/skills/text-classification in your project.
SKILL.md names no scripts, command-line tools or credentials: Text Classification is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Text Classification is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Text Classification: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and Andrej Karpathy (K-Dense-AI/mimeo, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.
Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.