Agent skill

NLP Engineer

by majiayu000 in majiayu000/claude-skill-registry

Expert in Natural Language Processing, designing systems for text classification, NER, translation, and LLM integration using Hugging Face, spaCy, and LangChain.

MITAuto-check passedAI & LLM Engineering

Install NLP Engineer

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill nlp-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry nlp-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-ml/nlp-engineer-skill .claude/skills/nlp-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nlp-engineer
GitHub stars
666
Used in
1 other repo
Token cost
~833 tokens
SKILL.md length
306 words
Files
2
Skills in repo
1,273
Repo updated
First seen
Licence
MIT

At a glance

Expert in Natural Language Processing, designing systems for text classification, NER, translation, and LLM integration using Hugging Face, spaCy, and LangChain.

  • Works in 3 steps: Text Classification Pipeline → NER System → Embedding-Based Search
  • Building NLP pipelines
  • SKILL.md covers Purpose, When to Use, Quick Start and Decision Framework, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

NLP Engineer is an agent skill from majiayu000/claude-skill-registry. Expert in Natural Language Processing, designing systems for text classification, NER, translation, and LLM integration using Hugging Face, spaCy, and LangChain. Use when building NLP pipelines, text analysis, or LLM-powered features. Triggers include "NLP", "text classification", "NER", "named entity", "sentiment analysis", "spaCy", "Hugging Face", "transformers".

Its SKILL.md is about 830 tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in AI & LLM Engineering, covering Natural language processing. It works with Hugging Face, LangChain and Transformers. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Building NLP pipelines
  • LLM-powered features
  • Text classification
  • Sentiment analysis

Example prompts

  • “text classification”
  • “named entity”
  • “sentiment analysis”
  • “/nlp-engineer”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Text Classification Pipeline
  2. NER System
  3. Embedding-Based Search

What it can do on your machine

Read from SKILL.md and the folder at commit 2d14a69. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

NLP Engineer loads about 833 tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 306 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~833

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 2d14a69, republished under its MIT licence (© majiayu000). 306 words, ~833 tokens.

Download SKILL.mdSave it as .claude/skills/nlp-engineer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
nlp-engineer
description
Expert in Natural Language Processing, designing systems for text classification, NER, translation, and LLM integration using Hugging Face, spaCy, and LangChain. Use when building NLP pipelines, text analysis, or LLM-powered features. Triggers include "NLP", "text classification", "NER", "named entity", "sentiment analysis", "spaCy", "Hugging Face", "transformers".

NLP Engineer

Purpose

Provides expertise in Natural Language Processing systems design and implementation. Specializes in text classification, named entity recognition, sentiment analysis, and integrating modern LLMs using frameworks like Hugging Face, spaCy, and LangChain.

When to Use

  • Building text classification systems
  • Implementing named entity recognition (NER)
  • Creating sentiment analysis pipelines
  • Fine-tuning transformer models
  • Designing LLM-powered features
  • Implementing text preprocessing pipelines
  • Building search and retrieval systems
  • Creating text generation applications

Quick Start

Invoke this skill when:

  • Building NLP pipelines (classification, NER, sentiment)
  • Fine-tuning transformer models
  • Implementing text preprocessing
  • Integrating LLMs for text tasks
  • Designing semantic search systems

Do NOT invoke when:

  • RAG architecture design → use /ai-engineer
  • LLM prompt optimization → use /prompt-engineer
  • ML model deployment → use /mlops-engineer
  • General data processing → use /data-engineer

Decision Framework

NLP Task Type?
├── Classification
│   ├── Simple → Fine-tuned BERT/DistilBERT
│   └── Zero-shot → LLM with prompting
├── NER
│   ├── Standard entities → spaCy
│   └── Custom entities → Fine-tuned model
├── Generation
│   └── LLM (GPT, Claude, Llama)
└── Semantic Search
    └── Embeddings + Vector store

Core Workflows

1. Text Classification Pipeline
  1. Collect and label training data
  2. Preprocess text (tokenization, cleaning)
  3. Select base model (BERT, RoBERTa)
  4. Fine-tune on labeled dataset
  5. Evaluate with appropriate metrics
  6. Deploy with inference optimization
2. NER System
  1. Define entity types for domain
  2. Create labeled training data
  3. Choose framework (spaCy, Hugging Face)
  4. Train custom NER model
  5. Evaluate precision, recall, F1
  6. Integrate with post-processing rules
  1. Select embedding model (sentence-transformers)
  2. Generate embeddings for corpus
  3. Index in vector database
  4. Implement query embedding
  5. Add hybrid search (keyword + semantic)
  6. Tune similarity thresholds

Best Practices

  • Start with pretrained models, fine-tune as needed
  • Use domain-specific preprocessing
  • Evaluate with task-appropriate metrics
  • Consider inference latency for production
  • Implement proper text cleaning pipelines
  • Use batching for efficient inference

Anti-Patterns

Anti-PatternProblemCorrect Approach
Training from scratchWastes data and computeFine-tune pretrained
No preprocessingNoisy inputs hurt performanceClean and normalize text
Wrong metricsMisleading evaluationTask-appropriate metrics
Ignoring class imbalanceBiased predictionsBalance or weight classes
Overfitting to eval setPoor generalizationProper train/val/test splits

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-ml/nlp-engineer-skill of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 2d14a69

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

NLP Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

NLP Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
NLP Engineer this skillmajiayu000/claude-skill-registry6661 repos~833Automated safety check: PassMIT
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.6kAutomated safety check: PassMIT
Hugging Face Transformers Usagedavila7/claude-code-templates32k12 repos~1.2kAutomated safety check: PassMIT
Transformers.jshuggingface/skills11k1 repos~6.2kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k7 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 3 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Transformers Usage

    davila7/claude-code-templates

    Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

    32k GitHub starsUsed in 12 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Transformers.js

    huggingface/skills

    Official

    Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks.

    11k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 7 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed

More from majiayu000/claude-skill-registry

All 1,273 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Pyzotero

    majiayu000/claude-skill-registry

    Interact with Zotero reference management libraries using the pyzotero Python client.

    666 GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check: notes
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed

Questions about NLP Engineer

What does NLP Engineer do?

Expert in Natural Language Processing, designing systems for text classification, NER, translation, and LLM integration using Hugging Face, spaCy, and LangChain. NLP Engineer is an agent skill from majiayu000/claude-skill-registry. Expert in Natural Language Processing, designing systems for text classification, NER, translation, and LLM integration using Hugging Face, spaCy, and LangChain.

When should I use NLP Engineer?

NLP Engineer fits situations like: building NLP pipelines; LLM-powered features; text classification; sentiment analysis.

How do I install NLP Engineer in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill nlp-engineer -a claude-code`. Or copy the skill folder (skills/ai-ml/nlp-engineer-skill in majiayu000/claude-skill-registry) into .claude/skills/nlp-engineer in your project. Claude Code loads it when a task matches its description.

How do I install NLP Engineer in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill nlp-engineer -a codex`. Or copy the skill folder (skills/ai-ml/nlp-engineer-skill in majiayu000/claude-skill-registry) into .agents/skills/nlp-engineer in your project. Codex loads it when a task matches its description.

Can I use NLP Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill nlp-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nlp-engineer, .gemini/skills/nlp-engineer, .github/skills/nlp-engineer and .opencode/skills/nlp-engineer in your project.

What does NLP Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: NLP Engineer is instructions for the agent only.

Does NLP Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is NLP Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does NLP Engineer use?

NLP Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does NLP Engineer use?

About 833 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to NLP Engineer?

Skills that share tags, products or a category with NLP Engineer: Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Transformers Usage (davila7/claude-code-templates, 32k stars), Transformers.js (huggingface/skills, 11k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains NLP Engineer?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 1,273 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.