Install the "nlp-engineering" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-ml/nlp-engineering into .claude/skills/nlp-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp-engineering", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "nlp-engineering" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-ml/nlp-engineering into .agents/skills/nlp-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp-engineering", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "nlp-engineering" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-ml/nlp-engineering into .cursor/skills/nlp-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp-engineering", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "nlp-engineering" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-ml/nlp-engineering into .gemini/skills/nlp-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp-engineering", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "nlp-engineering" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-ml/nlp-engineering into .github/skills/nlp-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp-engineering", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "nlp-engineering" agent skill from https://github.com/majiayu000/claude-skill-registry/tree/main/skills/ai-ml/nlp-engineering into .opencode/skills/nlp-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nlp-engineering", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
nlp-engineering
GitHub stars
666
Used in
1 other repo
Token cost
~4.4k tokens
SKILL.md length
1,350 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT
At a glance
A skill your agent uses when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization.
Works in 5 steps: Preprocessing is load-bearing - Garbage… → Match the model to the task - A… → Embed offline, search online -… → …
Building NLP pipelines
SKILL.md covers When to use this skill, Key principles, Core concepts and Common tasks, plus 4 more sections
Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
What it does
NLP Engineering is an agent skill from majiayu000/claude-skill-registry. Use this skill when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization. Triggers on text preprocessing, tokenization, embeddings, vector search, named entity recognition, sentiment analysis, text classification, summarization, and any task requiring natural language processing.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).
It sits in AI & LLM Engineering, covering Natural language processing and Embeddings. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.
When your agent uses it
Building NLP pipelines
Implementing text classification
Semantic search
Text preprocessing
Example prompts
“/nlp-engineering”
Requirements
Python 3
Workflow steps
5 steps, taken from the first numbered list in SKILL.md.
1Preprocessing is load-bearing - Garbage in, garbage out. Inconsistent
2Match the model to the task - A 66M-parameter sentence-transformer is
3Embed offline, search online - Pre-compute embeddings at index time.
4Chunk with overlap, not just length - Fixed-length chunking without
5Evaluate before you ship - Define offline metrics (precision@k, NDCG,
What it can do on your machine
Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
NLP Engineering loads about 4.4k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,350 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~87
When it runs· the whole SKILL.md, loaded when a task matches
~4.4k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/nlp-engineering/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
nlp-engineering
description
Use this skill when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization. Triggers on text preprocessing, tokenization, embeddings, vector search, named entity recognition, sentiment analysis, text classification, summarization, and any task requiring natural language processing.
When this skill is activated, always start your first response with the 🧢 emoji.
NLP Engineering
A practical framework for building production NLP systems. This skill covers the
full stack of natural language processing - from raw text ingestion through
tokenization, embedding, retrieval, classification, and generation - with an
emphasis on making the right architectural choices at each layer. Designed for
engineers who know Python and ML basics and need opinionated guidance on building
reliable, scalable text processing pipelines.
When to use this skill
Trigger this skill when the user:
Builds a text preprocessing or cleaning pipeline
Generates or stores embeddings for documents or queries
Implements semantic search or similarity-based retrieval
Classifies text into categories (sentiment, intent, topic, etc.)
Extracts named entities, relationships, or structured data from text
Summarizes long documents (extractive or abstractive)
Chunks documents for RAG (Retrieval-Augmented Generation) pipelines
Pure LLM prompt engineering or chain-of-thought with no text processing pipeline
Speech-to-text or image captioning (separate modalities with different toolchains)
Key principles
Preprocessing is load-bearing - Garbage in, garbage out. Inconsistent
casing, stray HTML, and unicode noise degrade every downstream component.
Invest in a reproducible cleaning pipeline before touching a model.
Match the model to the task - A 66M-parameter sentence-transformer is
often better than GPT-4 embeddings for a narrow domain retrieval task, and
100x cheaper. Pick the smallest model that hits your quality bar.
Embed offline, search online - Pre-compute embeddings at index time.
Doing embedding + vector search in the request path is an avoidable latency
sink. Only re-embed at write time (new docs) or on model upgrade.
Chunk with overlap, not just length - Fixed-length chunking without
overlap splits sentences at boundaries and degrades retrieval recall. Always
use a sliding window with 10-20% overlap and respect sentence boundaries.
Evaluate before you ship - Define offline metrics (precision@k, NDCG,
ROUGE, F1) before building. An NLP system without evals is a system you
cannot improve or regress-test.
Core concepts
Tokenization
Tokenization converts raw text into a sequence of tokens a model can process.
Modern models use subword tokenizers (BPE, WordPiece, SentencePiece) rather
than whitespace splitting, allowing them to handle out-of-vocabulary words
gracefully by decomposing them into known subword units.
Key considerations: token budget (LLMs have context windows), language coverage
(multilingual text needs a multilingual tokenizer), and domain vocabulary
(medical/legal/code text may have poor tokenization with general-purpose tokenizers).
Embeddings
An embedding is a dense vector representation of text that encodes semantic
meaning. Similar texts produce vectors with high cosine similarity. Embeddings
are the foundation of semantic search, clustering, and classification.
Two categories: encoding models (sentence-transformers, E5, BGE) are fast,
cheap, and purpose-built for retrieval. LLM embeddings (OpenAI
text-embedding-3, Cohere Embed) are convenient API calls but cost money per
token and introduce external latency.
Attention and transformers
Transformers process the full token sequence in parallel using self-attention,
letting every token attend to every other token. This gives transformer-based
models long-range context understanding that recurrent models lacked. For NLP
tasks, you almost never need to implement attention from scratch - use
HuggingFace transformers and fine-tune a pretrained checkpoint.
Vector similarity
Three distance metrics dominate:
Metric
Formula (conceptual)
Best for
Cosine similarity
angle between vectors
Normalized embeddings, most retrieval
Dot product
magnitude + angle
When vector magnitude carries information
Euclidean distance
straight-line distance
Rare; prefer cosine for NLP
Most vector stores (Pinecone, Weaviate, pgvector, FAISS) default to cosine or
dot product. Normalize your embeddings before storing them to make cosine and
dot product equivalent.
Common tasks
Text preprocessing pipeline
Build a reproducible cleaning pipeline before any modeling step. Apply in this
order: decode -> strip HTML -> normalize unicode -> lowercase -> remove noise ->
normalize whitespace.
python
import re
import unicodedata
from bs4 import BeautifulSoup
def preprocess(text: str, lowercase: bool = True) -> str:
# 1. Decode HTML entities and strip tags
text = BeautifulSoup(text, "html.parser").get_text(separator=" ")
# 2. Normalize unicode (NFD -> NFC, remove combining chars if needed)
text = unicodedata.normalize("NFC", text)
# 3. Lowercase
if lowercase:
text = text.lower()
# 4. Remove URLs, emails, special tokens
text = re.sub(r"https?://\S+|www\.\S+", " ", text)
text = re.sub(r"\S+@\S+\.\S+", " ", text)
# 5. Collapse whitespace
text = re.sub(r"\s+", " ", text).strip()
return text
# Usage
clean = preprocess("<p>Visit https://example.com for more info.</p>")
# -> "visit for more info."
Persist the preprocessing config (lowercase flag, regex patterns) alongside
your model so training and inference use identical transformations.
Generate embeddings
Use sentence-transformers for local, cost-free embeddings or the OpenAI API
for convenience. Always batch your calls.
python
# Option A: sentence-transformers (local, free, fast on GPU)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
documents = ["The quick brown fox", "Machine learning is fun", "NLP rocks"]
# encode() handles batching internally; show_progress_bar for large corpora
embeddings = model.encode(documents, normalize_embeddings=True, show_progress_bar=True)
# -> numpy array, shape (3, 384)
# Option B: OpenAI embeddings API
from openai import OpenAI
client = OpenAI()
def embed_batch(texts: list[str], model: str = "text-embedding-3-small") -> list[list[float]]:
# Strip newlines - they degrade embedding quality per OpenAI docs
texts = [t.replace("\n", " ") for t in texts]
response = client.embeddings.create(input=texts, model=model)
return [item.embedding for item in response.data]
Build semantic search
Index embeddings into a vector store and retrieve by cosine similarity at query
time. This example uses FAISS for local search and pgvector for PostgreSQL.
python
import numpy as np
import faiss
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
# --- Indexing ---
docs = ["Python is a programming language.", "The Eiffel Tower is in Paris.", ...]
doc_embeddings = model.encode(docs, normalize_embeddings=True).astype("float32")
# Inner product on normalized vectors = cosine similarity
index = faiss.IndexFlatIP(doc_embeddings.shape[1])
index.add(doc_embeddings)
# --- Retrieval ---
def search(query: str, top_k: int = 5) -> list[tuple[str, float]]:
q_emb = model.encode([query], normalize_embeddings=True).astype("float32")
scores, indices = index.search(q_emb, top_k)
return [(docs[i], float(scores[0][j])) for j, i in enumerate(indices[0])]
results = search("programming languages for data science")
# -> [("Python is a programming language.", 0.87), ...]
For production, use faiss.IndexIVFFlat (approximate, faster) or a managed
vector store (pgvector, Pinecone, Weaviate) rather than exact IndexFlatIP.
Text classification with transformers
Fine-tune a pretrained encoder for sequence classification. HuggingFace
transformers + datasets is the standard stack.
python
from datasets import Dataset
from transformers import (
AutoTokenizer,
AutoModelForSequenceClassification,
TrainingArguments,
Trainer,
)
import torch
MODEL_ID = "distilbert-base-uncased"
LABELS = ["negative", "neutral", "positive"]
id2label = {i: l for i, l in enumerate(LABELS)}
label2id = {l: i for i, l in enumerate(LABELS)}
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSequenceClassification.from_pretrained(
MODEL_ID, num_labels=len(LABELS), id2label=id2label, label2id=label2id
)
def tokenize(batch):
return tokenizer(batch["text"], truncation=True, padding="max_length", max_length=128)
# train_data: list of {"text": str, "label": int}
train_ds = Dataset.from_list(train_data).map(tokenize, batched=True)
args = TrainingArguments(
output_dir="./sentiment-model",
num_train_epochs=3,
per_device_train_batch_size=32,
evaluation_strategy="epoch",
save_strategy="best",
load_best_model_at_end=True,
)
trainer = Trainer(model=model, args=args, train_dataset=train_ds, eval_dataset=eval_ds)
trainer.train()
Use distilbert or roberta-base for most classification tasks. Only
escalate to larger models if the smaller ones underperform after fine-tuning.
NER pipeline
Use spaCy for fast rule-augmented NER or a HuggingFace token classification
model for custom entity types.
python
import spacy
from transformers import pipeline
# Option A: spaCy (fast, battle-tested for standard entities)
nlp = spacy.load("en_core_web_sm")
def extract_entities(text: str) -> list[dict]:
doc = nlp(text)
return [
{"text": ent.text, "label": ent.label_, "start": ent.start_char, "end": ent.end_char}
for ent in doc.ents
]
entities = extract_entities("Apple Inc. was founded by Steve Jobs in Cupertino.")
# -> [{"text": "Apple Inc.", "label": "ORG", ...}, {"text": "Steve Jobs", "label": "PERSON", ...}]
# Option B: HuggingFace token classification (custom entities, higher accuracy)
ner = pipeline(
"token-classification",
model="dslim/bert-base-NER",
aggregation_strategy="simple", # merges B-/I- tokens into spans
)
results = ner("OpenAI released GPT-4 in San Francisco.")
Extractive and abstractive summarization
Choose extractive for faithfulness (no hallucination risk) and abstractive for
fluency.
python
# --- Extractive: rank sentences by TF-IDF centrality ---
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
def extractive_summary(text: str, n_sentences: int = 3) -> str:
sentences = [s.strip() for s in text.split(".") if s.strip()]
tfidf = TfidfVectorizer().fit_transform(sentences)
sim_matrix = cosine_similarity(tfidf)
scores = sim_matrix.sum(axis=1)
top_indices = np.argsort(scores)[-n_sentences:][::-1]
return ". ".join(sentences[i] for i in sorted(top_indices)) + "."
# --- Abstractive: seq2seq model ---
from transformers import pipeline
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
def abstractive_summary(text: str, max_length: int = 130) -> str:
# BART has a 1024-token context window - chunk long documents first
result = summarizer(text, max_length=max_length, min_length=30, do_sample=False)
return result[0]["summary_text"]
Chunking strategies for long documents
Chunking is critical for RAG quality. Poor chunking is the single most common
cause of poor retrieval recall.
python
from langchain.text_splitter import RecursiveCharacterTextSplitter
def chunk_document(
text: str,
chunk_size: int = 512,
chunk_overlap: int = 64,
) -> list[dict]:
"""
Recursive splitter tries paragraph -> sentence -> word boundaries in order.
chunk_overlap ensures context continuity across chunk boundaries.
"""
splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size,
chunk_overlap=chunk_overlap,
separators=["\n\n", "\n", ". ", " ", ""],
)
chunks = splitter.split_text(text)
return [{"text": chunk, "chunk_index": i, "total_chunks": len(chunks)} for i, chunk in enumerate(chunks)]
# Semantic chunking (group sentences by embedding similarity instead of length)
from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai.embeddings import OpenAIEmbeddings
semantic_splitter = SemanticChunker(
OpenAIEmbeddings(),
breakpoint_threshold_type="percentile", # split where similarity drops sharply
breakpoint_threshold_amount=95,
)
semantic_chunks = semantic_splitter.create_documents([text])
Rule of thumb: chunk_size 256-512 tokens for precise retrieval, 512-1024 for
richer context. Always store chunk metadata (source doc ID, page, position)
alongside the embedding.
Show full SKILL.md (543 more words)Show less
Anti-patterns / common mistakes
Mistake
Why it's wrong
What to do instead
Embedding raw HTML or markdown
Markup tokens poison the semantic space
Strip all markup in preprocessing before embedding
Fixed-size chunks with no overlap
Splits sentences at boundaries, breaks coherence
Use recursive splitter with 10-20% overlap
Re-embedding at query time if corpus is static
Unnecessary latency on every request
Pre-compute all embeddings offline; embed only on writes
Using Euclidean distance for text similarity
Less meaningful than cosine for high-dimensional sparse-ish vectors
Normalize embeddings and use cosine/dot product
Fine-tuning a large model before trying a small pretrained one
Expensive, slow, often unnecessary
Benchmark a frozen small model first; fine-tune only if quality gap exists
Ignoring tokenizer mismatch between training and inference
Token boundaries differ, degrading model accuracy
Use the same tokenizer class and vocab for train and serve
Gotchas
Embedding model upgrades invalidate the entire index - Switching from BAAI/bge-small-en-v1.5 to text-embedding-3-small (or any other model) produces vectors in a different semantic space. Mixing embeddings from two different models in the same index causes meaningless similarity scores. When upgrading embedding models, you must re-embed and re-index every document in the corpus before the new model can be used in production.
Preprocessing applied at index time must be applied identically at query time - If you lowercase and strip HTML when building the index but forget to apply the same preprocessing to the query string, queries produce poor recall because the normalized index vectors don't match un-normalized query vectors. Encapsulate preprocessing in a shared function called by both the indexing pipeline and the query path.
FAISS IndexFlatIP does exact search but does not scale past ~1M vectors - For production corpora above ~500K documents, exact search latency becomes unacceptable. Use IndexIVFFlat (inverted file index, approximate) or a managed vector store. The tradeoff is recall (95-99% instead of 100%) for 10-100x faster search. Benchmark recall vs. latency before committing to an ANN approach.
HuggingFace pipeline() loads the full model on every call in scripts - Calling pipeline("token-classification", model="...") inside a request handler or loop re-loads the model weights from disk on every invocation, causing massive latency. Instantiate the pipeline once at module load time (or application startup) and reuse the same instance across all requests.
Sentence boundary detection matters more than chunk size for retrieval quality - Splitting text every N characters without checking for sentence boundaries creates chunks that start or end mid-sentence. These partial-sentence chunks retrieve poorly because their vectors average semantically incomplete text. Use RecursiveCharacterTextSplitter with sentence-aware separators (["\n\n", "\n", ". ", " "]) rather than a character-count-only splitter.
References
For detailed comparison tables and implementation guidance on specific topics,
read the relevant file from the references/ folder:
references/embedding-models.md - comparison of OpenAI, Cohere, sentence-transformers, E5, BGE with dimensions, benchmarks, and cost
Only load a references file if the current task requires it - they are long and
will consume context.
Companion check
On first activation of this skill in a conversation: check which companion skills are installed by running ls ~/.claude/skills/ ~/.agent/skills/ ~/.agents/skills/ .claude/skills/ .agent/skills/ .agents/skills/ 2>/dev/null. Compare the results against the recommended_skills field in this file's frontmatter. For any that are missing, mention them once and offer to install:
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.
NLP Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
NLP Engineering compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
NLP Engineering this skillmajiayu000/claude-skill-registry
Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
A skill your agent uses when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization. NLP Engineering is an agent skill from majiayu000/claude-skill-registry. Use this skill when building NLP pipelines, implementing text classification, semantic search, embeddings, or summarization.
When should I use NLP Engineering?
NLP Engineering fits situations like: building NLP pipelines; implementing text classification; semantic search; text preprocessing.
How do I install NLP Engineering in Claude Code?
Run `npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a claude-code`. Or copy the skill folder (skills/ai-ml/nlp-engineering in majiayu000/claude-skill-registry) into .claude/skills/nlp-engineering in your project. Claude Code loads it when a task matches its description.
How do I install NLP Engineering in Codex?
Run `npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a codex`. Or copy the skill folder (skills/ai-ml/nlp-engineering in majiayu000/claude-skill-registry) into .agents/skills/nlp-engineering in your project. Codex loads it when a task matches its description.
Can I use NLP Engineering in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill nlp-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nlp-engineering, .gemini/skills/nlp-engineering, .github/skills/nlp-engineering and .opencode/skills/nlp-engineering in your project.
What does NLP Engineering need to run?
SKILL.md names no scripts, command-line tools or credentials: NLP Engineering is instructions for the agent only. Our summary lists: Python 3.
Does NLP Engineering access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is NLP Engineering safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does NLP Engineering use?
NLP Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does NLP Engineering use?
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to NLP Engineering?
Skills that share tags, products or a category with NLP Engineering: Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars), Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Research (taishi-i/awesome-japanese-nlp-resources, 1k stars) and Search (taishi-i/awesome-japanese-nlp-resources, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains NLP Engineering?
majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.