Topic · AI & LLM Engineering
Best natural language processing skills for Claude Code, Codex and other agents.
- skills
- 141
- official
- 3
Natural language processing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Helps write and debug JavaScript or TypeScript that uses the compromise English NLP library for matching, entity extraction, tagging and sentence transforms. | spencermountain/ | 12k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 2 | Google API integration for blog performance: PageSpeed Insights, CrUX Core Web Vitals with 25-week history, Search Console performance, URL Inspection, Indexing API, GA4 organic traffic, NLP entity… | AgriciDaniel/ | 2.3k | 1 repo | ~3.3k | Automated safety check: Notes | MIT | 5 days ago |
| 3 | Guide for writing correct code with the compromise rule-based NLP library: tagging, match syntax, in-place transforms and common tasks like tense changes and redaction. | spencermountain/ | 12k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 4 | Three modes for CS-conference papers (CVPR/ICCV/ECCV vision, ACL/EMNLP/NAACL NLP, ICLR/NeurIPS/ICML/AAAI ML). | Spark-To-Paper-Skills/ | 1.2k | — | ~5.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 5 | Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking. | Orchestra-Research/ | 13k | 7 repos | ~3.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations. | maziyarpanahi/ | 5.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 7 | Tokenize, tag, and analyze natural language text using Apple's NaturalLanguage framework and translate between languages with the Translation framework. | dpearson2699/ | 1.2k | 1 repo | ~3.5k | Automated safety check: Pass | Unknown | 2 mo ago |
| 8 | Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems. | ModelCloud/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 9 | Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | Global-installable, project-level academic rebuttal strategy skill for AI/ML/CV/NLP/Robotics papers. | xiongqi123123/ | 305 | — | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 11 | Drive native macOS apps via interceptor macos : AX trees, background click/type/keys/drag/scroll, occluded or minimized window capture, browser chrome, URL bars, OS dialogs, Apple Events, trusted OS… | Hacker-Valley-Media/ | 514 | — | ~2k | Automated safety check: Pass | Unknown | 5 days ago |
| 12 | Search 2500+ curated ChatGPT and LLM open-source repositories. | taishi-i/ | 3.3k | — | ~3.8k | Automated safety check: Pass | CC0-1.0 | 3 days ago |
| 13 | Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics. | maziyarpanahi/ | 5.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 14 | Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols. | curiositech/ | 243 | 3 repos | ~3.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 15 | 15.Stellar Dev End-to-end Stellar development playbook. An agent skill from VelaPayments/vela-payments. | VelaPayments/ | 131 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 16 | 16.SEO Google Direct access to Google's own SEO data via Search Console (Search Analytics, URL Inspection, Sitemaps), PageSpeed Insights v5, CrUX field data with 25-week history, Indexing API v3, GA4 organic… | seranking/ | 160 | — | ~4.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm. | maziyarpanahi/ | 5.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs). | K-Dense-AI/ | 282 | — | ~1.9k | Automated safety check: Pass | MIT | 1 mo ago |
| 19 | Analyze text content using both traditional NLP and LLM-enhanced methods. | liangdabiao/ | 290 | 1 repo | ~1.7k | Automated safety check: Notes | No licence | 5 mo ago |
| 20 | Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 21 | This skill should be used when the user asks to "learn from Kaggle", "study Kaggle solutions", "analyze Kaggle competitions", or mentions Kaggle competition URLs. | Galaxy-Dawn/ | 5.7k | 2 repos | ~940 | Automated safety check: Pass | MIT | 14 days ago |
| 22 | 22.Compare Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as… | taishi-i/ | 1k | — | ~4.1k | Automated safety check: Notes | CC0-1.0 | yesterday |
| 23 | Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text. | Orchestra-Research/ | 13k | 3 repos | ~1.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 24 | Trains and uses SentencePiece tokenizers on raw text, with BPE or Unigram models, for multilingual and CJK projects that need a reproducible vocabulary. | Orchestra-Research/ | 13k | 3 repos | ~1.4k | Automated safety check: Notes | MIT | 3 mo ago |
| 25 | 25.Research Analyze current trends and challenges in Japanese NLP for a topic. | taishi-i/ | 1k | — | ~3.5k | Automated safety check: Notes | CC0-1.0 | yesterday |
| 26 | 學術研究實驗設計技能——從研究假設到可重現實驗計畫的完整流程。當使用者需要規劃實驗、設計 ablation study、選擇 baseline、確定評估指標,或問「我應該跑哪些實驗」時,一定要使用此技能。觸發詞包括:實驗設計、experiment design、ablation、baseline、跑什麼實驗、evaluation metric、如何驗證方法。適用於機器學習、NLP、CV… | voidful/ | 132 | — | ~1.2k | Automated safety check: Pass | MIT | 6 mo ago |
| 27 | Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets. | davila7/ | 32k | 12 repos | ~1.2k | Automated safety check: Pass | MIT | today |
| 28 | Uses the Deepgram Python SDK's Read API to analyze text for sentiment, summaries, topics and intents with client.read.v1.text.analyze, from raw text or a hosted URL. | deepgram/ | 469 | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 29 | Consolidates BERTopic, LDA or NMF topic output into a theory-driven classification framework and writes the final labels back to an Excel file. | TyrealQ/ | 108 | — | ~1k | Automated safety check: Pass | MIT | 14 days ago |
| 30 | Applies the reasoning, architectural principles, and AI philosophy of Christopher Manning (natural language processing expert, Stanford University, director of Stanford AI Lab). | K-Dense-AI/ | 282 | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 31 | Expert developer for Calcpad.Highlighter - tokenization, linting, content resolution, and language tooling. | imartincei/ | 109 | — | ~1k | Automated safety check: Notes | MIT | yesterday |
| 32 | 32.Search Search all Japanese NLP resources (libraries, models, datasets, tutorials, dictionaries, Hugging Face). | taishi-i/ | 1k | — | ~4.3k | Automated safety check: Notes | CC0-1.0 | yesterday |
| 33 | Add or modify a Lizard language reader. An agent skill from terryyin/lizard. | terryyin/ | 2.5k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 34 | 34.Discover Given a Japanese NLP GitHub repo/model/dataset (URL / owner/repo / tool name) OR a topic, find what's already in awesome-japanese-nlp-resources and discover related resources NOT yet listed… | taishi-i/ | 1k | — | ~6.5k | Automated safety check: Notes | CC0-1.0 | yesterday |
| 35 | Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks. | huggingface/ | 11k | 1 repo | ~6.2k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 36 | 나노바나나 프롬프트의 정확성을 한글 번역본과 비교하여 검증하고, WebSearch로 팩트체크한 후 이슈별로 사용자 확인을 거쳐 수정합니다. | team-attention/ | 294 | — | ~2.2k | Automated safety check: Pass | No licence | 7 mo ago |
| 37 | Azure AI Text Analytics SDK for sentiment analysis, entity recognition, key phrases, language detection, PII, and healthcare NLP. | microsoft/ | 3.1k | 6 repos | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 38 | Scores news, announcements and macro events with the LLM, stores them in an event CSV and blends the decaying event signal with technical signals in signal_engine.py. | HKUDS/ | 35k | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 39 | Build, inspect, prepare, generate and export Overmind datasets in Data Workshop. | overmind-core/ | 544 | — | ~875 | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 40 | 40.Gtars Supports Gtars for local genomic interval models and set algebra, overlaps and counts, consensus and coverage, tokenization, fragment processing, and refget/BEDbase planning across Python, Rust, and… | K-Dense-AI/ | 48k | 1 repo | ~3.8k | Automated safety check: Notes | MIT | 2 days ago |
| 41 | 41.Phee Query Query the PHEE pharmacovigilance event extraction dataset. An agent skill from QSong-github/DrugClaw. | QSong-github/ | 116 | 1 repo | ~694 | Automated safety check: Pass | No licence | 1 mo ago |
| 42 | Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label. | maziyarpanahi/ | 5.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 43 | Combine OpenMed clinical NLP with Microsoft Presidio, spaCy, or LangChain through OpenMed's built-in interop adapter registry (openmed.interop). | maziyarpanahi/ | 5.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 44 | Orient and bootstrap any project that uses OpenMed, the on-device clinical and biomedical NLP library, for named-entity recognition, PHI de-identification, FHIR export, and evaluation. | maziyarpanahi/ | 5.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 45 | Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. | maziyarpanahi/ | 5.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | Run clinical and biomedical named-entity recognition on medical text with OpenMed's analyzetext. | maziyarpanahi/ | 5.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 47 | Replace detected PHI with realistic, type-matched fake values in OpenMed so clinical notes stay readable and parseable instead of full of [REDACTED] markers. | maziyarpanahi/ | 5.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 48 | Extract arbitrary, custom entity types from clinical or biomedical text with no fine-tuning using OpenMed's GLiNER / GLiNER2 zero-shot support. | maziyarpanahi/ | 5.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
Questions, answered from the data.
What is the best natural language processing skill?
Compromise NLP Library from spencermountain/compromise ranks first of the 141 natural language processing skills listed here, with the highest score: its repository has 12k GitHub stars, its SKILL.md loads about 2k tokens and it passes the automated safety check with no findings. Next come Blog Google and Compromise NLP for JavaScript.
Which natural language processing skills are official?
3 of the 141 natural language processing skills are official, published by the vendor's own GitHub organization: Transformers.js, Azure AI Textanalytics Py and Azure AI Language Conversations Py.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Reinforcement learning67
- AI interpretability23