Agent skill

Legal NLP Guide

by wentorai in wentorai/research-plugins

NLP techniques for legal text analysis, case law mining, and contracts

MITAuto-check passedLegal & Compliance

Install Legal NLP Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill legal-nlp-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins legal-nlp-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/law/legal-nlp-guide .claude/skills/legal-nlp-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
legal-nlp-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
283 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

NLP techniques for legal text analysis, case law mining, and contracts

  • Tasks that involve Natural language processing
  • SKILL.md covers Legal Text Characteristics, Legal Text Classification, Named Entity Recognition and Contract Analysis, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Legal research

What it does

Legal NLP Guide is an agent skill from wentorai/research-plugins. NLP techniques for legal text analysis, case law mining, and contracts

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Legal & Compliance, covering Natural language processing and Legal research. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Natural language processing
  • Tasks that involve Legal research

Example prompts

  • “/legal-nlp-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Legal NLP Guide loads about 2.1k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 283 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 283 words, ~2,124 tokens.

Download SKILL.mdSave it as .claude/skills/legal-nlp-guide/SKILL.md (or your agent's skills folder).
name
legal-nlp-guide
description
NLP techniques for legal text analysis, case law mining, and contracts

A skill for applying natural language processing techniques to legal texts. Covers legal document classification, named entity recognition for legal entities, contract clause extraction, case law similarity search, and court opinion summarization using modern NLP tools.

Legal language presents unique NLP challenges:

  • Long documents: Court opinions average 5,000-20,000 tokens; contracts can exceed 50,000
  • Domain-specific vocabulary: Terms of art with precise legal meanings (e.g., "consideration", "estoppel")
  • Complex syntax: Multi-clause sentences with nested qualifications and cross-references
  • Citation networks: Dense cross-referencing between cases, statutes, and regulations
  • Temporal reasoning: Effective dates, amendments, and retroactivity
Document Type Classification
python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

# Legal-BERT: domain-adapted BERT for legal text
model_name = "nlpaueb/legal-bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
    model_name, num_labels=5
)

# Legal document categories
labels = ["contract", "court_opinion", "statute", "regulation", "brief"]

def classify_legal_document(text: str, max_length: int = 512) -> dict:
    """
    Classify a legal document into predefined categories.
    For long documents, use the first 512 tokens (typically the
    preamble/introduction which contains strong classification signals).
    """
    inputs = tokenizer(
        text, return_tensors="pt",
        max_length=max_length, truncation=True, padding=True
    )
    with torch.no_grad():
        logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1).squeeze()
    predicted = labels[probs.argmax().item()]
    return {
        "predicted_class": predicted,
        "confidence": probs.max().item(),
        "all_scores": {l: p.item() for l, p in zip(labels, probs)},
    }
Topic Classification for Case Law

Common topic taxonomies for legal research:

CategoryExamples
Constitutional LawDue process, equal protection, First Amendment
Criminal LawSentencing, evidence, plea bargaining
Contract LawBreach, formation, damages
Tort LawNegligence, product liability, defamation
Property LawReal property, intellectual property, zoning
Administrative LawAgency rulemaking, judicial review

Named Entity Recognition

Legal NER extends standard NER with domain-specific entity types:

python
import spacy

# Load a legal NER model (e.g., trained on the LegalNERo dataset)
# or fine-tune spaCy on legal annotations
nlp = spacy.load("en_legal_ner")

legal_entity_types = {
    "COURT": "Court or tribunal name",
    "JUDGE": "Judge or justice name",
    "PARTY": "Plaintiff, defendant, petitioner, respondent",
    "STATUTE": "Statute or regulation citation",
    "CASE_CITATION": "Case name and reporter citation",
    "DATE": "Dates of decisions, filings, events",
    "JURISDICTION": "Geographic or subject matter jurisdiction",
    "PROVISION": "Specific section or clause reference",
}

def extract_legal_entities(text: str) -> list[dict]:
    """Extract legal named entities from text."""
    doc = nlp(text)
    entities = []
    for ent in doc.ents:
        entities.append({
            "text": ent.text,
            "label": ent.label_,
            "start": ent.start_char,
            "end": ent.end_char,
            "description": legal_entity_types.get(ent.label_, ""),
        })
    return entities
Citation Extraction and Parsing
python
import re

# US case citation patterns (simplified)
CASE_CITE_PATTERN = re.compile(
    r"(?P<volume>\d+)\s+"
    r"(?P<reporter>U\.S\.|S\.\s?Ct\.|F\.\s?\d[dthsr]+|"
    r"F\.\s?Supp\.\s?\d*[dthsr]*)\s+"
    r"(?P<page>\d+)"
    r"(?:\s*,\s*(?P<pinpoint>\d+))?"
    r"(?:\s*\((?P<year>\d{4})\))?"
)

def parse_citations(text: str) -> list[dict]:
    """Extract and parse legal citations from text."""
    citations = []
    for match in CASE_CITE_PATTERN.finditer(text):
        citations.append({
            "full_match": match.group(),
            "volume": match.group("volume"),
            "reporter": match.group("reporter"),
            "page": match.group("page"),
            "pinpoint": match.group("pinpoint"),
            "year": match.group("year"),
        })
    return citations

Contract Analysis

Clause Extraction and Classification
python
def segment_contract_clauses(text: str) -> list[dict]:
    """
    Segment a contract into numbered clauses and classify them.
    Uses section numbering patterns as primary segmentation cues.
    """
    # Split on section/article numbering patterns
    section_pattern = re.compile(
        r"\n\s*(?:Section|Article|Clause|\d+\.)\s+\d+[\.\d]*\s*[:\.\-]?\s*",
        re.IGNORECASE,
    )
    sections = section_pattern.split(text)
    headers = section_pattern.findall(text)

    clause_types = {
        "indemnification": ["indemnif", "hold harmless", "defend and indemnify"],
        "termination": ["terminat", "cancel", "expir"],
        "confidentiality": ["confidential", "non-disclosure", "proprietary"],
        "limitation_of_liability": ["limit of liabilit", "limitation of liabilit",
                                     "aggregate liability", "consequential damages"],
        "governing_law": ["governing law", "governed by", "jurisdiction"],
        "force_majeure": ["force majeure", "act of god", "beyond reasonable control"],
        "assignment": ["assign", "transfer", "delegate"],
    }

    clauses = []
    for i, section in enumerate(sections[1:], 1):
        detected_type = "general"
        section_lower = section.lower()
        for ctype, keywords in clause_types.items():
            if any(kw in section_lower for kw in keywords):
                detected_type = ctype
                break
        clauses.append({
            "index": i,
            "header": headers[i - 1].strip() if i <= len(headers) else "",
            "type": detected_type,
            "text": section.strip()[:500],
        })
    return clauses
Embedding-Based Case Retrieval
python
from sentence_transformers import SentenceTransformer
import numpy as np

# Legal domain sentence embeddings
encoder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")

def build_case_index(case_summaries: list[str]) -> np.ndarray:
    """Encode case summaries into dense vector representations."""
    embeddings = encoder.encode(case_summaries, show_progress_bar=True)
    # L2 normalize for cosine similarity via dot product
    norms = np.linalg.norm(embeddings, axis=1, keepdims=True)
    return embeddings / norms

def search_similar_cases(query: str, index: np.ndarray,
                         case_ids: list[str], top_k: int = 10) -> list:
    """Find the most similar cases to a query."""
    query_vec = encoder.encode([query])
    query_vec = query_vec / np.linalg.norm(query_vec)
    scores = (index @ query_vec.T).squeeze()
    top_indices = np.argsort(scores)[::-1][:top_k]
    return [(case_ids[i], scores[i]) for i in top_indices]
  • CaseHOLD: Multiple-choice QA from case law holdings (Harvard)
  • LEDGAR: 100,000 contract provisions labeled with 12 clause types
  • ECtHR dataset: European Court of Human Rights case texts with violation labels
  • LegalBench: Multi-task benchmark for legal reasoning (Stanford)
  • CUAD (Contract Understanding Atticus Dataset): 510 contracts with 41 clause type annotations

Tools and Resources

  • Legal-BERT / CaseLaw-BERT: Domain-adapted transformer models
  • spaCy + Blackstone: Legal NER and text processing for UK law
  • Haystack: Open-source framework for legal document search and QA
  • CourtListener / RECAP: Free US case law and PACER document archive
  • LexNLP (ContraxSuite): Python library for legal text extraction

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/law/legal-nlp-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Legal NLP Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Legal NLP Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Legal NLP Guide this skillwentorai/research-plugins2981 repos~2.1kAutomated safety check: PassMIT
Tw Legal RAGaa0101181514/tw-legal-rag328—~580Automated safety check: PassCustom licence
Recallentireio/skills223—~1.6kAutomated safety check: PassMIT
Pseudonymization Riskmukul975/Privacy-Data-Protection-Skills301—~3.2kAutomated safety check: PassApache-2.0
Legal AI Model Router Stephane Boghossianlawve-ai/awesome-legal-skills847—~1.5kAutomated safety check: PassAGPL-3.0-or-later
Design Award SearchSeanJ1ang/design-judge-skills712—~3kAutomated safety check: PassApache-2.0

Similar skills

  • Tw Legal RAG

    aa0101181514/tw-legal-rag

    Retrieve real Taiwan court judgments with verifiable citations before answering any question about Taiwan law or case law.

    328 GitHub stars~580 tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Recall

    entireio/skills

    A skill your agent uses when the user describes a task and wants to know whether something similar has been done before, then turn the closest prior session into a task playbook.

    223 GitHub stars~1.6k tokensUpdated 12 days ago
    Legal & ComplianceAuto-check passed
  • Pseudonymization Risk

    mukul975/Privacy-Data-Protection-Skills

    Assessment of pseudonymization techniques and re-identification risk.

    301 GitHub stars~3.2k tokensUpdated 6 mo ago
    Legal & ComplianceAuto-check passed
  • Legal AI Model Router Stephane Boghossian

    lawve-ai/awesome-legal-skills

    Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty.

    847 GitHub stars~1.5k tokensUpdated 8 days ago
    Legal & ComplianceAuto-check passed
  • Design Award Search

    SeanJ1ang/design-judge-skills

    Find and verify award-winning designs in the same or adjacent functional category through eight explicit relevance dimensions: problem and user, core function, sensing technology, intervention…

    712 GitHub stars~3k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • China Lawyer Analyst

    CSlawyer1985/china-lawyer-analyst

    通过中国法律视角分析事件,运用成文法解释、指导案例参照、请求权基础分析等方法, 理解权利义务、评估责任风险、识别法律依据并推荐合规策略。

    196 GitHub stars~3.3k tokensUpdated 8 mo ago
    Legal & ComplianceAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Legal NLP Guide

What does Legal NLP Guide do?

NLP techniques for legal text analysis, case law mining, and contracts. Legal NLP Guide is an agent skill from wentorai/research-plugins.

When should I use Legal NLP Guide?

Legal NLP Guide fits situations like: tasks that involve Natural language processing; tasks that involve Legal research.

How do I install Legal NLP Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill legal-nlp-guide -a claude-code`. Or copy the skill folder (skills/domains/law/legal-nlp-guide in wentorai/research-plugins) into .claude/skills/legal-nlp-guide in your project. Claude Code loads it when a task matches its description.

How do I install Legal NLP Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill legal-nlp-guide -a codex`. Or copy the skill folder (skills/domains/law/legal-nlp-guide in wentorai/research-plugins) into .agents/skills/legal-nlp-guide in your project. Codex loads it when a task matches its description.

Can I use Legal NLP Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill legal-nlp-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/legal-nlp-guide, .gemini/skills/legal-nlp-guide, .github/skills/legal-nlp-guide and .opencode/skills/legal-nlp-guide in your project.

What does Legal NLP Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Legal NLP Guide is instructions for the agent only. Our summary lists: Python 3.

Does Legal NLP Guide access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Legal NLP Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Legal NLP Guide use?

Legal NLP Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Legal NLP Guide use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Legal NLP Guide?

Skills that share tags, products or a category with Legal NLP Guide: Tw Legal RAG (aa0101181514/tw-legal-rag, 328 stars), Recall (entireio/skills, 223 stars), Pseudonymization Risk (mukul975/Privacy-Data-Protection-Skills, 301 stars) and Legal AI Model Router Stephane Boghossian (lawve-ai/awesome-legal-skills, 847 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Legal NLP Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.