Agent skill

Open Semantic Search Guide

by wentorai in wentorai/research-plugins

Self-hosted semantic search and text mining platform. An agent skill from wentorai/research-plugins.

MITAuto-check passedAI & LLM Engineering

Install Open Semantic Search Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill open-semantic-search-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins open-semantic-search-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/literature/search/open-semantic-search-guide .claude/skills/open-semantic-search-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
open-semantic-search-guide
GitHub stars
298
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
125 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Self-hosted semantic search and text mining platform. An agent skill from wentorai/research-plugins.

  • Works in 5 steps: Paper search: Full-text search over… → Literature mining: Extract entities and… → Institutional repository: Campus-wide… → …
  • Tasks that involve Natural language processing
  • SKILL.md covers Overview, Installation, Architecture and Indexing Documents, plus 6 more sections
  • Calls curl, git and docker-compose; reaches github.com

What it does

Open Semantic Search Guide is an agent skill from wentorai/research-plugins. Self-hosted semantic search and text mining platform

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Natural language processing

Example prompts

  • “/open-semantic-search-guide”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Paper search: Full-text search over local paper collections
  2. Literature mining: Extract entities and relationships from papers
  3. Institutional repository: Campus-wide document search
  4. Due diligence: Search across legal/business document archives
  5. Investigative research: Cross-reference entities across documents

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • git
    • docker-compose

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • solr.apache.org
    • tika.apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Open Semantic Search Guide loads about 1.3k tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 125 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 125 words, ~1,341 tokens.

Download SKILL.mdSave it as .claude/skills/open-semantic-search-guide/SKILL.md (or your agent's skills folder).
name
open-semantic-search-guide
description
Self-hosted semantic search and text mining platform

Open Semantic Search Guide

Overview

Open Semantic Search is a self-hosted search and text mining platform that combines full-text search (Apache Solr) with semantic analysis — entity extraction, named entity recognition, text classification, and knowledge graph building. Process and search across documents (PDF, DOCX, emails) with faceted navigation and visual analytics. Ideal for researchers needing private, on-premise document search over large paper collections.

Installation

bash
# Docker deployment (recommended)
git clone https://github.com/opensemanticsearch/open-semantic-search.git
cd open-semantic-search
docker-compose up -d

# Access web UI at http://localhost:8080
# Admin panel at http://localhost:8080/admin

Architecture

Documents (PDF, DOCX, HTML, email)
         ↓
   Connector/Crawler (file system, web, IMAP)
         ↓
   ETL Pipeline
   ├── Text extraction (Apache Tika)
   ├── OCR (Tesseract, for scanned docs)
   ├── NER (spaCy, Stanford NER)
   ├── Entity linking (knowledge base)
   └── Classification (custom models)
         ↓
   Apache Solr (full-text index + facets)
         ↓
   Web UI (search, browse, visualize)

Indexing Documents

bash
# Index a directory of papers
curl -X POST "http://localhost:8080/api/index" \
  -H "Content-Type: application/json" \
  -d '{"path": "/data/papers/", "recursive": true}'

# Index single file
curl -X POST "http://localhost:8080/api/index" \
  -H "Content-Type: application/json" \
  -d '{"path": "/data/papers/attention.pdf"}'

# Schedule recurring index
# Add to crontab or use built-in scheduler

Search Features

markdown
### Full-Text Search
- Boolean queries: "attention mechanism" AND transformer
- Phrase search: "self-attention"
- Wildcard: transform*
- Proximity: "attention transformer"~5 (within 5 words)
- Field-specific: title:"attention" author:"Vaswani"

### Faceted Navigation
- Filter by: author, date, organization, topic, language
- Nested facets for hierarchical browsing
- Date range slider
- Entity type filters (person, organization, location)

### Semantic Features
- Named entity highlighting in results
- Related entity suggestions
- Concept co-occurrence visualization
- Auto-generated tag clouds

Python Client

python
import requests

SEARCH_URL = "http://localhost:8080/api/search"

def search_papers(query, filters=None, max_results=20):
    """Search indexed documents."""
    params = {
        "q": query,
        "rows": max_results,
        "fl": "title,author,content_type,date,score",
        "hl": "true",        # Highlight matches
        "hl.fl": "content",  # Highlight in content field
        "facet": "true",
        "facet.field": ["author", "organization", "topic"],
    }
    if filters:
        params["fq"] = filters

    resp = requests.get(SEARCH_URL, params=params)
    data = resp.json()

    results = data["response"]["docs"]
    facets = data.get("facet_counts", {}).get("facet_fields", {})

    return results, facets

# Search
results, facets = search_papers(
    "attention mechanism transformer",
    filters='date:[2023-01-01T00:00:00Z TO *]',
)

for doc in results:
    print(f"[{doc.get('date', 'N/A')}] {doc.get('title', 'Untitled')}")
    print(f"  Score: {doc['score']:.2f}")

Entity Extraction Configuration

json
{
  "ner": {
    "engines": ["spacy", "stanford"],
    "models": {
      "spacy": "en_core_web_lg",
      "stanford": "english.all.3class.caseless"
    },
    "entity_types": [
      "PERSON", "ORG", "GPE", "DATE",
      "WORK_OF_ART", "EVENT"
    ],
    "custom_entities": {
      "METHODOLOGY": ["transformer", "CNN", "RNN", "GAN"],
      "DATASET": ["ImageNet", "CIFAR", "MNIST", "COCO"]
    }
  },
  "classification": {
    "enabled": true,
    "model": "custom_topic_classifier",
    "categories": ["NLP", "CV", "RL", "Theory"]
  }
}

Knowledge Graph

python
# Query the auto-built knowledge graph
def get_entity_network(entity, depth=2):
    """Get co-occurring entities for a given entity."""
    resp = requests.get(
        f"{SEARCH_URL}/graph",
        params={"entity": entity, "depth": depth},
    )
    graph = resp.json()

    for node in graph["nodes"]:
        print(f"Entity: {node['label']} ({node['type']})")
    for edge in graph["edges"]:
        print(f"  {edge['source']} ↔ {edge['target']} "
              f"(co-occur: {edge['weight']})")

get_entity_network("Transformer")

Use Cases

  1. Paper search: Full-text search over local paper collections
  2. Literature mining: Extract entities and relationships from papers
  3. Institutional repository: Campus-wide document search
  4. Due diligence: Search across legal/business document archives
  5. Investigative research: Cross-reference entities across documents

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/literature/search/open-semantic-search-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Open Semantic Search Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Open Semantic Search Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Open Semantic Search Guide this skillwentorai/research-plugins2981 repos~1.3kAutomated safety check: PassMIT
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
OpenMed Model Card Writermaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Open Semantic Search Guide

What does Open Semantic Search Guide do?

Self-hosted semantic search and text mining platform. An agent skill from wentorai/research-plugins. Open Semantic Search Guide is an agent skill from wentorai/research-plugins.

When should I use Open Semantic Search Guide?

Open Semantic Search Guide fits situations like: tasks that involve Natural language processing.

How do I install Open Semantic Search Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill open-semantic-search-guide -a claude-code`. Or copy the skill folder (skills/literature/search/open-semantic-search-guide in wentorai/research-plugins) into .claude/skills/open-semantic-search-guide in your project. Claude Code loads it when a task matches its description.

How do I install Open Semantic Search Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill open-semantic-search-guide -a codex`. Or copy the skill folder (skills/literature/search/open-semantic-search-guide in wentorai/research-plugins) into .agents/skills/open-semantic-search-guide in your project. Codex loads it when a task matches its description.

Can I use Open Semantic Search Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill open-semantic-search-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/open-semantic-search-guide, .gemini/skills/open-semantic-search-guide, .github/skills/open-semantic-search-guide and .opencode/skills/open-semantic-search-guide in your project.

What does Open Semantic Search Guide need to run?

Going by SKILL.md and its folder, Open Semantic Search Guide needs the command-line tools its instructions call (curl, git and docker-compose). Our summary lists: Python 3; Docker.

Does Open Semantic Search Guide access the network?

SKILL.md names 3 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: solr.apache.org and tika.apache.org. This is read from the text; nothing was executed.

Is Open Semantic Search Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Open Semantic Search Guide use?

Open Semantic Search Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Open Semantic Search Guide use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Open Semantic Search Guide?

Skills that share tags, products or a category with Open Semantic Search Guide: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenMed Model Card Writer (maziyarpanahi/openmed, 5.5k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and Andrej Karpathy (K-Dense-AI/mimeo, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Open Semantic Search Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.