Agent skill

Solr Semantic Search

by griddynamics in griddynamics/rosetta

To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Solr Semantic Search

skills CLI
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install griddynamics/rosetta solr-semantic-search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .claude/skills/solr-semantic-search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
solr-semantic-search
GitHub stars
354
Token cost
~1.9k tokens
SKILL.md length
750 words
Files
10 (incl. references)
Skills in repo
12
Repo updated
First seen
Licence
Apache-2.0

At a glance

To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.

  • Works in 3 steps: Tagging — phrase → analyzed tokens →… → Graph — tags become edges, positions… → Query building — for each viable path,…
  • Tasks that involve Embeddings
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Solr Semantic Search is an agent skill from griddynamics/rosetta. To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `README.md`, `references/01-architecture.md` and `references/02-concept-indexing.md`).

It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: An instruction layer for AI coding agent. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Embeddings

Example prompts

  • “/solr-semantic-search”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Tagging — phrase → analyzed tokens → shingles (1..N) → lookup in the concept index → ProducedTag list (token, position, type, matched…
  2. Graph — tags become edges, positions become vertices; find K-shortest paths (= valid phrase interpretations) and resolve ambiguity by…
  3. Query building — for each viable path, build an abstract Sm query, apply dependency groups and min-should-match, then translate to a Solr…

What it can do on your machine

Read from SKILL.md and the folder at commit 5441232. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Solr Semantic Search loads about 1.9k tokens when it runs, and up to ~45k if it reads all its reference files. Until then it costs about 27 tokens; SKILL.md has 750 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~45k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from griddynamics/rosetta at commit 5441232, republished under its Apache-2.0 licence (© griddynamics). 750 words, ~1,867 tokens.

Download SKILL.mdSave it as .claude/skills/solr-semantic-search/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
solr-semantic-search
description
To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.
license
Apache-2.0
tags
solr, semantic-search, tagging, query-understanding, taxonomy
baseSchema
docs/schemas/skill.md
<solr-semantic-search>
<role>

You are a senior Apache Solr engineer who designs, builds, debugs, and extends phrase-tagging semantic search on Solr 9.x: decomposing natural-language queries into structured concepts via dictionary lookup, resolving path ambiguity in a tag graph, and assembling precise multi-field Solr queries. This is lexical, not vector/embedding, semantic search.

</role>

<when_to_use_skill>

Concept tagging, query understanding, taxonomy-driven search, structured Brand/Line/Model recognition, shingle-based matching, multi-word synonyms, path resolution, fuzzy-phrase-to-structured-query extraction. Traditional Solr query work and vector/kNN semantic search → solr-query skill. Custom plugins this architecture relies on → solr-extending skill.

</when_to_use_skill>

<core_concepts>

Three independently testable layers, separated by stable interfaces (ProducedTag, StagedTag, SmQuery):

  1. Tagging — phrase → analyzed tokens → shingles (1..N) → lookup in the concept index → ProducedTag list (token, position, type, matched fields+weights).
  2. Graph — tags become edges, positions become vertices; find K-shortest paths (= valid phrase interpretations) and resolve ambiguity by dropping weak alternatives.
  3. Query building — for each viable path, build an abstract Sm query, apply dependency groups and min-should-match, then translate to a Solr query against the catalog.

This SKILL.md is a router. For any non-trivial question, read the relevant references/ file before answering — references hold the examples, schemas, code, and decision tables and are not duplicated here.

</core_concepts>

<references>
When the user asks about…Read
Architecture overview, the three layers, data flowREAD SKILL FILE references/01-architecture.md
Concept collection schema, building it from source data, indexing handlerREAD SKILL FILE references/02-concept-indexing.md
Phrase tagging mechanics: shingles, lookup, scoring, multi-language, fuzzy/word-break/prefixREAD SKILL FILE references/03-tagging.md
Graph construction (JGraphT), vertices/edges, paths, quasi-positions for multi-word synsREAD SKILL FILE references/04-graph-paths.md
Ambiguity resolution between competing interpretations (Path vs Shingle resolvers)READ SKILL FILE references/05-ambiguity-resolution.md
Building the final Solr query from tagged paths, Sm query model, dependency groupsREAD SKILL FILE references/06-query-building.md
Adapting this to a new domain: schema design, concept sources, stages configREAD SKILL FILE references/07-applying-to-domain.md
Sm* query model implementation — full code for SmQuery/SmBoolean/SmTerm and the Solr translator fabricREAD SKILL FILE references/08-query-model-implementation.md
</references>

<when_to_choose>

This is a heavyweight architecture. It is the right tool when the domain has well-defined concepts (products, models, attributes) with known synonyms, queries must be understood structurally ("what is the Brand? Line? attribute?"), vector search yields too many false positives for the required precision, and authoritative taxonomies exist to extract concepts from.

It is the wrong tool when the domain is open-ended natural language (use embeddings), there are no curated concept dictionaries, or only fuzzy retrieval is needed without structural understanding.

</when_to_choose>

<mental_model>

USER PHRASE: "sony wh-1000xm5 ear pads"
   ──► LAYER 1 TAGGING: tokens → shingles → concept-index lookup → ProducedTag list
   ──► LAYER 2 GRAPH: tags→edges, positions→vertices; K-shortest paths; resolve ambiguity
   ──► LAYER 3 QUERY BUILDING: per path build Sm query, dependency groups, min-should-match → Solr query
   ──► SOLR SEARCH against the catalog ──► RESULTS

Why it beats naive eDisMax, three problems:

  • Ambiguous tokens — "air" may be a Model (MacBook Air, weight 100) or description text (weight 1). The tagger emits both tags; the path resolver picks the higher-weight interpretation instead of letting scores compete across qf.
  • Multi-word concepts — "ear pads" is two tokens but one category. As a multi-word synonym it produces a single MULTI_SYN tag spanning both positions, preserving the structure eDisMax pf loses.
  • Domain rules — "sony wh-1000xm5" must validate that Sony's WH line includes the 1000XM5 model. A BLM post-processor (e.g. BrandLineModelProcessor) checks recognized Brand/Line/Model tags against a canonical CatalogProvider, drops invalid combos, and turns valid ones into structured filters (brand_id_s:SONY AND line_id_s:WH AND model_id_s:WH-1000XM5).
Show full SKILL.md (235 more words)Show less

</mental_model>

<key_data_types>

Token        — analyzed phrase token (term + position + lang)
Shingle      — N consecutive tokens treated as a unit
ProducedTag  — recognized concept: token, start/end position, relation type, matched fields (with weights)
StagedTag    — ProducedTag enriched with staging info (fields, boosts, dependencies) for a search stage
SmQuery      — abstract semantic query (SmBoolean/SmTerm/SmBoost/…) translated to a Lucene/Solr Query
TagType      — CONCEPT | SYN | MULTI_SYN | SPELL | PREFIX | RECOGNIZED_PRODUCT (validated Brand/Line/Model)
StageConfig  — per-stage config (fields, min-should-match, min-pattern-score, ambiguity resolver, …)

The tagger is a Solr request handler at /semanticTagGraph (params: q, lang, source, fuzzy, wordBreak, prefix, maxShingleLength, debug, dot). It returns tokens, tags (each with token, start/end, relation, entryFields weights), unrecognized, and a graphviz tagsDot. Downstream runs ambiguity resolution → path finding → query building, then hits the catalog collection.

</key_data_types>

<anti_patterns>

  • Indexing arbitrary text as concepts — the concept collection holds curated terms (catalog identifiers, taxonomy names, validated synonyms), not free text, or everything matches everything.
  • Skipping ambiguity resolution — without it the path resolver returns dozens of paths and the query builder produces a massive boolean OR; latency explodes.
  • Hardcoding BLM-like logic in the tagger — domain validation belongs in a post-processor, not the generic tagger.
  • Wide maxShingleLength — shingles 1..10 over a 10-token phrase is O(N²); cap at 4–5.
  • Per-request synonym loading — load SynonymsStorage once at startup.
  • Ignoring path coverage — a path that does not span the full phrase is incomplete; reject in the query builder unless the stage allows partial matches.
  • Using the catalog core for concept lookup — concepts live in their own small, fast collection; mixing them with the catalog wrecks both.

</anti_patterns>

<solr_10_deltas>

The architecture is Solr 9.x-tested. On Solr 10: BlockJoinParentQParser API stable; JGraphT is an external dep — pin to your build; custom RequestHandler/SearchComponent base classes unchanged; concept indexing via TermsComponent works the same, with minor changes to the /admin/luke response shape. On Solr 9.x these differences will not bite.

</solr_10_deltas>

</solr-semantic-search>

© griddynamics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in plugins/core-claude/skills/solr-semantic-search of griddynamics/rosetta.

  • SKILL.md
  • README.md
  • references/01-architecture.md
  • references/02-concept-indexing.md
  • references/03-tagging.md
  • references/04-graph-paths.md
  • references/05-ambiguity-resolution.md
  • references/06-query-building.md
  • references/07-applying-to-domain.md
  • references/08-query-model-implementation.md

Open the folder on GitHubat commit 5441232

Compare with similar skills

Solr Semantic Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Solr Semantic Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Solr Semantic Search this skillgriddynamics/rosetta354—~1.9kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Mashup Mods

    rehan-remade/universal-modder

    Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.

    5.3k GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from griddynamics/rosetta

All 12 skills in this repo
  • Collect GitHub Stats

    griddynamics/rosetta

    Collect GitHub repo health/usage stats into merged JSON. An agent skill from griddynamics/rosetta.

    354 GitHub stars~415 tokensUpdated 9 days ago
    Auto-check passed
  • External Lib Flow

    griddynamics/rosetta

    Workflow for onboarding an external private library so AI can use it without source access.

    354 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Specflow Use

    griddynamics/rosetta

    To connect Rosetta with Grid Dynamics SpecFlow MCP; only when SpecFlow is mentioned and the MCP is installed.

    354 GitHub stars~780 tokensUpdated 9 days ago
    Auto-check passed
  • Harness

    griddynamics/rosetta

    To build an AI harness: run, observe, validate, automate repeated work faster — CLI/MCP actions, devcontainers, skills, subagents, hooks, pipelines, automations.

    354 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Coding Agents Farm

    griddynamics/rosetta

    To orchestrate parallel coding-agent farms (Claude, Codex, Copilot, Gemini, etc.) on isolated git worktrees.

    354 GitHub stars~2.4k tokensUpdated 9 days ago
    Auto-check passed
  • Coding Agents Hooks Authoring

    griddynamics/rosetta

    To author, register, and test Rosetta hooks, add a SemanticKind, or debug a hook that won't fire.

    354 GitHub stars~1.4k tokensUpdated 9 days ago
    Auto-check passed

Questions about Solr Semantic Search

What does Solr Semantic Search do?

To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths. Solr Semantic Search is an agent skill from griddynamics/rosetta. To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.

When should I use Solr Semantic Search?

Solr Semantic Search fits situations like: tasks that involve Embeddings.

How do I install Solr Semantic Search in Claude Code?

Run `npx skills add griddynamics/rosetta --skill solr-semantic-search -a claude-code`. Or copy the skill folder (plugins/core-claude/skills/solr-semantic-search in griddynamics/rosetta) into .claude/skills/solr-semantic-search in your project. Claude Code loads it when a task matches its description.

How do I install Solr Semantic Search in Codex?

Run `npx skills add griddynamics/rosetta --skill solr-semantic-search -a codex`. Or copy the skill folder (plugins/core-claude/skills/solr-semantic-search in griddynamics/rosetta) into .agents/skills/solr-semantic-search in your project. Codex loads it when a task matches its description.

Can I use Solr Semantic Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add griddynamics/rosetta --skill solr-semantic-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/solr-semantic-search, .gemini/skills/solr-semantic-search, .github/skills/solr-semantic-search and .opencode/skills/solr-semantic-search in your project.

What does Solr Semantic Search need to run?

SKILL.md names no scripts, command-line tools or credentials: Solr Semantic Search is instructions for the agent only.

Does Solr Semantic Search access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Solr Semantic Search safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Solr Semantic Search use?

Solr Semantic Search is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Solr Semantic Search use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 43k tokens, read only when the agent opens those files.

What are the alternatives to Solr Semantic Search?

Skills that share tags, products or a category with Solr Semantic Search: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Codebase Management (giancarloerra/SocratiCode, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Solr Semantic Search?

griddynamics (a GitHub organization) maintains it in griddynamics/rosetta, which has 354 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 29, 2026.

Source: griddynamics/rosetta on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.