Chroma Vector Database
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install griddynamics/rosetta solr-semantic-search --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .claude/skills/solr-semantic-search && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "solr-semantic-search" agent skill from https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-search into .claude/skills/solr-semantic-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "solr-semantic-search", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-searchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install griddynamics/rosetta solr-semantic-search --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .agents/skills/solr-semantic-search && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "solr-semantic-search" agent skill from https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-search into .agents/skills/solr-semantic-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "solr-semantic-search", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install griddynamics/rosetta solr-semantic-search --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .cursor/skills/solr-semantic-search && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "solr-semantic-search" agent skill from https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-search into .cursor/skills/solr-semantic-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "solr-semantic-search", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/griddynamics/rosetta.git --path plugins/core-claude/skills/solr-semantic-search--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install griddynamics/rosetta solr-semantic-search --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .gemini/skills/solr-semantic-search && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "solr-semantic-search" agent skill from https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-search into .gemini/skills/solr-semantic-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "solr-semantic-search", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install griddynamics/rosetta solr-semantic-searchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .github/skills/solr-semantic-search && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "solr-semantic-search" agent skill from https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-search into .github/skills/solr-semantic-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "solr-semantic-search", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add griddynamics/rosetta --skill solr-semantic-search -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install griddynamics/rosetta solr-semantic-search --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/griddynamics/rosetta.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/core-claude/skills/solr-semantic-search .opencode/skills/solr-semantic-search && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "solr-semantic-search" agent skill from https://github.com/griddynamics/rosetta/tree/main/plugins/core-claude/skills/solr-semantic-search into .opencode/skills/solr-semantic-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "solr-semantic-search", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
solr-semantic-searchTo build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.
Solr Semantic Search is an agent skill from griddynamics/rosetta. To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `README.md`, `references/01-architecture.md` and `references/02-concept-indexing.md`).
It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: An instruction layer for AI coding agent. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5441232. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Solr Semantic Search loads about 1.9k tokens when it runs, and up to ~45k if it reads all its reference files. Until then it costs about 27 tokens; SKILL.md has 750 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from griddynamics/rosetta at commit 5441232, republished under its Apache-2.0 licence (© griddynamics). 750 words, ~1,867 tokens.
.claude/skills/solr-semantic-search/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.<solr-semantic-search>
<role>
You are a senior Apache Solr engineer who designs, builds, debugs, and extends phrase-tagging semantic search on Solr 9.x: decomposing natural-language queries into structured concepts via dictionary lookup, resolving path ambiguity in a tag graph, and assembling precise multi-field Solr queries. This is lexical, not vector/embedding, semantic search.
</role>
<when_to_use_skill>
Concept tagging, query understanding, taxonomy-driven search, structured Brand/Line/Model recognition, shingle-based matching, multi-word synonyms, path resolution, fuzzy-phrase-to-structured-query extraction. Traditional Solr query work and vector/kNN semantic search → solr-query skill. Custom plugins this architecture relies on → solr-extending skill.
</when_to_use_skill>
<core_concepts>
Three independently testable layers, separated by stable interfaces (ProducedTag, StagedTag, SmQuery):
ProducedTag list (token, position, type, matched fields+weights).Sm query, apply dependency groups and min-should-match, then translate to a Solr query against the catalog.This SKILL.md is a router. For any non-trivial question, read the relevant references/ file before answering — references hold the examples, schemas, code, and decision tables and are not duplicated here.
</core_concepts>
<references>
| When the user asks about… | Read |
|---|---|
| Architecture overview, the three layers, data flow | READ SKILL FILE references/01-architecture.md |
| Concept collection schema, building it from source data, indexing handler | READ SKILL FILE references/02-concept-indexing.md |
| Phrase tagging mechanics: shingles, lookup, scoring, multi-language, fuzzy/word-break/prefix | READ SKILL FILE references/03-tagging.md |
| Graph construction (JGraphT), vertices/edges, paths, quasi-positions for multi-word syns | READ SKILL FILE references/04-graph-paths.md |
| Ambiguity resolution between competing interpretations (Path vs Shingle resolvers) | READ SKILL FILE references/05-ambiguity-resolution.md |
| Building the final Solr query from tagged paths, Sm query model, dependency groups | READ SKILL FILE references/06-query-building.md |
| Adapting this to a new domain: schema design, concept sources, stages config | READ SKILL FILE references/07-applying-to-domain.md |
| Sm* query model implementation — full code for SmQuery/SmBoolean/SmTerm and the Solr translator fabric | READ SKILL FILE references/08-query-model-implementation.md |
</references>
<when_to_choose>
This is a heavyweight architecture. It is the right tool when the domain has well-defined concepts (products, models, attributes) with known synonyms, queries must be understood structurally ("what is the Brand? Line? attribute?"), vector search yields too many false positives for the required precision, and authoritative taxonomies exist to extract concepts from.
It is the wrong tool when the domain is open-ended natural language (use embeddings), there are no curated concept dictionaries, or only fuzzy retrieval is needed without structural understanding.
</when_to_choose>
<mental_model>
USER PHRASE: "sony wh-1000xm5 ear pads"
──► LAYER 1 TAGGING: tokens → shingles → concept-index lookup → ProducedTag list
──► LAYER 2 GRAPH: tags→edges, positions→vertices; K-shortest paths; resolve ambiguity
──► LAYER 3 QUERY BUILDING: per path build Sm query, dependency groups, min-should-match → Solr query
──► SOLR SEARCH against the catalog ──► RESULTSWhy it beats naive eDisMax, three problems:
qf.MULTI_SYN tag spanning both positions, preserving the structure eDisMax pf loses.BrandLineModelProcessor) checks recognized Brand/Line/Model tags against a canonical CatalogProvider, drops invalid combos, and turns valid ones into structured filters (brand_id_s:SONY AND line_id_s:WH AND model_id_s:WH-1000XM5).</mental_model>
<key_data_types>
Token — analyzed phrase token (term + position + lang)
Shingle — N consecutive tokens treated as a unit
ProducedTag — recognized concept: token, start/end position, relation type, matched fields (with weights)
StagedTag — ProducedTag enriched with staging info (fields, boosts, dependencies) for a search stage
SmQuery — abstract semantic query (SmBoolean/SmTerm/SmBoost/…) translated to a Lucene/Solr Query
TagType — CONCEPT | SYN | MULTI_SYN | SPELL | PREFIX | RECOGNIZED_PRODUCT (validated Brand/Line/Model)
StageConfig — per-stage config (fields, min-should-match, min-pattern-score, ambiguity resolver, …)The tagger is a Solr request handler at /semanticTagGraph (params: q, lang, source, fuzzy, wordBreak, prefix, maxShingleLength, debug, dot). It returns tokens, tags (each with token, start/end, relation, entryFields weights), unrecognized, and a graphviz tagsDot. Downstream runs ambiguity resolution → path finding → query building, then hits the catalog collection.
</key_data_types>
<anti_patterns>
maxShingleLength — shingles 1..10 over a 10-token phrase is O(N²); cap at 4–5.SynonymsStorage once at startup.</anti_patterns>
<solr_10_deltas>
The architecture is Solr 9.x-tested. On Solr 10: BlockJoinParentQParser API stable; JGraphT is an external dep — pin to your build; custom RequestHandler/SearchComponent base classes unchanged; concept indexing via TermsComponent works the same, with minor changes to the /admin/luke response shape. On Solr 9.x these differences will not bite.
</solr_10_deltas>
</solr-semantic-search>
© griddynamics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (references) in plugins/core-claude/skills/solr-semantic-search of griddynamics/rosetta.
Open the folder on GitHubat commit 5441232
Solr Semantic Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Solr Semantic Search this skillgriddynamics/rosetta | 354 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Codebase Managementgiancarloerra/SocratiCode | 3.3k | 1 repos | ~1.8k | Automated safety check: Pass | AGPL-3.0 | |
| Sentence-Transformers Training Routerhuggingface/skills | 11k | 1 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
giancarloerra/SocratiCode
Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
rehan-remade/universal-modder
Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.
griddynamics/rosetta
Collect GitHub repo health/usage stats into merged JSON. An agent skill from griddynamics/rosetta.
griddynamics/rosetta
Workflow for onboarding an external private library so AI can use it without source access.
griddynamics/rosetta
To connect Rosetta with Grid Dynamics SpecFlow MCP; only when SpecFlow is mentioned and the MCP is installed.
griddynamics/rosetta
To build an AI harness: run, observe, validate, automate repeated work faster — CLI/MCP actions, devcontainers, skills, subagents, hooks, pipelines, automations.
griddynamics/rosetta
To orchestrate parallel coding-agent farms (Claude, Codex, Copilot, Gemini, etc.) on isolated git worktrees.
griddynamics/rosetta
To author, register, and test Rosetta hooks, add a SemanticKind, or debug a hook that won't fire.
Categories
To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths. Solr Semantic Search is an agent skill from griddynamics/rosetta. To build Solr phrase-tagging semantic search: concept tagging, taxonomy, graph paths.
Solr Semantic Search fits situations like: tasks that involve Embeddings.
Run `npx skills add griddynamics/rosetta --skill solr-semantic-search -a claude-code`. Or copy the skill folder (plugins/core-claude/skills/solr-semantic-search in griddynamics/rosetta) into .claude/skills/solr-semantic-search in your project. Claude Code loads it when a task matches its description.
Run `npx skills add griddynamics/rosetta --skill solr-semantic-search -a codex`. Or copy the skill folder (plugins/core-claude/skills/solr-semantic-search in griddynamics/rosetta) into .agents/skills/solr-semantic-search in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add griddynamics/rosetta --skill solr-semantic-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/solr-semantic-search, .gemini/skills/solr-semantic-search, .github/skills/solr-semantic-search and .opencode/skills/solr-semantic-search in your project.
SKILL.md names no scripts, command-line tools or credentials: Solr Semantic Search is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Solr Semantic Search is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 43k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Solr Semantic Search: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Codebase Management (giancarloerra/SocratiCode, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
griddynamics (a GitHub organization) maintains it in griddynamics/rosetta, which has 354 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 29, 2026.
Source: griddynamics/rosetta on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.