Agent skill

Chroma

by AlexAI-MCP in AlexAI-MCP/hermes-CCC

Open-source embedding database for RAG — store embeddings, vector search, metadata filtering.

MITAuto-check passedAI & LLM Engineering

Install Chroma

skills CLI
$ npx skills add AlexAI-MCP/hermes-CCC --skill chroma -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexAI-MCP/hermes-CCC chroma --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/chroma .claude/skills/chroma && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chroma
GitHub stars
135
Token cost
~2.3k tokens
SKILL.md length
739 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Open-source embedding database for RAG — store embeddings, vector search, metadata filtering.

  • Tasks that involve Embeddings
  • SKILL.md covers Purpose, Install, Basic Setup and Create a Collection, plus 23 more sections
  • Calls pip and docker

What it does

Chroma is an agent skill from AlexAI-MCP/hermes-CCC. Open-source embedding database for RAG — store embeddings, vector search, metadata filtering. Simple API, scales from notebook to production.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Embeddings. The repository describes itself as: Hermes Agent ported to Claude Code Channel — 46 native skills, no OAuth, no external process. The licence is MIT.

When your agent uses it

  • Tasks that involve Embeddings

Example prompts

  • “/chroma”

Requirements

  • Python 3
  • Docker
  • A credential in YOUR_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 8107e89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chroma loads about 2.3k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 739 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexAI-MCP/hermes-CCC at commit 8107e89, republished under its MIT licence (© AlexAI-MCP). 739 words, ~2,279 tokens.

Download SKILL.mdSave it as .claude/skills/chroma/SKILL.md (or your agent's skills folder).
name
chroma
description
Open-source embedding database for RAG — store embeddings, vector search, metadata filtering. Simple API, scales from notebook to production.
version
1.0.0
author
hermes-CCC (ported from Hermes Agent by NousResearch)
license
MIT

Chroma

Purpose

  • Use this skill to store embeddings and perform semantic retrieval for RAG systems.
  • Prefer Chroma when you want a simple local developer experience with an easy Python API.
  • Chroma works well for notebooks, prototypes, local apps, and moderate production deployments.
  • It supports metadata filtering, server mode, and pluggable embedding functions.

Install

bash
pip install chromadb sentence-transformers
  • Add any extra embedding provider packages you need, such as the OpenAI SDK.

Basic Setup

python
import chromadb

client = chromadb.PersistentClient(path="./chroma_db")
  • PersistentClient stores data on disk.
  • It is a good default for local development and single-node setups.

Create a Collection

python
collection = client.create_collection(name="docs")
  • Use descriptive collection names like docs, papers, tickets, or kb_chunks.
  • A collection is the logical container for your embeddings and payload metadata.

Add Documents

  • Minimal add call:
python
collection.add(
    documents=["Chroma is useful for local RAG.", "Vector search retrieves semantically similar text."],
    metadatas=[{"source": "note1"}, {"source": "note2"}],
    ids=["doc-1", "doc-2"],
)
  • Required arrays must align by index.
  • Keep ids stable if you plan to update or delete records later.

Query

python
results = collection.query(
    query_texts=["How do I store embeddings for RAG?"],
    n_results=5,
)

print(results["documents"])
print(results["metadatas"])
  • query_texts=['...'] is the most common path when Chroma is managing embeddings for you.
  • Start with n_results=5 or 10 for most retrieval experiments.

Basic Collection Lifecycle

  • Create collection
  • Add documents
  • Query for nearest neighbors
  • Update documents when content changes
  • Delete stale records

Get or Create

  • For idempotent startup flows, prefer get_or_create_collection:
python
collection = client.get_or_create_collection(name="docs")
  • This is helpful in services that initialize storage on boot.

Embedding Functions

  • Chroma supports multiple embedding strategies.
  • Common choices include:
  • DefaultEmbeddingFunction
  • SentenceTransformerEmbeddingFunction
  • OpenAIEmbeddingFunction

Default Embedding Function

python
from chromadb.utils.embedding_functions import DefaultEmbeddingFunction

embedding_fn = DefaultEmbeddingFunction()
collection = client.get_or_create_collection(
    name="default-embeddings",
    embedding_function=embedding_fn,
)
  • The default function is convenient for quick experiments.
  • For production, you usually want to control the embedding model explicitly.

SentenceTransformers Embedding Function

python
from chromadb.utils.embedding_functions import SentenceTransformerEmbeddingFunction

embedding_fn = SentenceTransformerEmbeddingFunction(
    model_name="sentence-transformers/all-MiniLM-L6-v2",
)

collection = client.get_or_create_collection(
    name="st-docs",
    embedding_function=embedding_fn,
)
  • This is a good choice for local semantic search with no external API dependency.

OpenAI Embedding Function

python
from chromadb.utils.embedding_functions import OpenAIEmbeddingFunction

embedding_fn = OpenAIEmbeddingFunction(
    api_key="YOUR_API_KEY",
    model_name="text-embedding-3-small",
)

collection = client.get_or_create_collection(
    name="openai-docs",
    embedding_function=embedding_fn,
)
  • Use this when you want consistent hosted embeddings across environments.
  • Track model versions because embedding changes can invalidate similarity behavior.

Add Structured Documents

python
collection.add(
    documents=[
        "Retrieval augmented generation combines retrieval with generation.",
        "Chroma collections can store metadata for filtering.",
    ],
    metadatas=[
        {"source": "paper1", "section": "intro", "year": 2024},
        {"source": "paper1", "section": "methods", "year": 2024},
    ],
    ids=["paper1-intro", "paper1-methods"],
)
  • Include source, section, title, author, or timestamp metadata if you need filtered retrieval.

Metadata Filtering

  • Filter by metadata with where:
python
results = collection.query(
    query_texts=["What does the paper say about filtering?"],
    n_results=5,
    where={"source": "paper1"},
)
  • Example requested pattern:

  • where={'source': 'paper1'}

  • Metadata filters are critical for multi-tenant or source-restricted RAG.

  • Chroma also supports document-side text filtering with where_document.
  • The commonly used operator form is:
python
results = collection.query(
    query_texts=["database"],
    n_results=5,
    where_document={"$contains": "keyword"},
)
  • Some examples on the internet simplify this idea informally as where_document={'...': 'keyword'}.
  • In practice, use the explicit operator form your installed Chroma version documents.

Update Documents

  • Update existing records by ID:
python
collection.update(
    ids=["doc-1"],
    documents=["Chroma stores embeddings persistently for local RAG systems."],
    metadatas=[{"source": "note1", "updated": True}],
)
  • Use stable IDs so updates remain deterministic.

Delete Documents

python
collection.delete(ids=["doc-2"])
  • You can also delete by filter in many workflows:
python
collection.delete(where={"source": "note1"})
  • Deletes are useful for document re-indexing and source cleanup jobs.

Inspect Data

  • Fetch records directly:
python
items = collection.get(ids=["doc-1"])
print(items)
  • Use this for debugging chunk contents, metadata shape, and embedding lifecycle issues.

HTTP Client for Server Mode

  • Chroma can run as a server and be accessed over HTTP.
  • Python client example:
python
import chromadb

client = chromadb.HttpClient(host="localhost", port=8000)
collection = client.get_or_create_collection(name="docs")
  • This is useful when multiple apps need to share one Chroma instance.

Docker Server

  • Start a server with Docker:
bash
docker run -p 8000:8000 chromadb/chroma
  • Pair that with chromadb.HttpClient(host='localhost', port=8000) from Python.
  • For persistent data in Docker, mount a volume rather than relying on container-local storage.
Show full SKILL.md (282 more words)Show less

Retrieval Design Tips

  • Chunk documents before indexing.
  • Keep chunk size and overlap consistent during experiments.
  • Store source metadata so retrieved chunks can be traced back to originals.
  • Log your embedding model name and collection schema.

Chroma Strengths

  • Very easy local setup
  • Clean Python API
  • Good fit for notebook and single-service workflows
  • Flexible embedding function support

Chroma Limitations

  • It is simpler than some production-first vector engines.
  • Large-scale distributed deployments may need a more specialized backend.
  • You should benchmark real workloads before using it as a high-scale production default.

Common Failure Modes

  • Empty search results:

  • confirm documents were added

  • confirm the embedding function is configured as expected

  • lower filtering constraints

  • Embedding mismatch:

  • avoid changing embedding models inside the same collection without re-indexing

  • document the embedding function used for each collection

  • Duplicate records:

  • choose deterministic IDs from source path plus chunk index

  • upsert or update intentionally instead of re-adding blind

  • Server connectivity issues:

  • verify the Docker container is running

  • confirm localhost:8000 is reachable

  • switch from PersistentClient to HttpClient only when appropriate

  • Start local with PersistentClient.
  • Use sentence-transformers for quick offline prototypes.
  • Add metadata filters early if you know you need scoped retrieval.
  • Move to server mode when multiple apps or services need shared access.

When To Use This Skill

  • You are building a local RAG system.
  • You need an easy vector store API inside Python.
  • You want metadata-aware retrieval without a heavy infrastructure stack.
  • You need a bridge from notebook experimentation to a small production deployment.

Quick Reference

  • Install: pip install chromadb sentence-transformers
  • Local client: chromadb.PersistentClient(path="./chroma_db")
  • Add docs: collection.add(documents=[...], metadatas=[...], ids=[...])
  • Query: collection.query(query_texts=["..."], n_results=5)
  • Metadata filter: where={"source": "paper1"}
  • Document text filter: where_document={"$contains": "keyword"}
  • Server client: chromadb.HttpClient(host="localhost", port=8000)
  • Docker: docker run -p 8000:8000 chromadb/chroma

© AlexAI-MCP, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/chroma of AlexAI-MCP/hermes-CCC.

Open the folder on GitHubat commit 8107e89

Compare with similar skills

Chroma next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chroma compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chroma this skillAlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Mashup Mods

    rehan-remade/universal-modder

    Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.

    6.5k GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from AlexAI-MCP/hermes-CCC

All 44 skills in this repo
  • GitHub Code Review

    AlexAI-MCP/hermes-CCC

    Review GitHub pull requests with a findings-first engineering mindset.

    135 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • GitHub PR Workflow

    AlexAI-MCP/hermes-CCC

    Run a disciplined GitHub pull request workflow from branch creation through merge.

    135 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Memory

    AlexAI-MCP/hermes-CCC

    Manage durable project memory for Claude Code. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Route

    AlexAI-MCP/hermes-CCC

    Route Claude Code work by complexity, risk, and tool needs. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Skill

    AlexAI-MCP/hermes-CCC

    Create, improve, inventory, and audit Claude Code skills. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Traj

    AlexAI-MCP/hermes-CCC

    Capture Claude Code interaction trajectories in training-friendly formats.

    135 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Chroma

What does Chroma do?

Open-source embedding database for RAG — store embeddings, vector search, metadata filtering. Chroma is an agent skill from AlexAI-MCP/hermes-CCC. Open-source embedding database for RAG — store embeddings, vector search, metadata filtering.

When should I use Chroma?

Chroma fits situations like: tasks that involve Embeddings.

How do I install Chroma in Claude Code?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill chroma -a claude-code`. Or copy the skill folder (skills/chroma in AlexAI-MCP/hermes-CCC) into .claude/skills/chroma in your project. Claude Code loads it when a task matches its description.

How do I install Chroma in Codex?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill chroma -a codex`. Or copy the skill folder (skills/chroma in AlexAI-MCP/hermes-CCC) into .agents/skills/chroma in your project. Codex loads it when a task matches its description.

Can I use Chroma in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexAI-MCP/hermes-CCC --skill chroma -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chroma, .gemini/skills/chroma, .github/skills/chroma and .opencode/skills/chroma in your project.

What does Chroma need to run?

Going by SKILL.md and its folder, Chroma needs the command-line tools its instructions call (pip and docker). Our summary lists: Python 3; Docker; A credential in YOUR_API_KEY.

Does Chroma access the network?

SKILL.md contains no URLs. Its commands use pip and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Chroma safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chroma use?

Chroma is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chroma use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Chroma?

Skills that share tags, products or a category with Chroma: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chroma?

AlexAI-MCP (a GitHub user) maintains it in AlexAI-MCP/hermes-CCC, which has 135 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on April 8, 2026.

Source: AlexAI-MCP/hermes-CCC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.