Agent skill

Chroma Vector Database

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

MITAuto-check passedAI & LLM Engineering

Install Chroma Vector Database

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill chroma -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs chroma --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/15-rag/chroma .claude/skills/chroma && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chroma
GitHub stars
13k
Used in
7 other repos
Token cost
~2.3k tokens
SKILL.md length
237 words
Files
2 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

  • Works in 6 steps: Create collection → Add documents → Query (similarity search) → …
  • Building a RAG prototype that needs a local vector store
  • SKILL.md covers When to use Chroma, Quick start, Core operations and Persistent storage, plus 8 more sections
  • Calls pip and npm

What it does

Chroma is an open-source embedding database, and the skill describes when it fits: RAG applications, a local or self-hosted vector store, prototyping in notebooks, semantic search over documents and storing embeddings with metadata. It also names alternatives for other needs, such as Pinecone for managed scaling, FAISS for pure similarity search, Weaviate and Qdrant.

Setup is `pip install chromadb` followed by a Python client. The core operations are creating a collection, adding documents with or without IDs and metadata, querying by text for similar documents with metadata filters, getting by ID, updating content and deleting. A persistent client with a local path keeps data on disk. Documents are embedded with Sentence Transformers by default, and the skill shows how to switch to OpenAI or HuggingFace embedding functions or write a custom one. A reference file covers integration with other tools.

When your agent uses it

  • Building a RAG prototype that needs a local vector store
  • Adding semantic search over a folder of documents
  • Persisting embeddings between runs and filtering results by metadata

Example prompts

  • “Set up a persistent Chroma collection and add these markdown docs with source metadata.”
  • “Query the collection for the five closest chunks to this question and filter by source.”
  • “Switch the collection to OpenAI embeddings instead of the default model.”

Requirements

  • Python with the `chromadb` package
  • Optional: an OpenAI or Hugging Face API key for those embedding functions

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Create collection
  2. Add documents
  3. Query (similarity search)
  4. Get documents
  5. Update documents
  6. Delete documents

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • docs.trychroma.com
    • discord.gg

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chroma Vector Database loads about 2.3k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 237 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 237 words, ~2,291 tokens.

Download SKILL.mdSave it as .claude/skills/chroma/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
chroma
description
Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects.
version
1.0.0
author
Orchestra Research
license
MIT
tags
RAG, Chroma, Vector Database, Embeddings, Semantic Search, Open Source, Self-Hosted, Document Retrieval, Metadata Filtering
dependencies
chromadb, sentence-transformers

Chroma - Open-Source Embedding Database

The AI-native database for building LLM applications with memory.

When to use Chroma

Use Chroma when:

  • Building RAG (retrieval-augmented generation) applications
  • Need local/self-hosted vector database
  • Want open-source solution (Apache 2.0)
  • Prototyping in notebooks
  • Semantic search over documents
  • Storing embeddings with metadata

Metrics:

  • 24,300+ GitHub stars
  • 1,900+ forks
  • v1.3.3 (stable, weekly releases)
  • Apache 2.0 license

Use alternatives instead:

  • Pinecone: Managed cloud, auto-scaling
  • FAISS: Pure similarity search, no metadata
  • Weaviate: Production ML-native database
  • Qdrant: High performance, Rust-based

Quick start

Installation
bash
# Python
pip install chromadb

# JavaScript/TypeScript
npm install chromadb @chroma-core/default-embed
Basic usage (Python)
python
import chromadb

# Create client
client = chromadb.Client()

# Create collection
collection = client.create_collection(name="my_collection")

# Add documents
collection.add(
    documents=["This is document 1", "This is document 2"],
    metadatas=[{"source": "doc1"}, {"source": "doc2"}],
    ids=["id1", "id2"]
)

# Query
results = collection.query(
    query_texts=["document about topic"],
    n_results=2
)

print(results)

Core operations

1. Create collection
python
# Simple collection
collection = client.create_collection("my_docs")

# With custom embedding function
from chromadb.utils import embedding_functions

openai_ef = embedding_functions.OpenAIEmbeddingFunction(
    api_key="your-key",
    model_name="text-embedding-3-small"
)

collection = client.create_collection(
    name="my_docs",
    embedding_function=openai_ef
)

# Get existing collection
collection = client.get_collection("my_docs")

# Delete collection
client.delete_collection("my_docs")
2. Add documents
python
# Add with auto-generated IDs
collection.add(
    documents=["Doc 1", "Doc 2", "Doc 3"],
    metadatas=[
        {"source": "web", "category": "tutorial"},
        {"source": "pdf", "page": 5},
        {"source": "api", "timestamp": "2025-01-01"}
    ],
    ids=["id1", "id2", "id3"]
)

# Add with custom embeddings
collection.add(
    embeddings=[[0.1, 0.2, ...], [0.3, 0.4, ...]],
    documents=["Doc 1", "Doc 2"],
    ids=["id1", "id2"]
)
python
# Basic query
results = collection.query(
    query_texts=["machine learning tutorial"],
    n_results=5
)

# Query with filters
results = collection.query(
    query_texts=["Python programming"],
    n_results=3,
    where={"source": "web"}
)

# Query with metadata filters
results = collection.query(
    query_texts=["advanced topics"],
    where={
        "$and": [
            {"category": "tutorial"},
            {"difficulty": {"$gte": 3}}
        ]
    }
)

# Access results
print(results["documents"])      # List of matching documents
print(results["metadatas"])      # Metadata for each doc
print(results["distances"])      # Similarity scores
print(results["ids"])            # Document IDs
4. Get documents
python
# Get by IDs
docs = collection.get(
    ids=["id1", "id2"]
)

# Get with filters
docs = collection.get(
    where={"category": "tutorial"},
    limit=10
)

# Get all documents
docs = collection.get()
5. Update documents
python
# Update document content
collection.update(
    ids=["id1"],
    documents=["Updated content"],
    metadatas=[{"source": "updated"}]
)
6. Delete documents
python
# Delete by IDs
collection.delete(ids=["id1", "id2"])

# Delete with filter
collection.delete(
    where={"source": "outdated"}
)

Persistent storage

python
# Persist to disk
client = chromadb.PersistentClient(path="./chroma_db")

collection = client.create_collection("my_docs")
collection.add(documents=["Doc 1"], ids=["id1"])

# Data persisted automatically
# Reload later with same path
client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_collection("my_docs")

Embedding functions

Default (Sentence Transformers)
python
# Uses sentence-transformers by default
collection = client.create_collection("my_docs")
# Default model: all-MiniLM-L6-v2
OpenAI
python
from chromadb.utils import embedding_functions

openai_ef = embedding_functions.OpenAIEmbeddingFunction(
    api_key="your-key",
    model_name="text-embedding-3-small"
)

collection = client.create_collection(
    name="openai_docs",
    embedding_function=openai_ef
)
HuggingFace
python
huggingface_ef = embedding_functions.HuggingFaceEmbeddingFunction(
    api_key="your-key",
    model_name="sentence-transformers/all-mpnet-base-v2"
)

collection = client.create_collection(
    name="hf_docs",
    embedding_function=huggingface_ef
)
Custom embedding function
python
from chromadb import Documents, EmbeddingFunction, Embeddings

class MyEmbeddingFunction(EmbeddingFunction):
    def __call__(self, input: Documents) -> Embeddings:
        # Your embedding logic
        return embeddings

my_ef = MyEmbeddingFunction()
collection = client.create_collection(
    name="custom_docs",
    embedding_function=my_ef
)

Metadata filtering

python
# Exact match
results = collection.query(
    query_texts=["query"],
    where={"category": "tutorial"}
)

# Comparison operators
results = collection.query(
    query_texts=["query"],
    where={"page": {"$gt": 10}}  # $gt, $gte, $lt, $lte, $ne
)

# Logical operators
results = collection.query(
    query_texts=["query"],
    where={
        "$and": [
            {"category": "tutorial"},
            {"difficulty": {"$lte": 3}}
        ]
    }  # Also: $or
)

# Contains
results = collection.query(
    query_texts=["query"],
    where={"tags": {"$in": ["python", "ml"]}}
)

LangChain integration

python
from langchain_chroma import Chroma
from langchain_openai import OpenAIEmbeddings
from langchain.text_splitter import RecursiveCharacterTextSplitter

# Split documents
text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000)
docs = text_splitter.split_documents(documents)

# Create Chroma vector store
vectorstore = Chroma.from_documents(
    documents=docs,
    embedding=OpenAIEmbeddings(),
    persist_directory="./chroma_db"
)

# Query
results = vectorstore.similarity_search("machine learning", k=3)

# As retriever
retriever = vectorstore.as_retriever(search_kwargs={"k": 5})

LlamaIndex integration

python
from llama_index.vector_stores.chroma import ChromaVectorStore
from llama_index.core import VectorStoreIndex, StorageContext
import chromadb

# Initialize Chroma
db = chromadb.PersistentClient(path="./chroma_db")
collection = db.get_or_create_collection("my_collection")

# Create vector store
vector_store = ChromaVectorStore(chroma_collection=collection)
storage_context = StorageContext.from_defaults(vector_store=vector_store)

# Create index
index = VectorStoreIndex.from_documents(
    documents,
    storage_context=storage_context
)

# Query
query_engine = index.as_query_engine()
response = query_engine.query("What is machine learning?")

Server mode

python
# Run Chroma server
# Terminal: chroma run --path ./chroma_db --port 8000

# Connect to server
import chromadb
from chromadb.config import Settings

client = chromadb.HttpClient(
    host="localhost",
    port=8000,
    settings=Settings(anonymized_telemetry=False)
)

# Use as normal
collection = client.get_or_create_collection("my_docs")

Best practices

  1. Use persistent client - Don't lose data on restart
  2. Add metadata - Enables filtering and tracking
  3. Batch operations - Add multiple docs at once
  4. Choose right embedding model - Balance speed/quality
  5. Use filters - Narrow search space
  6. Unique IDs - Avoid collisions
  7. Regular backups - Copy chroma_db directory
  8. Monitor collection size - Scale up if needed
  9. Test embedding functions - Ensure quality
  10. Use server mode for production - Better for multi-user

Performance

OperationLatencyNotes
Add 100 docs~1-3sWith embedding
Query (top 10)~50-200msDepends on collection size
Metadata filter~10-50msFast with proper indexing

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in 15-rag/chroma of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/integration.md

Open the folder on GitHubat commit 773a529

Used in 7 other repositories

We found 8 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 7 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Chroma Vector Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chroma Vector Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chroma Vector Database this skillOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Langchain RAGlangchain-ai/langchain-skills1.3k—~3.9kAutomated safety check: PassMIT
Retail Product Search Agentgoogle/adk-recipes10k—~3kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k—~2kAutomated safety check: PassMIT
DBoracle/skills876—~1.4kAutomated safety check: PassUPL-1.0
LLM Opsdavila7/claude-code-templates32k3 repos~2kAutomated safety check: PassMIT

Similar skills

  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Retail Product Search Agent

    google/adk-recipes

    Official

    Builds a retail product search agent on Google Cloud, from catalog ingestion into BigQuery and Vector Search to ADK scaffolding, evaluation and Cloud Run deployment.

    10k GitHub stars~3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • DB

    oracle/skills

    Official

    Oracle Database guidance for SQL, PL/SQL, SQLcl, ORDS, Oracle Vector SDK, administration, app development, performance, security, migrations, and agent-safe database workflows.

    876 GitHub stars~1.4k tokensUpdated 3 days ago
    DatabasesAuto-check passed
  • LLM Ops

    davila7/claude-code-templates

    LLM Operations -- RAG, embeddings, vector databases, fine-tuning, prompt engineering avancado, custos de LLM, evals de qualidade e arquiteturas de IA para producao.

    32k GitHub starsUsed in 3 repos~2k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Skills

    llama-farm/llamafarm

    RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.

    836 GitHub stars~1.3k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 6 repos~1.3k tokens
    Auto-check passed

Questions about Chroma Vector Database

What does Chroma Vector Database do?

Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects. Chroma is an open-source embedding database, and the skill describes when it fits: RAG applications, a local or self-hosted vector store, prototyping in notebooks, semantic search over documents and storing embeddings with metadata. It also names alternatives for other needs, such as Pinecone for managed scaling, FAISS for pure similarity search, Weaviate and Qdrant.

When should I use Chroma Vector Database?

Chroma Vector Database fits situations like: building a RAG prototype that needs a local vector store; adding semantic search over a folder of documents; persisting embeddings between runs and filtering results by metadata.

How do I install Chroma Vector Database in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill chroma -a claude-code`. Or copy the skill folder (15-rag/chroma in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/chroma in your project. Claude Code loads it when a task matches its description.

How do I install Chroma Vector Database in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill chroma -a codex`. Or copy the skill folder (15-rag/chroma in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/chroma in your project. Codex loads it when a task matches its description.

Can I use Chroma Vector Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill chroma -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chroma, .gemini/skills/chroma, .github/skills/chroma and .opencode/skills/chroma in your project.

What does Chroma Vector Database need to run?

Going by SKILL.md and its folder, Chroma Vector Database needs the command-line tools its instructions call (pip and npm). Our summary lists: Python with the `chromadb` package; Optional: an OpenAI or Hugging Face API key for those embedding functions.

Does Chroma Vector Database access the network?

SKILL.md names 3 domains. As links in the text: github.com, docs.trychroma.com and discord.gg. This is read from the text; nothing was executed.

Is Chroma Vector Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chroma Vector Database use?

Chroma Vector Database is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chroma Vector Database use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 193 tokens, read only when the agent opens those files.

What are the alternatives to Chroma Vector Database?

Skills that share tags, products or a category with Chroma Vector Database: Langchain RAG (langchain-ai/langchain-skills, 1.3k stars), Retail Product Search Agent (google/adk-recipes, 10k stars), RAG Architect (Jeffallan/claude-skills, 12k stars) and DB (oracle/skills, 876 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chroma Vector Database?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.