Chroma Vector Database
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Walks through turning on semantic search in Open Second Brain: embedding key, sqlite-vec extension, first reindex and an optional periodic refresh, starting from o2b search check.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install itechmeat/open-second-brain embeddings-setup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/embeddings-setup .claude/skills/embeddings-setup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "embeddings-setup" agent skill from https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setup into .claude/skills/embeddings-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "embeddings-setup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install itechmeat/open-second-brain embeddings-setup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/embeddings-setup .agents/skills/embeddings-setup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "embeddings-setup" agent skill from https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setup into .agents/skills/embeddings-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "embeddings-setup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install itechmeat/open-second-brain embeddings-setup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/embeddings-setup .cursor/skills/embeddings-setup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "embeddings-setup" agent skill from https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setup into .cursor/skills/embeddings-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "embeddings-setup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/itechmeat/open-second-brain.git --path skills/embeddings-setup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install itechmeat/open-second-brain embeddings-setup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/embeddings-setup .gemini/skills/embeddings-setup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "embeddings-setup" agent skill from https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setup into .gemini/skills/embeddings-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "embeddings-setup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install itechmeat/open-second-brain embeddings-setupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/embeddings-setup .github/skills/embeddings-setup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "embeddings-setup" agent skill from https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setup into .github/skills/embeddings-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "embeddings-setup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install itechmeat/open-second-brain embeddings-setup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/embeddings-setup .opencode/skills/embeddings-setup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "embeddings-setup" agent skill from https://github.com/itechmeat/open-second-brain/tree/main/skills/embeddings-setup into .opencode/skills/embeddings-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "embeddings-setup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
embeddings-setupWalks through turning on semantic search in Open Second Brain: embedding key, sqlite-vec extension, first reindex and an optional periodic refresh, starting from o2b search check.
Open Second Brain always has keyword search; semantic search over embedded vectors is opt-in. The skill acts proactively when you mention embeddings or semantic search, or when `o2b search check` reports missing pieces. That command comes first, and the agent reads both the report and its recommendations block before branching: a missing embedding key leads to provider setup, an unavailable vector extension leads to a macOS or a Linux fix, and a healthy setup with semantic search still off leads to the first indexing run.
For the provider step it asks which one you want. The default is OpenAI's `text-embedding-3-small` at about $0.02 per 1M tokens, and any OpenAI-compatible endpoint works, including Groq, Together or a local LM Studio server. Settings go into `~/.hermes/.env` or the configured env file and never into a tracked file. The base URL must use https, plain http is accepted only for loopback hosts, and a server on another machine without TLS needs an explicit opt-in. Conceptual questions about embeddings are answered directly.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0f0c9a7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
brewbunsqlite3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.together.xyzFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPEN_SECOND_BRAIN_EMBEDDING_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Open Second Brain Embeddings Setup loads about 2.6k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,341 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
Required env vars (write to `~/.hermes/.env` or the configured envOPEN_SECOND_BRAIN_EMBEDDING_BASE_URL=http://100.64.0.5:1234/v1Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from itechmeat/open-second-brain at commit 0f0c9a7, republished under its MIT licence (© itechmeat). 1,341 words, ~2,619 tokens.
.claude/skills/embeddings-setup/SKILL.md (or your agent's skills folder).Open Second Brain ships with two search paths: keyword-only (always
on, no credentials) and semantic via embedded vectors (opt-in). This
SKILL walks the activation flow for the semantic path. The flow is
proactive — when the user mentions semantic search or o2b search check surfaces missing pieces, take the user through this list
rather than waiting for explicit instruction.
o2b search checko2b search checkThe output names every missing piece and ends with a
recommendations: block listing the exact commands to fix each.
Read both before suggesting next steps. The recommendations field
is also present in the --json output for headless callers.
Branch on what the report shows:
embedding_key: MISSING → go to step 2. The key_sources_checked:
line names where the key was looked for, and key_present_under:
names any other registered provider profile whose env key is set. If
the key sits under another profile, the fix may be pointing
embedding_provider at that profile rather than adding a key.vec_extension: unavailable on macOS → go to step 3.vec_extension: unavailable on Linux → go to step 4.semantic_enabled: false (no embeddings yet) →
go to step 5.Ask the user which provider they want. The default is
text-embedding-3-small from OpenAI (about $0.02 per 1M tokens,
which covers tens of thousands of vault pages). Any OpenAI-compatible
endpoint works — Groq, Together, a local LM Studio server, etc.
Required env vars (write to ~/.hermes/.env or the configured env
file, never to a tracked file):
OPEN_SECOND_BRAIN_EMBEDDING_PROVIDER=openai-compat
OPEN_SECOND_BRAIN_EMBEDDING_MODEL=text-embedding-3-small
OPEN_SECOND_BRAIN_EMBEDDING_KEY=<placeholder; user pastes the key>
# Optional — only when not using OpenAI:
# OPEN_SECOND_BRAIN_EMBEDDING_BASE_URL=https://api.together.xyz/v1The base URL must be https://. Plain http:// works only for a
loopback host (localhost, 127.0.0.1, ::1). When the user's
embedding server runs on ANOTHER machine without TLS - LM Studio on the
Windows host reached from WSL, Ollama on a LAN box, a tailnet 100.x
address - the endpoint is refused until the operator opts in for it
explicitly:
OPEN_SECOND_BRAIN_EMBEDDING_BASE_URL=http://100.64.0.5:1234/v1
OPEN_SECOND_BRAIN_EMBEDDING_ALLOW_INSECURE_HTTP=true(or embedding_allow_insecure_http: true in the o2b config; the
reranker has its own search_rerank_allow_insecure_http). Tell the user
what it means before setting it: chunk text and the API key cross the
network unencrypted, so it belongs only on a network they trust. The
opt-out never applies to a base URL that comes from a provider profile
registered with o2b search provider add; set the URL in config or env
instead. Every run that uses it prints one stderr warning per endpoint.
Never invent or echo the key. Write a placeholder, then ask the
user to paste their key in place of it. Recheck with o2b search check after the user confirms.
Spend estimates name their price source: builtin (the built-in table;
the local provider is free), operator (declared) or unknown. A
model the table does not list (most local servers, newer or niche
models) is unknown, and its estimates print price unknown. When the
user has set a positive embedding_cost_gate_usd, an unknown price
refuses embedding runs with EMBEDDING_COST_UNPRICED until the price is
declared, and it also refuses the query embeds of searches that do not
run at local reach (MCP tools and hooks): a hybrid search falls back to
keyword-only and names semantic-cost-unpriced in its trail. Ask the
user for the provider's rate (USD per million tokens)
and write the pair beside the other embedding settings:
OPEN_SECOND_BRAIN_EMBEDDING_PRICE_MODEL=nomic-embed-text:latest
OPEN_SECOND_BRAIN_EMBEDDING_PRICE_USD_PER_MTOK=0(or embedding_price_model and embedding_price_usd_per_mtok in the
o2b config). Set both or neither. 0 declares the model free; only use
it when the user confirms the endpoint costs nothing, because a
loopback URL can still be a paid proxy. The pair names one model, so
after a model switch o2b search check flags the stale declaration
until it is updated. Declaring a price never triggers a reindex.
A search query longer than the model's input window is cut to it before
the embed and the trail names semantic-query-truncated. The window
comes from the curated model table; for a model it does not list the
window is unknown and nothing is cut, so a serving stack that truncates
silently or refuses long inputs does so unannounced. When the user knows
the model's window (in its own tokens), declare it:
OPEN_SECOND_BRAIN_EMBEDDING_INPUT_WINDOW_TOKENS=512(or embedding_input_window_tokens in the o2b config). It is refused
for the local provider, which has no window.
When the user's serving stack needs a request field the provider does
not send, put a JSON object in embedding_extra_body (env
OPEN_SECOND_BRAIN_EMBEDDING_EXTRA_BODY). Only openai-compat sends it,
and model, input and encoding_format are refused by name. A
dimensions field needs embedding_dimension set to the same width:
OPEN_SECOND_BRAIN_EMBEDDING_EXTRA_BODY='{"dimensions": 512}'
OPEN_SECOND_BRAIN_EMBEDDING_DIM=512The extra body is not part of the embedding identity, so changing it
never triggers a reindex; a changed embedding_dimension does. Some
fields shape the vectors the provider returns (a provider task,
input_type, normalize or truncate switch): after changing such a
field, rebuild the vectors with o2b search reindex --embeddings so new
query vectors stay in the space of the stored passages (o2b search index --force keeps the stored vector of every unchanged chunk).
Apple ships /usr/lib/libsqlite3.dylib with
SQLITE_OMIT_LOAD_EXTENSION, so the optional sqlite-vec extension
cannot load against the system SQLite. Homebrew's sqlite formula
is built with extension loading enabled.
brew install sqliteThe o2b wrapper auto-detects Homebrew SQLite on Darwin and exports
DYLD_LIBRARY_PATH for bun:sqlite on the next invocation. No
manual patching of the wrapper script is needed (the v0.10.5 shim
scripts/_macos-sqlite.sh handles it). Verify with:
o2b search checkExpect vec_extension: loaded.
sqlite-vec ships as an optional dependency. If the install path
missed it, force a rebuild:
bun pm ls | grep sqlite-vec
bun install --forceIf the package is present but still does not load, the Linux build
of libsqlite3 may have been compiled without LOAD_EXTENSION.
Capture sqlite3 :memory: "PRAGMA compile_options;" and surface it
to the user — they will know whether their distro's SQLite is
unusual.
o2b search reindex --embeddingsThis walks the vault, chunks every Markdown file, and computes embeddings. Cost scales linearly with vault size — a 150-file vault typically lands in a few seconds.
When the index already exists and only some chunks lack vectors, price the work first with the dry run, then apply it:
o2b search vector-backfill
o2b search vector-backfill --apply--path <prefix> (repeatable) limits the census, the estimate and the
spend to one part of the vault. For example,
o2b search vector-backfill --path Brain/preferences/ --path Brain/retired/ --apply
embeds just the belief notes, which is what the semantic query mode of
brain_context_pack reads. Later edits re-embed only the chunks that
changed; unchanged chunks keep their vectors.
Verify semantic search works:
o2b search "preferences for code review"The output rows carry (semantic) in the source column when the
match came from the vector index.
The vault keeps changing — other agents add notes, the user edits existing ones. Without periodic refresh the embeddings drift behind the keyword index.
o2b search reindex --cron-templatePrints a watchdog script (a heredoc that, when the operator runs
the printed block, lands at ~/.local/bin/osb-reindex.sh), a
native crontab line, and a hermes cron create command. Pure
stdout — the verb writes nothing on its own; pick the path that
matches the host and paste the relevant section into your
shell / crontab. Recommended cadence:
Override with --interval 6h (or any <N>m|h|d value below 60m / 24h
/ unlimited days).
Only the agent that runs reindex needs
OPEN_SECOND_BRAIN_EMBEDDING_KEY. Read-only consumers (sibling
Claude Code / Codex / OpenClaw sessions on the same vault) query
the already-computed vectors and never see the key. When designating
the reindex owner, the canonical choice is Hermes (it carries the
cron infrastructure); the others stay credential-free.
brew install on the user's behalf. The SKILL
prints the command; the user runs it.--cron-template recipe is the endorsed automation path.© itechmeat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/embeddings-setup of itechmeat/open-second-brain.
Open the folder on GitHubat commit 0f0c9a7
Open Second Brain Embeddings Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Open Second Brain Embeddings Setup this skillitechmeat/open-second-brain | 430 | — | ~2.6k | Automated safety check: Warn | MIT | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Codebase Managementgiancarloerra/SocratiCode | 3.3k | 1 repos | ~1.8k | Automated safety check: Pass | AGPL-3.0 | |
| RAG ArchitectJeffallan/claude-skills | 12k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Langchain RAGlangchain-ai/langchain-skills | 1.3k | 1 repos | ~3.9k | Automated safety check: Pass | MIT | |
| Cortexdbliliang-cn/cortexdb | 273 | — | ~18k | Automated safety check: Warn | MIT |
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
giancarloerra/SocratiCode
Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
liliang-cn/cortexdb
Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.
rrezartprebreza/spring-boot-skills
A skill your agent uses when integrating LLMs, chat clients, embeddings, RAG pipelines, or AI agents into Spring Boot.
itechmeat/open-second-brain
Sets up, evaluates or turns off Open Second Brain's optional decision-model feature, starting from the o2b decision-model check command and a provider route you choose.
itechmeat/open-second-brain
Teaches the agent to detect whether the codegraph symbol graph is available and use it for callers, callees and impact questions, without installing or writing anything for it.
itechmeat/open-second-brain
Adds, renames, reviews and explains Brain schema tokens, aliases, prefixes and link types in Open Second Brain vaults, previewing each change before applying it.
itechmeat/open-second-brain
Lets an agent read, write and maintain its own memory in a Brain folder of an Obsidian-compatible vault, while keeping the user's own notes read-only.
Categories
Walks through turning on semantic search in Open Second Brain: embedding key, sqlite-vec extension, first reindex and an optional periodic refresh, starting from o2b search check. Open Second Brain always has keyword search; semantic search over embedded vectors is opt-in. The skill acts proactively when you mention embeddings or semantic search, or when `o2b search check` reports missing pieces.
Open Second Brain Embeddings Setup fits situations like: enabling semantic search in an Open Second Brain vault; fixing an embedding key or sqlite-vec problem reported by o2b search check; choosing an embedding provider or a local OpenAI-compatible endpoint.
Run `npx skills add itechmeat/open-second-brain --skill embeddings-setup -a claude-code`. Or copy the skill folder (skills/embeddings-setup in itechmeat/open-second-brain) into .claude/skills/embeddings-setup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add itechmeat/open-second-brain --skill embeddings-setup -a codex`. Or copy the skill folder (skills/embeddings-setup in itechmeat/open-second-brain) into .agents/skills/embeddings-setup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add itechmeat/open-second-brain --skill embeddings-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/embeddings-setup, .gemini/skills/embeddings-setup, .github/skills/embeddings-setup and .opencode/skills/embeddings-setup in your project.
Going by SKILL.md and its folder, Open Second Brain Embeddings Setup needs the command-line tools its instructions call (brew, bun and sqlite3) and credentials named OPEN_SECOND_BRAIN_EMBEDDING_KEY. Our summary lists: The o2b command-line tool from Open Second Brain; An embedding API key or an OpenAI-compatible embedding endpoint.
SKILL.md names 1 domain. In commands or code: api.together.xyz; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): links to a raw public ip address. Read the flagged lines before installing; the check is not a guarantee either way.
Open Second Brain Embeddings Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Open Second Brain Embeddings Setup: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars), RAG Architect (Jeffallan/claude-skills, 12k stars) and Langchain RAG (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
itechmeat (a GitHub user) maintains it in itechmeat/open-second-brain, which has 430 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.
Source: itechmeat/open-second-brain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.