Agent skill

Open Second Brain Embeddings Setup

by itechmeat in itechmeat/open-second-brain

Walks through turning on semantic search in Open Second Brain: embedding key, sqlite-vec extension, first reindex and an optional periodic refresh, starting from o2b search check.

MITAuto-check: warningsAI & LLM Engineering

Install Open Second Brain Embeddings Setup

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add itechmeat/open-second-brain --skill embeddings-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install itechmeat/open-second-brain embeddings-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/itechmeat/open-second-brain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/embeddings-setup .claude/skills/embeddings-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
embeddings-setup
GitHub stars
430
Token cost
~2.6k tokens
SKILL.md length
1,341 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Walks through turning on semantic search in Open Second Brain: embedding key, sqlite-vec extension, first reindex and an optional periodic refresh, starting from o2b search check.

  • Works in 6 steps: Always start with o2b search check → Provider and API key → macOS: install Homebrew SQLite → …
  • Enabling semantic search in an Open Second Brain vault
  • SKILL.md covers Step 1 — Always start with o2b…, Step 2 — Provider and API key, Step 3 — macOS: install… and Step 4 — Linux: reconfirm the…, plus 4 more sections
  • Calls brew, bun and sqlite3; reaches api.together.xyz; needs OPEN_SECOND_BRAIN_EMBEDDING_KEY

What it does

Open Second Brain always has keyword search; semantic search over embedded vectors is opt-in. The skill acts proactively when you mention embeddings or semantic search, or when `o2b search check` reports missing pieces. That command comes first, and the agent reads both the report and its recommendations block before branching: a missing embedding key leads to provider setup, an unavailable vector extension leads to a macOS or a Linux fix, and a healthy setup with semantic search still off leads to the first indexing run.

For the provider step it asks which one you want. The default is OpenAI's `text-embedding-3-small` at about $0.02 per 1M tokens, and any OpenAI-compatible endpoint works, including Groq, Together or a local LM Studio server. Settings go into `~/.hermes/.env` or the configured env file and never into a tracked file. The base URL must use https, plain http is accepted only for loopback hosts, and a server on another machine without TLS needs an explicit opt-in. Conceptual questions about embeddings are answered directly.

When your agent uses it

  • Enabling semantic search in an Open Second Brain vault
  • Fixing an embedding key or sqlite-vec problem reported by o2b search check
  • Choosing an embedding provider or a local OpenAI-compatible endpoint

Example prompts

  • “Set up semantic search for my Open Second Brain vault.”
  • “o2b search check says vec_extension is unavailable on Linux; walk me through the fix.”
  • “Point the embeddings at my local LM Studio server and run the first reindex.”

Requirements

  • The o2b command-line tool from Open Second Brain
  • An embedding API key or an OpenAI-compatible embedding endpoint

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Always start with o2b search check
  2. Provider and API key
  3. macOS: install Homebrew SQLite
  4. Linux: reconfirm the optional dependency
  5. Compute the first vectors
  6. Offer periodic reindex

What it can do on your machine

Read from SKILL.md and the folder at commit 0f0c9a7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • brew
    • bun
    • sqlite3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.together.xyz

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPEN_SECOND_BRAIN_EMBEDDING_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Open Second Brain Embeddings Setup loads about 2.6k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,341 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • NoteMentions a .env fileSKILL.md:48
    Required env vars (write to `~/.hermes/.env` or the configured env
  • WarningLinks to a raw public IP addressSKILL.md:67
    OPEN_SECOND_BRAIN_EMBEDDING_BASE_URL=http://100.64.0.5:1234/v1

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from itechmeat/open-second-brain at commit 0f0c9a7, republished under its MIT licence (© itechmeat). 1,341 words, ~2,619 tokens.

Download SKILL.mdSave it as .claude/skills/embeddings-setup/SKILL.md (or your agent's skills folder).
name
embeddings-setup
description
Bring Open Second Brain semantic search online — embedding API key, sqlite-vec extension, first reindex, optional periodic refresh. INVOKE when the user mentions "embeddings", "semantic search", "vector index", or after `o2b search check` reports `vec_extension: unavailable`, `embedding_key: missing`, or any line in its `recommendations:` block. SKIP when the user is asking conceptual questions about embeddings or comparing providers without intent to set up — answer those directly.

Embeddings setup

Open Second Brain ships with two search paths: keyword-only (always on, no credentials) and semantic via embedded vectors (opt-in). This SKILL walks the activation flow for the semantic path. The flow is proactive — when the user mentions semantic search or o2b search check surfaces missing pieces, take the user through this list rather than waiting for explicit instruction.

Step 1 — Always start with o2b search check

bash
o2b search check

The output names every missing piece and ends with a recommendations: block listing the exact commands to fix each. Read both before suggesting next steps. The recommendations field is also present in the --json output for headless callers.

Branch on what the report shows:

  • embedding_key: MISSING → go to step 2. The key_sources_checked: line names where the key was looked for, and key_present_under: names any other registered provider profile whose env key is set. If the key sits under another profile, the fix may be pointing embedding_provider at that profile rather than adding a key.
  • A recommendation saying the embedding model has no known price, or that the price pair names another model → see "Declare the model's price" in step 2.
  • vec_extension: unavailable on macOS → go to step 3.
  • vec_extension: unavailable on Linux → go to step 4.
  • Everything OK but semantic_enabled: false (no embeddings yet) → go to step 5.

Step 2 — Provider and API key

Ask the user which provider they want. The default is text-embedding-3-small from OpenAI (about $0.02 per 1M tokens, which covers tens of thousands of vault pages). Any OpenAI-compatible endpoint works — Groq, Together, a local LM Studio server, etc.

Required env vars (write to ~/.hermes/.env or the configured env file, never to a tracked file):

bash
OPEN_SECOND_BRAIN_EMBEDDING_PROVIDER=openai-compat
OPEN_SECOND_BRAIN_EMBEDDING_MODEL=text-embedding-3-small
OPEN_SECOND_BRAIN_EMBEDDING_KEY=<placeholder; user pastes the key>
# Optional — only when not using OpenAI:
# OPEN_SECOND_BRAIN_EMBEDDING_BASE_URL=https://api.together.xyz/v1

The base URL must be https://. Plain http:// works only for a loopback host (localhost, 127.0.0.1, ::1). When the user's embedding server runs on ANOTHER machine without TLS - LM Studio on the Windows host reached from WSL, Ollama on a LAN box, a tailnet 100.x address - the endpoint is refused until the operator opts in for it explicitly:

bash
OPEN_SECOND_BRAIN_EMBEDDING_BASE_URL=http://100.64.0.5:1234/v1
OPEN_SECOND_BRAIN_EMBEDDING_ALLOW_INSECURE_HTTP=true

(or embedding_allow_insecure_http: true in the o2b config; the reranker has its own search_rerank_allow_insecure_http). Tell the user what it means before setting it: chunk text and the API key cross the network unencrypted, so it belongs only on a network they trust. The opt-out never applies to a base URL that comes from a provider profile registered with o2b search provider add; set the URL in config or env instead. Every run that uses it prints one stderr warning per endpoint.

Never invent or echo the key. Write a placeholder, then ask the user to paste their key in place of it. Recheck with o2b search check after the user confirms.

Declare the model's price

Spend estimates name their price source: builtin (the built-in table; the local provider is free), operator (declared) or unknown. A model the table does not list (most local servers, newer or niche models) is unknown, and its estimates print price unknown. When the user has set a positive embedding_cost_gate_usd, an unknown price refuses embedding runs with EMBEDDING_COST_UNPRICED until the price is declared, and it also refuses the query embeds of searches that do not run at local reach (MCP tools and hooks): a hybrid search falls back to keyword-only and names semantic-cost-unpriced in its trail. Ask the user for the provider's rate (USD per million tokens) and write the pair beside the other embedding settings:

bash
OPEN_SECOND_BRAIN_EMBEDDING_PRICE_MODEL=nomic-embed-text:latest
OPEN_SECOND_BRAIN_EMBEDDING_PRICE_USD_PER_MTOK=0

(or embedding_price_model and embedding_price_usd_per_mtok in the o2b config). Set both or neither. 0 declares the model free; only use it when the user confirms the endpoint costs nothing, because a loopback URL can still be a paid proxy. The pair names one model, so after a model switch o2b search check flags the stale declaration until it is updated. Declaring a price never triggers a reindex.

Declare the model's input window (uncurated models)

A search query longer than the model's input window is cut to it before the embed and the trail names semantic-query-truncated. The window comes from the curated model table; for a model it does not list the window is unknown and nothing is cut, so a serving stack that truncates silently or refuses long inputs does so unannounced. When the user knows the model's window (in its own tokens), declare it:

bash
OPEN_SECOND_BRAIN_EMBEDDING_INPUT_WINDOW_TOKENS=512

(or embedding_input_window_tokens in the o2b config). It is refused for the local provider, which has no window.

Extra request fields (OpenAI-compatible endpoints)

When the user's serving stack needs a request field the provider does not send, put a JSON object in embedding_extra_body (env OPEN_SECOND_BRAIN_EMBEDDING_EXTRA_BODY). Only openai-compat sends it, and model, input and encoding_format are refused by name. A dimensions field needs embedding_dimension set to the same width:

bash
OPEN_SECOND_BRAIN_EMBEDDING_EXTRA_BODY='{"dimensions": 512}'
OPEN_SECOND_BRAIN_EMBEDDING_DIM=512

The extra body is not part of the embedding identity, so changing it never triggers a reindex; a changed embedding_dimension does. Some fields shape the vectors the provider returns (a provider task, input_type, normalize or truncate switch): after changing such a field, rebuild the vectors with o2b search reindex --embeddings so new query vectors stay in the space of the stored passages (o2b search index --force keeps the stored vector of every unchanged chunk).

Show full SKILL.md (493 more words)Show less

Step 3 — macOS: install Homebrew SQLite

Apple ships /usr/lib/libsqlite3.dylib with SQLITE_OMIT_LOAD_EXTENSION, so the optional sqlite-vec extension cannot load against the system SQLite. Homebrew's sqlite formula is built with extension loading enabled.

bash
brew install sqlite

The o2b wrapper auto-detects Homebrew SQLite on Darwin and exports DYLD_LIBRARY_PATH for bun:sqlite on the next invocation. No manual patching of the wrapper script is needed (the v0.10.5 shim scripts/_macos-sqlite.sh handles it). Verify with:

bash
o2b search check

Expect vec_extension: loaded.

Step 4 — Linux: reconfirm the optional dependency

sqlite-vec ships as an optional dependency. If the install path missed it, force a rebuild:

bash
bun pm ls | grep sqlite-vec
bun install --force

If the package is present but still does not load, the Linux build of libsqlite3 may have been compiled without LOAD_EXTENSION. Capture sqlite3 :memory: "PRAGMA compile_options;" and surface it to the user — they will know whether their distro's SQLite is unusual.

Step 5 — Compute the first vectors

bash
o2b search reindex --embeddings

This walks the vault, chunks every Markdown file, and computes embeddings. Cost scales linearly with vault size — a 150-file vault typically lands in a few seconds.

When the index already exists and only some chunks lack vectors, price the work first with the dry run, then apply it:

bash
o2b search vector-backfill
o2b search vector-backfill --apply

--path <prefix> (repeatable) limits the census, the estimate and the spend to one part of the vault. For example, o2b search vector-backfill --path Brain/preferences/ --path Brain/retired/ --apply embeds just the belief notes, which is what the semantic query mode of brain_context_pack reads. Later edits re-embed only the chunks that changed; unchanged chunks keep their vectors.

Verify semantic search works:

bash
o2b search "preferences for code review"

The output rows carry (semantic) in the source column when the match came from the vector index.

Step 6 — Offer periodic reindex

The vault keeps changing — other agents add notes, the user edits existing ones. Without periodic refresh the embeddings drift behind the keyword index.

bash
o2b search reindex --cron-template

Prints a watchdog script (a heredoc that, when the operator runs the printed block, lands at ~/.local/bin/osb-reindex.sh), a native crontab line, and a hermes cron create command. Pure stdout — the verb writes nothing on its own; pick the path that matches the host and paste the relevant section into your shell / crontab. Recommended cadence:

  • 30 minutes when the vault sees active multi-agent work.
  • 6 hours when changes are sporadic.

Override with --interval 6h (or any <N>m|h|d value below 60m / 24h / unlimited days).

Multi-agent note

Only the agent that runs reindex needs OPEN_SECOND_BRAIN_EMBEDDING_KEY. Read-only consumers (sibling Claude Code / Codex / OpenClaw sessions on the same vault) query the already-computed vectors and never see the key. When designating the reindex owner, the canonical choice is Hermes (it carries the cron infrastructure); the others stay credential-free.

What this SKILL does NOT do

  • Does not invoke brew install on the user's behalf. The SKILL prints the command; the user runs it.
  • Does not paste an API key into the env file silently. Always use a placeholder, then ask the user to substitute.
  • Does not commit env files. They live outside the tracked tree.
  • Does not launch a long-running watcher inside OSB — the --cron-template recipe is the endorsed automation path.

© itechmeat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/embeddings-setup of itechmeat/open-second-brain.

Open the folder on GitHubat commit 0f0c9a7

Compare with similar skills

Open Second Brain Embeddings Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Open Second Brain Embeddings Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Open Second Brain Embeddings Setup this skillitechmeat/open-second-brain430—~2.6kAutomated safety check: WarnMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
Langchain RAGlangchain-ai/langchain-skills1.3k1 repos~3.9kAutomated safety check: PassMIT
Cortexdbliliang-cn/cortexdb273—~18kAutomated safety check: WarnMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub starsUsed in 1 repo~3.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Cortexdb

    liliang-cn/cortexdb

    Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.

    273 GitHub stars~18k tokensUpdated today
    Knowledge ManagementAuto-check: warnings
  • Spring AI Integration

    rrezartprebreza/spring-boot-skills

    A skill your agent uses when integrating LLMs, chat clients, embeddings, RAG pipelines, or AI agents into Spring Boot.

    296 GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed

More from itechmeat/open-second-brain

  • Decision Model Setup

    itechmeat/open-second-brain

    Sets up, evaluates or turns off Open Second Brain's optional decision-model feature, starting from the o2b decision-model check command and a provider route you choose.

    430 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Codegraph Detection Partner

    itechmeat/open-second-brain

    Teaches the agent to detect whether the codegraph symbol graph is available and use it for callers, callees and impact questions, without installing or writing anything for it.

    430 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Brain Schema Author

    itechmeat/open-second-brain

    Adds, renames, reviews and explains Brain schema tokens, aliases, prefixes and link types in Open Second Brain vaults, previewing each change before applying it.

    430 GitHub stars~513 tokensUpdated today
    Auto-check passed
  • Open Second Brain

    itechmeat/open-second-brain

    Lets an agent read, write and maintain its own memory in a Brain folder of an Obsidian-compatible vault, while keeping the user's own notes read-only.

    430 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Open Second Brain Embeddings Setup

What does Open Second Brain Embeddings Setup do?

Walks through turning on semantic search in Open Second Brain: embedding key, sqlite-vec extension, first reindex and an optional periodic refresh, starting from o2b search check. Open Second Brain always has keyword search; semantic search over embedded vectors is opt-in. The skill acts proactively when you mention embeddings or semantic search, or when `o2b search check` reports missing pieces.

When should I use Open Second Brain Embeddings Setup?

Open Second Brain Embeddings Setup fits situations like: enabling semantic search in an Open Second Brain vault; fixing an embedding key or sqlite-vec problem reported by o2b search check; choosing an embedding provider or a local OpenAI-compatible endpoint.

How do I install Open Second Brain Embeddings Setup in Claude Code?

Run `npx skills add itechmeat/open-second-brain --skill embeddings-setup -a claude-code`. Or copy the skill folder (skills/embeddings-setup in itechmeat/open-second-brain) into .claude/skills/embeddings-setup in your project. Claude Code loads it when a task matches its description.

How do I install Open Second Brain Embeddings Setup in Codex?

Run `npx skills add itechmeat/open-second-brain --skill embeddings-setup -a codex`. Or copy the skill folder (skills/embeddings-setup in itechmeat/open-second-brain) into .agents/skills/embeddings-setup in your project. Codex loads it when a task matches its description.

Can I use Open Second Brain Embeddings Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add itechmeat/open-second-brain --skill embeddings-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/embeddings-setup, .gemini/skills/embeddings-setup, .github/skills/embeddings-setup and .opencode/skills/embeddings-setup in your project.

What does Open Second Brain Embeddings Setup need to run?

Going by SKILL.md and its folder, Open Second Brain Embeddings Setup needs the command-line tools its instructions call (brew, bun and sqlite3) and credentials named OPEN_SECOND_BRAIN_EMBEDDING_KEY. Our summary lists: The o2b command-line tool from Open Second Brain; An embedding API key or an OpenAI-compatible embedding endpoint.

Does Open Second Brain Embeddings Setup access the network?

SKILL.md names 1 domain. In commands or code: api.together.xyz; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Open Second Brain Embeddings Setup safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): links to a raw public ip address. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Open Second Brain Embeddings Setup use?

Open Second Brain Embeddings Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Open Second Brain Embeddings Setup use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Open Second Brain Embeddings Setup?

Skills that share tags, products or a category with Open Second Brain Embeddings Setup: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars), RAG Architect (Jeffallan/claude-skills, 12k stars) and Langchain RAG (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Open Second Brain Embeddings Setup?

itechmeat (a GitHub user) maintains it in itechmeat/open-second-brain, which has 430 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.

Source: itechmeat/open-second-brain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.