Chroma Vector Database
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Cloudflare Vectorize vector database for semantic search and RAG.
$ npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install secondsky/claude-skills cloudflare-vectorize --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cloudflare-vectorize/skills/cloudflare-vectorize .claude/skills/cloudflare-vectorize && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cloudflare-vectorize" agent skill from https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorize into .claude/skills/cloudflare-vectorize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cloudflare-vectorize", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorizeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install secondsky/claude-skills cloudflare-vectorize --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/cloudflare-vectorize/skills/cloudflare-vectorize .agents/skills/cloudflare-vectorize && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cloudflare-vectorize" agent skill from https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorize into .agents/skills/cloudflare-vectorize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cloudflare-vectorize", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install secondsky/claude-skills cloudflare-vectorize --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/cloudflare-vectorize/skills/cloudflare-vectorize .cursor/skills/cloudflare-vectorize && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cloudflare-vectorize" agent skill from https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorize into .cursor/skills/cloudflare-vectorize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cloudflare-vectorize", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/secondsky/claude-skills.git --path plugins/cloudflare-vectorize/skills/cloudflare-vectorize--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install secondsky/claude-skills cloudflare-vectorize --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/cloudflare-vectorize/skills/cloudflare-vectorize .gemini/skills/cloudflare-vectorize && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cloudflare-vectorize" agent skill from https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorize into .gemini/skills/cloudflare-vectorize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cloudflare-vectorize", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install secondsky/claude-skills cloudflare-vectorizeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/cloudflare-vectorize/skills/cloudflare-vectorize .github/skills/cloudflare-vectorize && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cloudflare-vectorize" agent skill from https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorize into .github/skills/cloudflare-vectorize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cloudflare-vectorize", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install secondsky/claude-skills cloudflare-vectorize --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/cloudflare-vectorize/skills/cloudflare-vectorize .opencode/skills/cloudflare-vectorize && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cloudflare-vectorize" agent skill from https://github.com/secondsky/claude-skills/tree/main/plugins/cloudflare-vectorize/skills/cloudflare-vectorize into .opencode/skills/cloudflare-vectorize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cloudflare-vectorize", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cloudflare-vectorizeCloudflare Vectorize vector database for semantic search and RAG.
Cloudflare Vectorize is an agent skill from secondsky/claude-skills. Cloudflare Vectorize vector database for semantic search and RAG. Use for vector indexes, embeddings, similarity search, or encountering dimension mismatches, filter errors.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `references/embedding-models.md`, `references/index-operations.md` and `references/integration-openai-embeddings.md`).
It sits in AI & LLM Engineering, covering Vector databases, Embeddings and Retrieval-augmented generation. It works with Cloudflare, Cloudflare Workers and Workers AI. The repository describes itself as: Production-ready skills for Claude Code CLI - Cloudflare, React, Tailwind v4, and AI integrations. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8837836. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (TypeScript), which the agent can run.
Shell commands in SKILL.md call:
bunxnpmFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
developers.cloudflare.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cloudflare Vectorize loads about 3.3k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 781 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from secondsky/claude-skills at commit 8837836, republished under its MIT licence (© secondsky). 781 words, ~3,262 tokens.
.claude/skills/cloudflare-vectorize/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Complete implementation guide for Cloudflare Vectorize - a globally distributed vector database for building semantic search, RAG (Retrieval Augmented Generation), and AI-powered applications with Cloudflare Workers.
Status: Production Ready ✅ Last Updated: 2025-11-21 Dependencies: cloudflare-worker-base (for Worker setup), cloudflare-workers-ai (for embeddings) Latest Versions: wrangler@4.81.0, @cloudflare/workers-types@4.20260408.0 Token Savings: ~65% Errors Prevented: 8 Dev Time Saved: ~3 hours
# 1. Create the index with FIXED dimensions and metric
bunx wrangler vectorize create my-index \
--dimensions=768 \
--metric=cosine
# 2. Create metadata indexes IMMEDIATELY (before inserting vectors!)
bunx wrangler vectorize create-metadata-index my-index \
--property-name=category \
--type=string
bunx wrangler vectorize create-metadata-index my-index \
--property-name=timestamp \
--type=numberWhy: Metadata indexes MUST exist before vectors are inserted. Vectors added before a metadata index was created won't be filterable on that property.
# Dimensions MUST match your embedding model output:
# - Workers AI @cf/baai/bge-base-en-v1.5: 768 dimensions
# - OpenAI text-embedding-3-small: 1536 dimensions
# - OpenAI text-embedding-3-large: 3072 dimensions
# Metrics determine similarity calculation:
# - cosine: Best for normalized embeddings (most common)
# - euclidean: Absolute distance between vectors
# - dot-product: For non-normalized vectorswrangler.jsonc:
{
"name": "my-vectorize-worker",
"main": "src/index.ts",
"compatibility_date": "2025-10-21",
"vectorize": [
{
"binding": "VECTORIZE_INDEX",
"index_name": "my-index"
}
],
"ai": {
"binding": "AI"
}
}export interface Env {
VECTORIZE_INDEX: VectorizeIndex;
AI: Ai;
}
interface VectorizeVector {
id: string;
values: number[] | Float32Array | Float64Array;
namespace?: string;
metadata?: Record<string, string | number | boolean | string[]>;
}
interface VectorizeMatches {
matches: Array<{
id: string;
score: number;
values?: number[];
metadata?: Record<string, any>;
namespace?: string;
}>;
count: number;
}| Operation | Method | Key Point |
|---|---|---|
| Insert | insert([...]) | Keeps first if ID exists |
| Upsert | upsert([...]) | Overwrites if ID exists (use for updates) |
| Query | query(vector, { topK, filter }) | Returns similar vectors |
| Delete | deleteByIds([...]) | Remove by ID array |
| Get | getByIds([...]) | Retrieve specific vectors |
| Operator | Example | Description |
|---|---|---|
$eq | { category: "docs" } | Equality (implicit) |
$ne | { status: { $ne: "archived" } } | Not equal |
$in | { category: { $in: ["a", "b"] } } | In array |
$nin | { category: { $nin: ["x"] } } | Not in array |
$gte/$lt | { timestamp: { $gte: 123 } } | Range queries |
📄 Full operations guide: Load references/vector-operations.md for complete insert/upsert/query/delete examples with code.
| Model | Provider | Dimensions | Best For |
|---|---|---|---|
@cf/baai/bge-base-en-v1.5 | Workers AI | 768 | Free, general purpose |
text-embedding-3-small | OpenAI | 1536 | Balance quality/cost |
text-embedding-3-large | OpenAI | 3072 | Highest quality |
📄 Integration guides:
references/integration-workers-ai-bge-base.md for Workers AI setupreferences/integration-openai-embeddings.md for OpenAI integration| Limit | Value |
|---|---|
| Max metadata indexes | 10 per index |
| Max metadata size | 10 KiB per vector |
| String index | First 64 bytes (UTF-8) |
| Filter size | Max 2048 bytes |
Keys cannot: be empty, contain . (reserved for nesting), contain ", or start with $.
📄 Complete metadata guide: Load references/metadata-guide.md for cardinality best practices, nested metadata, and advanced filtering patterns.
export default {
async fetch(request: Request, env: Env): Promise<Response> {
const { question } = await request.json();
// 1. Generate embedding for user question
const questionEmbedding = await env.AI.run('@cf/baai/bge-base-en-v1.5', {
text: question
});
// 2. Search vector database for similar content
const results = await env.VECTORIZE_INDEX.query(
questionEmbedding.data[0],
{
topK: 3,
returnMetadata: 'all',
filter: { type: "documentation" }
}
);
// 3. Build context from retrieved documents
const context = results.matches
.map(m => m.metadata.content)
.join('\n\n---\n\n');
// 4. Generate answer with LLM using context
const answer = await env.AI.run('@cf/meta/llama-3-8b-instruct', {
messages: [
{
role: "system",
content: `Answer based on this context:\n\n${context}`
},
{
role: "user",
content: question
}
]
});
return Response.json({
answer: answer.response,
sources: results.matches.map(m => m.metadata.title)
});
}
};Recommended chunk sizes: 300-500 characters for semantic coherence.
Key metadata for chunks:
doc_id: Parent document IDchunk_index: Position in documentcontent: Text for retrieval display📄 Full chunking implementation: See templates/document-ingestion.ts for complete chunking pipeline.
Problem: Filtering doesn't work on existing vectors
Solution: Delete and re-insert vectors OR create metadata indexes BEFORE insertingProblem: "Vector dimensions do not match index configuration"
Solution: Ensure embedding model output matches index dimensions:
- Workers AI bge-base: 768
- OpenAI small: 1536
- OpenAI large: 3072Problem: "Invalid metadata key"
Solution: Keys cannot:
- Be empty
- Contain . (dot)
- Contain " (quote)
- Start with $ (dollar sign)Problem: "Filter exceeds 2048 bytes"
Solution: Simplify filter or split into multiple queriesProblem: Slow queries or reduced accuracy
Solution: Use lower cardinality fields for range queries, or use seconds instead of milliseconds for timestampsProblem: Updates not reflecting in index
Solution: Use upsert() to overwrite existing vectors, not insert()Problem: "VECTORIZE_INDEX is not defined"
Solution: Add [[vectorize]] binding to wrangler.jsoncProblem: Unclear when to use namespace vs metadata filtering
Solution:
- Namespace: Partition key, applied BEFORE metadata filters
- Metadata: Flexible key-value filtering within namespaceEssential commands:
# Create index (dimensions/metric are PERMANENT)
bunx wrangler vectorize create <name> --dimensions=768 --metric=cosine
# Create metadata index (MUST be before inserting vectors!)
bunx wrangler vectorize create-metadata-index <name> --property-name=category --type=string
# Get index info
bunx wrangler vectorize info <name>📄 Full CLI reference: Load references/wrangler-commands.md for all vectorize commands.
returnValues: true when needed (saves bandwidth)✅ Use Vectorize when:
❌ Don't use Vectorize for:
| Reference File | Load When... |
|---|---|
references/vector-operations.md | Need full insert/upsert/query/delete code examples |
references/metadata-guide.md | Setting up metadata indexes, filtering best practices |
references/wrangler-commands.md | Using Vectorize CLI commands |
references/integration-workers-ai-bge-base.md | Integrating Workers AI embeddings |
references/integration-openai-embeddings.md | Integrating OpenAI embeddings |
references/embedding-models.md | Comparing embedding model options |
references/index-operations.md | Index lifecycle management |
| Template | Purpose |
|---|---|
templates/basic-search.ts | Simple vector search |
templates/rag-chat.ts | Complete RAG chatbot |
templates/document-ingestion.ts | Document chunking pipeline |
templates/metadata-filtering.ts | Advanced filtering |
When installing vector database packages, follow supply chain security best practices:
npm config set ignore-scripts true (or Bun: disabled by default)socket package score npm <pkg> or use socket npm install <pkg> to check packagesLoad the dependency-upgrade skill for full security configuration including Socket CLI integration, cooldown setup, lockfile validation, and CI enforcement.
Version: 1.0.0 Status: Production Ready ✅ Token Savings: ~65% Errors Prevented: 8 major categories Dev Time Saved: ~2.5 hours per implementation
© secondsky, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (references) in plugins/cloudflare-vectorize/skills/cloudflare-vectorize of secondsky/claude-skills.
Open the folder on GitHubat commit 8837836
Cloudflare Vectorize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cloudflare Vectorize this skillsecondsky/claude-skills | 227 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| AI SDK Developmenttrypostit/trypost | 676 | 2 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Retail Product Search Agentgoogle/adk-recipes | 10k | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | 1 repos | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| RAG ArchitectJeffallan/claude-skills | 12k | 1 repos | ~2k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
trypostit/trypost
TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.
google/adk-recipes
Builds a retail product search agent on Google Cloud, from catalog ingestion into BigQuery and Vector Search to ADK scaffolding, evaluation and Cloud Run deployment.
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
secondsky/claude-skills
TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.
secondsky/claude-skills
AutoAnimate (@formkit/auto-animate) zero-config animations for React.
secondsky/claude-skills
MUI Base UI unstyled React components with Floating UI. An agent skill from secondsky/claude-skills.
secondsky/claude-skills
This skill should be used when the user asks to "upload images to Cloudflare", "implement direct creator upload", "configure image transformations", "optimize WebP/AVIF", "create image variants"…
secondsky/claude-skills
Deploy Next.js to Cloudflare Workers via the OpenNext adapter (@opennextjs/cloudflare).
secondsky/claude-skills
Cloudflare Sandboxes SDK for secure code execution in Linux containers at edge.
Works with
Categories
Cloudflare Vectorize vector database for semantic search and RAG. Cloudflare Vectorize is an agent skill from secondsky/claude-skills. Cloudflare Vectorize vector database for semantic search and RAG.
Cloudflare Vectorize fits situations like: similarity search; encountering dimension mismatches.
Run `npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a claude-code`. Or copy the skill folder (plugins/cloudflare-vectorize/skills/cloudflare-vectorize in secondsky/claude-skills) into .claude/skills/cloudflare-vectorize in your project. Claude Code loads it when a task matches its description.
Run `npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a codex`. Or copy the skill folder (plugins/cloudflare-vectorize/skills/cloudflare-vectorize in secondsky/claude-skills) into .agents/skills/cloudflare-vectorize in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add secondsky/claude-skills --skill cloudflare-vectorize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cloudflare-vectorize, .gemini/skills/cloudflare-vectorize, .github/skills/cloudflare-vectorize and .opencode/skills/cloudflare-vectorize in your project.
Going by SKILL.md and its folder, Cloudflare Vectorize needs TypeScript for the scripts in its folder and the command-line tools its instructions call (bunx and npm). Our summary lists: Node.js.
SKILL.md names 1 domain. As links in the text: developers.cloudflare.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cloudflare Vectorize is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Cloudflare Vectorize: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), AI SDK Development (trypostit/trypost, 676 stars), Retail Product Search Agent (google/adk-recipes, 10k stars) and Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
secondsky (a GitHub user) maintains it in secondsky/claude-skills, which has 227 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 28, 2026.
Source: secondsky/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.