Agent skill

Vector Database Ops

by majiayu000 in majiayu000/claude-skill-registry

Deploy, manage, and optimize vector databases for AI applications.

MITAuto-check passedDatabases

Install Vector Database Ops

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill vector-database-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry vector-database-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-ml/vector-database-ops .claude/skills/vector-database-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vector-database-ops
GitHub stars
666
Used in
3 other repos
Token cost
~2.1k tokens
SKILL.md length
263 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT

At a glance

Deploy, manage, and optimize vector databases for AI applications.

  • Tasks that involve Vector databases
  • SKILL.md covers When to Use This Skill, Vector Database Comparison, Qdrant — Production Deployment and Qdrant Collection Management, plus 8 more sections
  • Calls docker, curl and pg_dump; needs POSTGRES_PASSWORD and WEAVIATE_API_KEY

What it does

Vector Database Ops is an agent skill from majiayu000/claude-skill-registry. Deploy, manage, and optimize vector databases for AI applications. Covers Qdrant, Weaviate, pgvector, and Pinecone — collection management, indexing strategies, backup, and performance tuning for production RAG and semantic search workloads.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Databases, covering Vector databases. It works with Qdrant, pgvector, Weaviate and Pinecone. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve Vector databases

Example prompts

  • “/vector-database-ops”

Requirements

  • Python 3
  • Docker
  • A credential in WEAVIATE_API_KEY
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • curl
    • pg_dump

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • POSTGRES_PASSWORD
    • WEAVIATE_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vector Database Ops loads about 2.1k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 263 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 263 words, ~2,127 tokens.

Download SKILL.mdSave it as .claude/skills/vector-database-ops/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vector-database-ops
description
Deploy, manage, and optimize vector databases for AI applications. Covers Qdrant, Weaviate, pgvector, and Pinecone — collection management, indexing strategies, backup, and performance tuning for production RAG and semantic search workloads.
license
MIT
metadata.author
devops-skills
metadata.version
1.0

Vector Database Operations

Run production vector databases for AI-powered search, RAG, and recommendation systems.

When to Use This Skill

Use this skill when:

  • Setting up a vector database for a RAG or semantic search application
  • Choosing between Qdrant, Weaviate, pgvector, or Pinecone
  • Managing collections, indexes, and data migrations
  • Optimizing query performance and indexing for production loads
  • Implementing multi-tenant vector search with namespace isolation

Vector Database Comparison

DatabaseBest ForHostingFilteringScale
QdrantHigh-performance, rich filtering, self-hostedSelf / CloudExcellentVery High
WeaviateSchema-first, hybrid search, multi-modalSelf / CloudGoodHigh
pgvectorAlready on Postgres, simple use casesSelfGoodMedium
PineconeZero-ops managed, serverlessManaged onlyGoodVery High
ChromaLocal dev, prototypingSelf onlyBasicLow-Medium

Qdrant — Production Deployment

bash
# Docker (single node)
docker run -d \
  --name qdrant \
  -p 6333:6333 \
  -p 6334:6334 \
  -v $(pwd)/qdrant-data:/qdrant/storage \
  qdrant/qdrant:latest

# With custom config
docker run -d \
  --name qdrant \
  -p 6333:6333 \
  -v $(pwd)/qdrant-data:/qdrant/storage \
  -v $(pwd)/qdrant-config.yaml:/qdrant/config/production.yaml \
  qdrant/qdrant:latest
yaml
# qdrant-config.yaml
storage:
  storage_path: /qdrant/storage
  on_disk_payload: true          # store payload on disk (saves RAM)

service:
  max_request_size_mb: 32

hnsw_index:
  m: 16                          # graph connections per node
  ef_construct: 100              # accuracy vs build time trade-off
  full_scan_threshold: 10000     # switch to brute force below this

quantization:
  scalar:
    type: int8
    quantile: 0.99
    always_ram: true             # keep quantized index in RAM

telemetry_disabled: true

Qdrant Collection Management

python
from qdrant_client import QdrantClient
from qdrant_client.models import (
    Distance, VectorParams, HnswConfigDiff,
    ScalarQuantizationConfig, ScalarType, QuantizationConfig
)

client = QdrantClient("http://localhost:6333")

# Create optimized collection
client.create_collection(
    collection_name="documents",
    vectors_config=VectorParams(
        size=1536,                         # OpenAI ada-002 / text-embedding-3-small
        distance=Distance.COSINE,
        on_disk=True,                      # save RAM — vectors stored on disk
    ),
    hnsw_config=HnswConfigDiff(
        m=32,                              # higher = better recall, more RAM
        ef_construct=200,
        on_disk=False,                     # keep HNSW graph in RAM for speed
    ),
    quantization_config=QuantizationConfig(
        scalar=ScalarQuantizationConfig(
            type=ScalarType.INT8,
            quantile=0.99,
            always_ram=True,
        )
    ),
)

# Create payload index for fast filtering
client.create_payload_index(
    collection_name="documents",
    field_name="tenant_id",
    field_schema="keyword",
)
client.create_payload_index(
    collection_name="documents",
    field_name="created_at",
    field_schema="datetime",
)

# Collection info
info = client.get_collection("documents")
print(f"Vectors: {info.vectors_count}, Status: {info.status}")
python
from qdrant_client.models import Filter, FieldCondition, MatchValue, Range

# Tenant-isolated search (multi-tenant RAG)
results = client.query_points(
    collection_name="documents",
    query=query_embedding,
    query_filter=Filter(
        must=[
            FieldCondition(key="tenant_id", match=MatchValue(value="acme-corp")),
            FieldCondition(key="doc_type", match=MatchValue(value="contract")),
        ],
        should=[
            FieldCondition(key="created_at", range=Range(gte="2024-01-01")),
        ],
    ),
    limit=10,
    with_payload=True,
)

pgvector — PostgreSQL Extension

sql
-- Enable extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Create table with vector column
CREATE TABLE documents (
    id          UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    content     TEXT NOT NULL,
    embedding   VECTOR(1536),
    metadata    JSONB DEFAULT '{}',
    tenant_id   TEXT NOT NULL,
    created_at  TIMESTAMPTZ DEFAULT NOW()
);

-- Create HNSW index (faster queries, more memory)
CREATE INDEX ON documents
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);

-- Create IVFFlat index (less memory, slower build)
-- CREATE INDEX ON documents
-- USING ivfflat (embedding vector_cosine_ops)
-- WITH (lists = 100);

-- Semantic search with metadata filtering
SELECT id, content, metadata,
       1 - (embedding <=> $1::vector) AS similarity
FROM documents
WHERE tenant_id = 'acme-corp'
  AND metadata->>'doc_type' = 'contract'
ORDER BY embedding <=> $1::vector
LIMIT 10;
bash
# Deploy pgvector via Docker
docker run -d \
  --name pgvector \
  -e POSTGRES_PASSWORD=secret \
  -e POSTGRES_DB=vectordb \
  -p 5432:5432 \
  -v pgvector-data:/var/lib/postgresql/data \
  pgvector/pgvector:pg16

Weaviate Deployment

yaml
# docker-compose for Weaviate
services:
  weaviate:
    image: semitechnologies/weaviate:latest
    ports:
      - "8080:8080"
      - "50051:50051"
    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: "false"
      AUTHENTICATION_APIKEY_ENABLED: "true"
      AUTHENTICATION_APIKEY_ALLOWED_KEYS: "${WEAVIATE_API_KEY}"
      AUTHENTICATION_APIKEY_USERS: "admin"
      PERSISTENCE_DATA_PATH: /var/lib/weaviate
      ENABLE_MODULES: text2vec-openai,generative-openai
      OPENAI_APIKEY: "${OPENAI_API_KEY}"
      CLUSTER_HOSTNAME: node1
    volumes:
      - weaviate-data:/var/lib/weaviate
    restart: unless-stopped

volumes:
  weaviate-data:

Backup and Restore

bash
# Qdrant — snapshot backup
curl -X POST "http://localhost:6333/collections/documents/snapshots"
# Download snapshot
curl -O "http://localhost:6333/collections/documents/snapshots/documents-snapshot.snapshot"
# Restore
curl -X POST "http://localhost:6333/collections/documents/snapshots/recover" \
  -H "Content-Type: application/json" \
  -d '{"location": "/qdrant/snapshots/documents-snapshot.snapshot"}'

# pgvector — standard pg_dump
pg_dump -h localhost -U postgres -d vectordb \
  --table=documents --format=custom > documents-backup.dump

# Restore
pg_restore -h localhost -U postgres -d vectordb documents-backup.dump

Performance Tuning

python
# Qdrant — optimize collection after bulk load
client.update_collection(
    collection_name="documents",
    optimizer_config={"indexing_threshold": 0},  # force indexing now
)

# Wait for optimization to complete
import time
while True:
    info = client.get_collection("documents")
    if info.status.value == "green":
        break
    time.sleep(5)
    print(f"Optimizing... segments: {info.segments_count}")

Common Issues

IssueCauseFix
Slow queriesNo HNSW index built yetWait for indexing; check status == green
High RAM usageVectors in memoryEnable on_disk=True for vectors
Poor recallLow ef search paramIncrease ef in search request (at query time)
pgvector slowUsing IVFFlat without vacuumRun VACUUM ANALYZE documents
Weaviate OOMToo many objectsEnable async indexing; increase heap

Best Practices

  • Use cosine distance for normalized embeddings; dot product for unnormalized.
  • Always create payload indexes on filter fields (tenant_id, doc_type).
  • For datasets >10M vectors, use on_disk vectors + always_ram quantization.
  • Benchmark with your actual query patterns before choosing IVFFlat vs HNSW.
  • Snapshot before any bulk delete or migration operation.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-ml/vector-database-ops of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 000116a

Used in 3 other repositories

We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vector Database Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vector Database Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vector Database Ops this skillmajiayu000/claude-skill-registry6663 repos~2.1kAutomated safety check: PassMIT
Vector Database Engineeraiskillstore/marketplace4307 repos~563Automated safety check: PassNone
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
Agentsop Multi Tenant RAGagentsope/SkillAlchemy457—~9.8kAutomated safety check: PassMIT
Vector DBericrisco/rsc-harness156—~2.8kAutomated safety check: PassMIT

Similar skills

  • Vector Database Engineer

    aiskillstore/marketplace

    Expert in vector databases, embedding strategies, and semantic search implementation.

    430 GitHub starsUsed in 7 repos~563 tokens
    DatabasesAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agentsop Multi Tenant RAG

    agentsope/SkillAlchemy

    Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.

    457 GitHub stars~9.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vector DB

    ericrisco/rsc-harness

    A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

    156 GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed

More from majiayu000/claude-skill-registry

All 971 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed
  • Open Notebook

    majiayu000/claude-skill-registry

    Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

    666 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed

Categories

Questions about Vector Database Ops

What does Vector Database Ops do?

Deploy, manage, and optimize vector databases for AI applications. Vector Database Ops is an agent skill from majiayu000/claude-skill-registry. Deploy, manage, and optimize vector databases for AI applications.

When should I use Vector Database Ops?

Vector Database Ops fits situations like: tasks that involve Vector databases.

How do I install Vector Database Ops in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill vector-database-ops -a claude-code`. Or copy the skill folder (skills/ai-ml/vector-database-ops in majiayu000/claude-skill-registry) into .claude/skills/vector-database-ops in your project. Claude Code loads it when a task matches its description.

How do I install Vector Database Ops in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill vector-database-ops -a codex`. Or copy the skill folder (skills/ai-ml/vector-database-ops in majiayu000/claude-skill-registry) into .agents/skills/vector-database-ops in your project. Codex loads it when a task matches its description.

Can I use Vector Database Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill vector-database-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vector-database-ops, .gemini/skills/vector-database-ops, .github/skills/vector-database-ops and .opencode/skills/vector-database-ops in your project.

What does Vector Database Ops need to run?

Going by SKILL.md and its folder, Vector Database Ops needs the command-line tools its instructions call (docker, curl and pg_dump) and credentials named POSTGRES_PASSWORD, WEAVIATE_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; Docker; A credential in WEAVIATE_API_KEY; A credential in OPENAI_API_KEY.

Does Vector Database Ops access the network?

SKILL.md contains no URLs. Its commands use docker and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vector Database Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vector Database Ops use?

Vector Database Ops is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vector Database Ops use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vector Database Ops?

Skills that share tags, products or a category with Vector Database Ops: Vector Database Engineer (aiskillstore/marketplace, 430 stars), RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars) and Agentsop Multi Tenant RAG (agentsope/SkillAlchemy, 457 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vector Database Ops?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.