Pgvector Semantic Search
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/neo4j-vector-index-skill .claude/skills/neo4j-vector-index-skill && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "neo4j-vector-index-skill" agent skill from https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skill into .claude/skills/neo4j-vector-index-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "neo4j-vector-index-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/neo4j-vector-index-skill .agents/skills/neo4j-vector-index-skill && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "neo4j-vector-index-skill" agent skill from https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skill into .agents/skills/neo4j-vector-index-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "neo4j-vector-index-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/neo4j-vector-index-skill .cursor/skills/neo4j-vector-index-skill && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "neo4j-vector-index-skill" agent skill from https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skill into .cursor/skills/neo4j-vector-index-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "neo4j-vector-index-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/neo4j-contrib/neo4j-skills.git --path neo4j-vector-index-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/neo4j-vector-index-skill .gemini/skills/neo4j-vector-index-skill && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "neo4j-vector-index-skill" agent skill from https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skill into .gemini/skills/neo4j-vector-index-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "neo4j-vector-index-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/neo4j-vector-index-skill .github/skills/neo4j-vector-index-skill && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "neo4j-vector-index-skill" agent skill from https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skill into .github/skills/neo4j-vector-index-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "neo4j-vector-index-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/neo4j-vector-index-skill .opencode/skills/neo4j-vector-index-skill && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "neo4j-vector-index-skill" agent skill from https://github.com/neo4j-contrib/neo4j-skills/tree/main/neo4j-vector-index-skill into .opencode/skills/neo4j-vector-index-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "neo4j-vector-index-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
neo4j-vector-index-skillCreate and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…
Neo4j Vector Index Skill is an agent skill from neo4j-contrib/neo4j-skills. Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or db.index.vector.queryNodes() procedure (deprecated 2026.04, still works on 2025.x), configure HNSW and quantization options, pick similarity function and embedding provider dimensions, and batch-update embeddings. Use when tasks involve CREATE VECTOR INDEX, vector.dimensions, cosine/euclidean search, embedding ingestion pipelines, semantic or…
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/hybrid-search.md`). Compatibility notes: Neo4j = 2025.01; SEARCH clause requires 2026.01+
It sits in AI & LLM Engineering, covering Embeddings, Vector databases and Knowledge graphs. It works with Neo4j. The repository describes itself as: Neo4j Skills for Coding and other Agents including Cypher. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bb30e1f. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashWebFetchFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are cypher, bash and python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
neo4j.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NEO4J_PASSWORDFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Neo4j >= 2025.01; SEARCH clause requires 2026.01+
From compatibility in the SKILL.md frontmatter.
Neo4j Vector Index Skill loads about 5.6k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 227 tokens; SKILL.md has 1,619 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, WebFetchAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from neo4j-contrib/neo4j-skills at commit bb30e1f, republished under its MIT licence (© neo4j-contrib). 1,619 words, ~5,632 tokens.
.claude/skills/neo4j-vector-index-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.CREATE VECTOR INDEX) on nodes or relationshipsSEARCH clause (2026.01+) or db.index.vector.queryNodes() (2025.x)neo4j-graphrag-skill)neo4j-graphrag-skilldb.index.fulltext.queryNodes) → neo4j-cypher-skillneo4j-gds-skillneo4j-cypher-skillDrives syntax choice:
CALL dbms.components() YIELD versions RETURN versions[0] AS neo4j_version| Version | Use |
|---|---|
2026.01 or higher | SEARCH clause (in-index filtering, preferred) |
2025.x | db.index.vector.queryNodes() procedure (deprecated 2026.04 — use SEARCH when on 2026.x) |
Node index (single label):
CYPHER 25
CREATE VECTOR INDEX chunk_embedding IF NOT EXISTS
FOR (c:Chunk) ON (c.embedding)
OPTIONS {
indexConfig: {
`vector.dimensions`: 1536,
`vector.similarity_function`: 'cosine',
`vector.quantization.type`: 'SCALAR',
`vector.hnsw.m`: 16,
`vector.hnsw.ef_construction`: 100
}
}Node index with filterable properties [2026.01+] — WITH declares which properties can be used in SEARCH ... WHERE:
CYPHER 25
CREATE VECTOR INDEX chunk_embedding IF NOT EXISTS
FOR (c:Chunk) ON (c.embedding)
WITH [c.source, c.lang, c.published_year] // stored as metadata; filterable in SEARCH WHERE
OPTIONS { indexConfig: { `vector.dimensions`: 1536, `vector.similarity_function`: 'cosine' } }Multi-label index with filterable properties [2026.01+]:
CYPHER 25
CREATE VECTOR INDEX doc_embedding IF NOT EXISTS
FOR (n:Document|Article) ON n.embedding
WITH [n.author, n.published_year, n.lang]
OPTIONS { indexConfig: { `vector.dimensions`: 1536, `vector.similarity_function`: 'cosine' } }Relationship index:
CYPHER 25
CREATE VECTOR INDEX rel_embedding IF NOT EXISTS
FOR ()-[r:HAS_CHUNK]-() ON (r.embedding)
OPTIONS { indexConfig: { `vector.dimensions`: 768, `vector.similarity_function`: 'cosine' } }WITH property types — only scalar types allowed: INTEGER, FLOAT, STRING, BOOLEAN, DATE, ZONED DATETIME, LOCAL DATETIME, ZONED TIME, LOCAL TIME, DURATION. Not allowed: LIST, POINT, or the vector property itself.
Index config reference:
| Parameter | Type | Default | Notes |
|---|---|---|---|
vector.dimensions | INTEGER 1–4096 | none | Required; must match embedding model exactly |
vector.similarity_function | STRING | 'cosine' | 'cosine' or 'euclidean' |
vector.quantization.type | STRING | 'binary' ('scalar' before 2026.08) | 'none', 'scalar', 'binary' [2026.06+, GA 2026.07]; reduces storage; binary smallest (1 bit per dimension), most aggressive; needs vector-2.0+ (5.18+); set explicitly for version-independent behavior |
vector.quantization.enabled | BOOLEAN | true | Deprecated 2026.06 — use vector.quantization.type; false without vector.quantization.type: 'none' fails index creation before 2026.07 |
vector.default_search_expansion_factor | FLOAT 1.0–10000.0 | 1.0 none / 1.5 scalar / 3.0 binary (was 2.0 before 2026.07) | [2026.06+, GA 2026.07]; value >1.0 on quantized vectors enables automatic rescoring with full-precision vectors (High-Fidelity Quantized search, HFQ); not settable at query time; existing indexes keep their stored value until rebuilt |
vector.hnsw.m | INTEGER 1–512 | 16 | HNSW graph connections; higher = better recall, more memory |
vector.hnsw.ef_construction | INTEGER 1–3200 | 100 | Build-time candidates; higher = better recall, slower build |
Unquantized vectors: set vector.quantization.type: 'none' alone — on 2026.06, vector.quantization.enabled: false without it errors (fixed 2026.07).
Similarity function choice:
| Use case | Function |
|---|---|
| Normalized embeddings (OpenAI, Cohere, Voyage, Google) | 'cosine' |
| Unnormalized / raw distance matters | 'euclidean' |
Index providers — latest selected automatically; not specifiable in Cypher 25. Check with SHOW VECTOR INDEXES YIELD name, indexProvider:
| Provider | Quantization support |
|---|---|
vector-2026.08 | High-Fidelity Quantized search for binary; 'binary' default |
vector-2026.07 | High-Fidelity Quantized search for scalar and binary |
vector-2026.06 | scalar and binary; required for binary + rescoring |
vector-3.0 (2025.09+) / vector-2.0 (5.18+) | scalar |
Changing quantization type, expansion factor, or provider requires drop + re-create + re-population. Indexes built before 2026.08 and rarely updated: re-create for 2026.09 rescored-binary/scalar speedup (~5× latency and throughput).
Memory — vector index files live in OS filesystem cache, not page cache. Leave RAM for: HNSW graph ≈ 8 B × vectors × vector.hnsw.m plus vector values (full precision ≈ 4 B × dims × vectors; scalar ÷4; binary ÷32). Size page cache for vector properties only if queries return or re-rank them. Details → Vector index memory configuration.
Index builds asynchronously — do NOT query until ONLINE:
SHOW VECTOR INDEXES YIELD name, state, populationPercent
WHERE name = 'chunk_embedding'
RETURN name, state, populationPercentPoll every 5s until state = 'ONLINE' and populationPercent = 100.0. If state = 'FAILED' → stop, check logs.
Shell poll (cypher-shell):
until cypher-shell -u neo4j -p "$NEO4J_PASSWORD" \
"SHOW VECTOR INDEXES YIELD name, state WHERE name='chunk_embedding' RETURN state" \
| grep -q ONLINE; do
sleep 5
doneBatch UNWIND pattern (use for > 100 nodes — never one-node-per-transaction):
from neo4j import GraphDatabase
driver = GraphDatabase.driver(uri, auth=(user, password))
def embed_batch(texts: list[str]) -> list[list[float]]:
response = openai_client.embeddings.create(
model="text-embedding-3-small", input=texts
)
return [r.embedding for r in response.data]
def store_embeddings(records: list[dict], batch_size: int = 500):
expected_dim = 1536 # must match vector.dimensions
texts = [r["text"] for r in records]
embeddings = embed_batch(texts)
for emb in embeddings:
assert len(emb) == expected_dim, f"Dim mismatch: {len(emb)} != {expected_dim}"
rows = [{"id": r["id"], "embedding": emb}
for r, emb in zip(records, embeddings)]
for i in range(0, len(rows), batch_size):
driver.execute_query(
"UNWIND $rows AS row MATCH (c:Chunk {id: row.id}) SET c.embedding = row.embedding",
rows=rows[i:i+batch_size]
)❌ Never create index after embeddings are already stored — always create index first. ✅ Create index → poll ONLINE → ingest embeddings.
CYPHER 25
MATCH (c:Chunk)
SEARCH c IN (
VECTOR INDEX chunk_embedding
FOR $queryEmbedding
LIMIT 10
) SCORE AS score
RETURN c.text, score
ORDER BY score DESCWith in-index filter [2026.01+] — properties must be declared in WITH at index creation:
// Index must have been created with: WITH [c.source, c.lang, c.published_year]
CYPHER 25
MATCH (c:Chunk)
SEARCH c IN (
VECTOR INDEX chunk_embedding
FOR $queryEmbedding
WHERE c.source = $source AND c.lang = 'en' AND c.published_year >= 2024
LIMIT 10
) SCORE AS score
RETURN c.text, c.source, score
ORDER BY score DESCFiltering strategy — choose one:
| Strategy | When to use | Tradeoff |
|---|---|---|
In-index WHERE [2026.01+] | Filters on pre-declared WITH properties; known at index design time | Fast, consistent latency; properties must be declared upfront |
| Post-filter (MATCH + procedure) | Arbitrary Cypher predicates, graph traversal, OR/NOT | Full flexibility; may over-fetch then discard |
| Pre-filter (MATCH first, then SEARCH) | Small known candidate set; exact nearest-neighbor within subset | Deterministic; slow on large candidate sets |
In-index WHERE hard limits [2026.01+]:
WITH [...] at index creation — undeclared properties silently fall back to post-filteringIN list membership allowed [2026.06+]INTEGER, FLOAT, STRING, BOOLEAN, temporal types — not VECTOR/LIST/POINTCYPHER 25
CALL db.index.vector.queryNodes('chunk_embedding', 50, $queryEmbedding)
YIELD node AS c, score
WHERE c.source = $source // post-filter: fetch more, then filter
RETURN c.text, score
ORDER BY score DESC LIMIT 10Relationship index procedure:
CYPHER 25
CALL db.index.vector.queryRelationships('rel_embedding', 5, $queryEmbedding)
YIELD relationship AS r, score
RETURN r.text, scoreSEARCH clause hard limits (all versions):
$indexName not allowed — use literal string)Vector search as entry point, then graph hop:
CYPHER 25
MATCH (c:Chunk)
SEARCH c IN (
VECTOR INDEX chunk_embedding
FOR $queryEmbedding
LIMIT 10
) SCORE AS score
MATCH (c)<-[:HAS_CHUNK]-(a:Article)
OPTIONAL MATCH (a)-[:MENTIONS]->(org:Organization)
RETURN c.text, a.title, score, collect(DISTINCT org.name) AS organizations
ORDER BY score DESCFor full retrieval_query pipelines, HybridCypherRetriever, or neo4j-graphrag library → delegate to neo4j-graphrag-skill.
Use hybrid search when one signal misses useful candidates: semantic vectors miss exact terms, lexical fulltext misses paraphrases, structural graph signals find topology not present in text. The common pattern is vector + fulltext, but the same approach works for several vector indexes, GDS-written embeddings, graph traversal scores, or any two+ ranked/scored sources. Load references/hybrid-search.md and apply its query shape.
Rules:
score DESC, stable_id ASC.UNION ALL branch returns same columns: matched node + contribution.sourceK > finalK; combine before final limiting.wrrf DESC, stable_id ASC.UNION ALL branches and new sourceWeights keys.| Provider / Model | Dimensions | Similarity | Notes |
|---|---|---|---|
| OpenAI text-embedding-3-small | 1536 | cosine | Default; reducible to 256–1536 via dimensions= param |
| OpenAI text-embedding-3-large | 3072 | cosine | Reducible to 256–3072 |
| OpenAI text-embedding-ada-002 | 1536 | cosine | Legacy; prefer 3-small |
| Cohere embed-v3 (English) | 1024 | cosine | Use input_type='search_document' at ingest, 'search_query' at query |
| Voyage voyage-3-large | 1024 | cosine | High quality; needs voyage-ai package |
| Google text-embedding-004 | 768 | cosine | Via Vertex AI |
| Ollama nomic-embed-text | 768 | cosine | Local dev/testing |
| Ollama mxbai-embed-large | 1024 | cosine | Local; production-quality |
vector.dimensions must exactly match model output — no auto-truncation.
Ad-hoc similarity (not for kNN search — use index for that):
MATCH (a:Chunk {id: $id1}), (b:Chunk {id: $id2})
RETURN vector.similarity.cosine(a.embedding, b.embedding) AS sim
// vector.similarity.euclidean(a, b) — same signature, 0–1 range
// vector_distance (2025.10+) — metrics: EUCLIDEAN, EUCLIDEAN_SQUARED, MANHATTAN, COSINE, DOT, HAMMING
// Returns distance (lower = more similar, inverse of similarity)
RETURN vector_distance(a.embedding, b.embedding, 'COSINE') AS dist
// vector_dimension_count (2025.10+)
RETURN vector_dimension_count(n.embedding) AS dims
// vector_norm (2025.20+) — metrics: EUCLIDEAN, MANHATTAN
RETURN vector_norm(n.embedding, 'EUCLIDEAN') AS normConvert LIST to typed VECTOR:
// vector(value, dimension, coordinateType)
// coordinateType: FLOAT64, FLOAT32, INTEGER8/16/32/64
WITH vector([1.0, 2.0, 3.0], 3, 'FLOAT32') AS v
RETURN vector_dimension_count(v)// Show all vector indexes with config
SHOW VECTOR INDEXES YIELD name, state, populationPercent,
labelsOrTypes, properties, indexConfig
RETURN name, state, populationPercent, labelsOrTypes, properties, indexConfig;
// Drop (node data unchanged — only index structure removed)
DROP INDEX chunk_embedding IF EXISTS;
// No ALTER VECTOR INDEX — to change dimensions or similarity function:
// 1. DROP INDEX old_index IF EXISTS
// 2. CREATE VECTOR INDEX new_index ... with new OPTIONS
// 3. Re-generate all embeddings with new model
// 4. Poll until ONLINE| Error | Cause | Fix |
|---|---|---|
IllegalArgumentException: Index dimension mismatch | Stored embedding dim ≠ vector.dimensions | Fix embed generation; drop + recreate index with correct dim |
| Search returns incomplete results | Index still POPULATING | Poll until state = 'ONLINE' |
Unknown procedure db.index.vector.queryNodes | Neo4j < 5.11 | No vector index support below 5.11; upgrade |
SEARCH clause not available | Neo4j < 2026.01 | Use queryNodes() procedure |
OR/NOT not allowed in SEARCH WHERE | SEARCH in-index filter restriction | Move complex predicates to outer WHERE after SEARCH |
| Zero results from correct query | Wrong similarity function or all-zeros embedding | Verify with vector.similarity.cosine(); check embed call succeeded |
| Score always 1.0 | All-zeros or identical vectors | Embedding generation failed; add dimension assertion before ingest |
vector.quantization.enabled / .type option rejected | provider vector-1.0 (Neo4j < 5.18) | Omit quantization option or upgrade to 5.18+ |
BINARY quantization rejected | provider older than vector-2026.06 | Upgrade to 2026.06+; SHOW VECTOR INDEXES YIELD name, indexProvider to check (vector-2026.07 adds High-Fidelity Quantized search) |
vector.dimensions matches embedding model output exactlycosine for normalized, euclidean for distance-based)state = 'ONLINE' before first querySEARCH clause on Neo4j >= 2026.01 (preferred); procedure fallback only on 2025.x (deprecated 2026.04)WHERE uses AND-only predicates with scalar types// Syntax (requires CYPHER 25)
CYPHER 25
// ai.text.embed(resource :: STRING, provider :: STRING, configuration :: MAP) :: VECTORProvider strings are lowercase ('openai', 'vertexai', 'bedrock-titan', 'azure-openai'). Full provider config → neo4j-genai-plugin-skill.
Full query pattern — embed at query time, search immediately (procedure fallback for 2025.x):
CYPHER 25
WITH ai.text.embed(
"What are good open source projects",
"openai",
{ token: $openaiKey, model: 'text-embedding-3-small' }) AS userEmbedding
CALL db.index.vector.queryNodes('chunk_embedding', 6, userEmbedding) // deprecated 2026.04
YIELD node AS c, score
RETURN c.text, score
ORDER BY score DESCWith SEARCH clause (2026.01+):
CYPHER 25
WITH ai.text.embed("my query", "openai", { token: $openaiKey, model: 'text-embedding-3-small' }) AS userEmbedding
MATCH (c:Chunk)
SEARCH c IN (VECTOR INDEX chunk_embedding FOR userEmbedding LIMIT 6) SCORE AS score
RETURN c.text, score
ORDER BY score DESCNever pass API key as string literal — use $openaiKey parameter (inject via driver params) or apoc.static.get().
Rule: Use same model at ingest time and query time — embeddings from different models are not comparable.
Deprecated (still works but do not use in new code):
genai.vector.encode() [deprecated] → use ai.text.embed() [2025.12]genai.vector.encodeBatch() [deprecated] → use CALL ai.text.embedBatch() [2025.12]genai.vector.listEncodingProviders() [deprecated] → use CALL ai.text.embed.providers() [2025.12]For full ai.text.* reference (completion, structured output, chat, tokenization) → neo4j-genai-plugin-skill.
Set vector property via Cypher (e.g. during LOAD CSV or MERGE pipeline):
LOAD CSV WITH HEADERS FROM 'https://example.com/data.csv' AS row
MERGE (q:Question {text: row.question})
WITH q, row
CALL db.create.setNodeVectorProperty(q, 'embedding', apoc.convert.fromJsonList(row.question_embedding))apoc.convert.fromJsonList() converts "[0.1,0.2,...]" STRING to LIST<FLOAT>. Python-generated embeddings → UNWIND batch pattern (Step 3).
Choose based on training loss function:
'cosine''euclidean''cosine' (all major hosted APIs use it)Common pitfall — wrong similarity function:
❌ Created index with 'euclidean' but model outputs L2-normalized vectors
→ scores are mathematically correct but rankings differ from expected cosine order
→ no error thrown; wrong results silently returned
✅ Verify: run vector.similarity.cosine(a.embedding, b.embedding) manually on known
similar pairs — score should be > 0.9 for near-duplicate textSanity check query after index creation:
MATCH (c:Chunk) WITH c LIMIT 2
WITH collect(c) AS nodes
RETURN vector.similarity.cosine(nodes[0].embedding, nodes[1].embedding) AS cosine_check,
vector.similarity.euclidean(nodes[0].embedding, nodes[1].embedding) AS euclidean_checkIf both return null → embeddings not set. If cosine returns 1.0 → identical vectors (embed call failed).
| Gotcha | Detail | Fix |
|---|---|---|
| Index not ONLINE at ingest time | Inserting nodes before index exists is valid — index auto-populates. But querying during POPULATING returns partial results | Always poll state = 'ONLINE' before first query |
| Wrong dimensions — silent failure | Stored vector dim ≠ vector.dimensions → IllegalArgumentException at query time, not at ingest time | Assert len(emb) == expected_dim before every SET c.embedding |
| Different models at ingest vs query | No error; cosine scores ~0.3–0.5 for clearly similar text | Use same model string/version for both; store model name as node metadata |
| Missing model at query | ai.text.embed returns null silently if provider config wrong | Test encode call standalone; check CYPHER 25 RETURN ai.text.embed(...) before embedding into pipeline |
| Large single-transaction ingest | One transaction for 10k nodes → OOM or timeout | Use UNWIND $rows ... CALL IN TRANSACTIONS OF 500 ROWS or Python batch loop |
| Chunk overlap not set | Adjacent chunks with no overlap → context at boundaries lost → poor recall for cross-paragraph queries | Set chunk_overlap ≥ 10% of chunk_size |
Load on demand:
genai.vector.encode()© neo4j-contrib, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in neo4j-vector-index-skill of neo4j-contrib/neo4j-skills.
Open the folder on GitHubat commit bb30e1f
Neo4j Vector Index Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Neo4j Vector Index Skill this skillneo4j-contrib/neo4j-skills | 114 | — | ~5.6k | Automated safety check: Notes | MIT | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Qdrant Search Qualitygithub/awesome-copilot | 40k | 1 repos | ~336 | Automated safety check: Pass | MIT | |
| Scholar RAGjoshzyj/open-scholar-skill | 168 | — | ~7.4k | Automated safety check: Notes | Custom licence | |
| Cortexdbliliang-cn/cortexdb | 274 | — | ~18k | Automated safety check: Warn | MIT | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~2.3k | Automated safety check: Pass | MIT |
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
github/awesome-copilot
Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.
joshzyj/open-scholar-skill
Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.
liliang-cn/cortexdb
Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
shuyu-labs/WebCode
Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.
neo4j-contrib/neo4j-skills
Manages Neo4j Aura Agents via the v2beta1 REST API — create, list, get, update, delete, and invoke Aura agents backed by an AuraDB instance.
neo4j-contrib/neo4j-skills
Generates, optimizes, and validates Cypher 25 queries for Neo4j 2025.x and 2026.x.
neo4j-contrib/neo4j-skills
Serverless Aura Graph Analytics (AGA) GDS Sessions — covers GdsSessions, AuraGraphDataScience, AuraAPICredentials, DbmsConnectionInfo, SessionMemory, getorcreate, remote graph projection with…
neo4j-contrib/neo4j-skills
Orchestrates zero-to-running-app in 8 stages — prerequisites → context → provision → model → load → explore → query → build.
neo4j-contrib/neo4j-skills
Provisions and manages Neo4j Aura instances via CLI (aura-cli v1.7+) or REST API.
neo4j-contrib/neo4j-skills
Neo4j .NET Driver v6 — IDriver lifecycle, DI registration (singleton), ExecutableQuery fluent API, ExecuteReadAsync/ExecuteWriteAsync managed transactions, IResultCursor (FetchAsync/ ToListAsync)…
Works with
Categories
Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…. Neo4j Vector Index Skill is an agent skill from neo4j-contrib/neo4j-skills.x), configure HNSW and quantization options, pick similarity function and embedding provider dimensions, and batch-update embeddings.
Neo4j Vector Index Skill fits situations like: tasks involve CREATE VECTOR INDEX; vector.dimensions; cosine/euclidean search; embedding ingestion pipelines.
Run `npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a claude-code`. Or copy the skill folder (neo4j-vector-index-skill in neo4j-contrib/neo4j-skills) into .claude/skills/neo4j-vector-index-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a codex`. Or copy the skill folder (neo4j-vector-index-skill in neo4j-contrib/neo4j-skills) into .agents/skills/neo4j-vector-index-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/neo4j-vector-index-skill, .gemini/skills/neo4j-vector-index-skill, .github/skills/neo4j-vector-index-skill and .opencode/skills/neo4j-vector-index-skill in your project.
Going by SKILL.md and its folder, Neo4j Vector Index Skill needs credentials named NEO4J_PASSWORD. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, WebFetch. Compatibility (from SKILL.md): Neo4j >= 2025.01; SEARCH clause requires 2026.01+.
SKILL.md names 1 domain. As links in the text: neo4j.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Neo4j Vector Index Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Neo4j Vector Index Skill: Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Qdrant Search Quality (github/awesome-copilot, 40k stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars) and Cortexdb (liliang-cn/cortexdb, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
neo4j-contrib (a GitHub organization) maintains it in neo4j-contrib/neo4j-skills, which has 114 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 9, 2026.
Source: neo4j-contrib/neo4j-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.