Agent skill

Neo4j Vector Index Skill

by neo4j-contrib in neo4j-contrib/neo4j-skills

Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…

MITAuto-check: notesAI & LLM Engineering

Install Neo4j Vector Index Skill

skills CLI
$ npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install neo4j-contrib/neo4j-skills neo4j-vector-index-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/neo4j-contrib/neo4j-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/neo4j-vector-index-skill .claude/skills/neo4j-vector-index-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
neo4j-vector-index-skill
GitHub stars
114
Token cost
~5.6k tokens
SKILL.md length
1,619 words
Files
3 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…

  • Works in 6 steps: Create Vector Index → Wait for Index ONLINE → Ingest Embeddings → …
  • Tasks involve CREATE VECTOR INDEX
  • SKILL.md covers When to Use, When NOT to Use, Pre-flight — Determine Version and Step 1 — Create Vector Index, plus 15 more sections
  • Needs NEO4J_PASSWORD

What it does

Neo4j Vector Index Skill is an agent skill from neo4j-contrib/neo4j-skills. Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or db.index.vector.queryNodes() procedure (deprecated 2026.04, still works on 2025.x), configure HNSW and quantization options, pick similarity function and embedding provider dimensions, and batch-update embeddings. Use when tasks involve CREATE VECTOR INDEX, vector.dimensions, cosine/euclidean search, embedding ingestion pipelines, semantic or…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/hybrid-search.md`). Compatibility notes: Neo4j = 2025.01; SEARCH clause requires 2026.01+

It sits in AI & LLM Engineering, covering Embeddings, Vector databases and Knowledge graphs. It works with Neo4j. The repository describes itself as: Neo4j Skills for Coding and other Agents including Cypher. The licence is MIT.

When your agent uses it

  • Tasks involve CREATE VECTOR INDEX
  • Vector.dimensions
  • Cosine/euclidean search
  • Embedding ingestion pipelines

Example prompts

  • “/neo4j-vector-index-skill”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Neo4j >= 2025.01; SEARCH clause requires 2026.01+
  • Pre-approved tools (allowed-tools): Bash, WebFetch

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Create Vector Index
  2. Wait for Index ONLINE
  3. Ingest Embeddings
  4. Run Vector Search
  5. Combine with Graph Traversal (simple cases)
  6. Hybrid Search

What it can do on your machine

Read from SKILL.md and the folder at commit bb30e1f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • WebFetch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are cypher, bash and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • neo4j.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NEO4J_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Neo4j >= 2025.01; SEARCH clause requires 2026.01+

    From compatibility in the SKILL.md frontmatter.

Context cost

Neo4j Vector Index Skill loads about 5.6k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 227 tokens; SKILL.md has 1,619 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~227
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, WebFetch

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from neo4j-contrib/neo4j-skills at commit bb30e1f, republished under its MIT licence (© neo4j-contrib). 1,619 words, ~5,632 tokens.

Download SKILL.mdSave it as .claude/skills/neo4j-vector-index-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
neo4j-vector-index-skill
description
Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or db.index.vector.queryNodes() procedure (deprecated 2026.04, still works on 2025.x), configure HNSW and quantization options, pick similarity function and embedding provider dimensions, and batch-update embeddings. Use when tasks involve CREATE VECTOR INDEX, vector.dimensions, cosine/euclidean search, embedding ingestion pipelines, semantic or structural nearest-neighbor lookup, or hybrid search (vector + fulltext, multiple vector sources, or graph-derived scores). Does NOT handle GraphRAG retrieval_query graph traversal — use neo4j-graphrag-skill. Does NOT handle fulltext-only/keyword-only search — use neo4j-cypher-skill. Does NOT compute GDS graph embeddings (FastRP, Node2Vec) — use neo4j-gds-skill.
allowed-tools
Bash, WebFetch
compatibility
Neo4j >= 2025.01; SEARCH clause requires 2026.01+
version
1.0.16

When to Use

  • Creating a vector index (CREATE VECTOR INDEX) on nodes or relationships
  • Running vector similarity / nearest-neighbor search
  • Storing embeddings on graph nodes during ingestion
  • Indexing/querying embeddings already written by GDS algorithms
  • Choosing similarity function, dimensions, HNSW params, or quantization
  • Using SEARCH clause (2026.01+) or db.index.vector.queryNodes() (2025.x)
  • Batch-updating embeddings after model change
  • Combining vector results with immediate graph neighborhood (full retrieval_query pipelines → neo4j-graphrag-skill)
  • Hybrid search that combines vector results with fulltext or other ranked sources

When NOT to Use

  • GraphRAG pipelines (VectorCypherRetriever, HybridCypherRetriever, retrieval_query) → neo4j-graphrag-skill
  • Fulltext-only / keyword-only search (FULLTEXT INDEX, db.index.fulltext.queryNodes) → neo4j-cypher-skill
  • Computing GDS graph embeddings (FastRP, Node2Vec, GraphSAGE) → neo4j-gds-skill
  • Index admin (list all indexes, drop range/text/lookup indexes) → neo4j-cypher-skill

Pre-flight — Determine Version

Drives syntax choice:

cypher
CALL dbms.components() YIELD versions RETURN versions[0] AS neo4j_version
VersionUse
2026.01 or higherSEARCH clause (in-index filtering, preferred)
2025.xdb.index.vector.queryNodes() procedure (deprecated 2026.04 — use SEARCH when on 2026.x)

Step 1 — Create Vector Index

Node index (single label):

cypher
CYPHER 25
CREATE VECTOR INDEX chunk_embedding IF NOT EXISTS
FOR (c:Chunk) ON (c.embedding)
OPTIONS {
  indexConfig: {
    `vector.dimensions`: 1536,
    `vector.similarity_function`: 'cosine',
    `vector.quantization.type`: 'SCALAR',
    `vector.hnsw.m`: 16,
    `vector.hnsw.ef_construction`: 100
  }
}

Node index with filterable properties [2026.01+] — WITH declares which properties can be used in SEARCH ... WHERE:

cypher
CYPHER 25
CREATE VECTOR INDEX chunk_embedding IF NOT EXISTS
FOR (c:Chunk) ON (c.embedding)
WITH [c.source, c.lang, c.published_year]  // stored as metadata; filterable in SEARCH WHERE
OPTIONS { indexConfig: { `vector.dimensions`: 1536, `vector.similarity_function`: 'cosine' } }

Multi-label index with filterable properties [2026.01+]:

cypher
CYPHER 25
CREATE VECTOR INDEX doc_embedding IF NOT EXISTS
FOR (n:Document|Article) ON n.embedding
WITH [n.author, n.published_year, n.lang]
OPTIONS { indexConfig: { `vector.dimensions`: 1536, `vector.similarity_function`: 'cosine' } }

Relationship index:

cypher
CYPHER 25
CREATE VECTOR INDEX rel_embedding IF NOT EXISTS
FOR ()-[r:HAS_CHUNK]-() ON (r.embedding)
OPTIONS { indexConfig: { `vector.dimensions`: 768, `vector.similarity_function`: 'cosine' } }

WITH property types — only scalar types allowed: INTEGER, FLOAT, STRING, BOOLEAN, DATE, ZONED DATETIME, LOCAL DATETIME, ZONED TIME, LOCAL TIME, DURATION. Not allowed: LIST, POINT, or the vector property itself.

Index config reference:

ParameterTypeDefaultNotes
vector.dimensionsINTEGER 1–4096noneRequired; must match embedding model exactly
vector.similarity_functionSTRING'cosine''cosine' or 'euclidean'
vector.quantization.typeSTRING'binary' ('scalar' before 2026.08)'none', 'scalar', 'binary' [2026.06+, GA 2026.07]; reduces storage; binary smallest (1 bit per dimension), most aggressive; needs vector-2.0+ (5.18+); set explicitly for version-independent behavior
vector.quantization.enabledBOOLEANtrueDeprecated 2026.06 — use vector.quantization.type; false without vector.quantization.type: 'none' fails index creation before 2026.07
vector.default_search_expansion_factorFLOAT 1.0–10000.01.0 none / 1.5 scalar / 3.0 binary (was 2.0 before 2026.07)[2026.06+, GA 2026.07]; value >1.0 on quantized vectors enables automatic rescoring with full-precision vectors (High-Fidelity Quantized search, HFQ); not settable at query time; existing indexes keep their stored value until rebuilt
vector.hnsw.mINTEGER 1–51216HNSW graph connections; higher = better recall, more memory
vector.hnsw.ef_constructionINTEGER 1–3200100Build-time candidates; higher = better recall, slower build

Unquantized vectors: set vector.quantization.type: 'none' alone — on 2026.06, vector.quantization.enabled: false without it errors (fixed 2026.07).

Similarity function choice:

Use caseFunction
Normalized embeddings (OpenAI, Cohere, Voyage, Google)'cosine'
Unnormalized / raw distance matters'euclidean'

Index providers — latest selected automatically; not specifiable in Cypher 25. Check with SHOW VECTOR INDEXES YIELD name, indexProvider:

ProviderQuantization support
vector-2026.08High-Fidelity Quantized search for binary; 'binary' default
vector-2026.07High-Fidelity Quantized search for scalar and binary
vector-2026.06scalar and binary; required for binary + rescoring
vector-3.0 (2025.09+) / vector-2.0 (5.18+)scalar

Changing quantization type, expansion factor, or provider requires drop + re-create + re-population. Indexes built before 2026.08 and rarely updated: re-create for 2026.09 rescored-binary/scalar speedup (~5× latency and throughput).

Memory — vector index files live in OS filesystem cache, not page cache. Leave RAM for: HNSW graph ≈ 8 B × vectors × vector.hnsw.m plus vector values (full precision ≈ 4 B × dims × vectors; scalar ÷4; binary ÷32). Size page cache for vector properties only if queries return or re-rank them. Details → Vector index memory configuration.


Step 2 — Wait for Index ONLINE

Index builds asynchronously — do NOT query until ONLINE:

cypher
SHOW VECTOR INDEXES YIELD name, state, populationPercent
WHERE name = 'chunk_embedding'
RETURN name, state, populationPercent

Poll every 5s until state = 'ONLINE' and populationPercent = 100.0. If state = 'FAILED' → stop, check logs.

Shell poll (cypher-shell):

bash
until cypher-shell -u neo4j -p "$NEO4J_PASSWORD" \
  "SHOW VECTOR INDEXES YIELD name, state WHERE name='chunk_embedding' RETURN state" \
  | grep -q ONLINE; do
  sleep 5
done

Step 3 — Ingest Embeddings

Batch UNWIND pattern (use for > 100 nodes — never one-node-per-transaction):

python
from neo4j import GraphDatabase

driver = GraphDatabase.driver(uri, auth=(user, password))

def embed_batch(texts: list[str]) -> list[list[float]]:
    response = openai_client.embeddings.create(
        model="text-embedding-3-small", input=texts
    )
    return [r.embedding for r in response.data]

def store_embeddings(records: list[dict], batch_size: int = 500):
    expected_dim = 1536  # must match vector.dimensions
    texts = [r["text"] for r in records]
    embeddings = embed_batch(texts)
    for emb in embeddings:
        assert len(emb) == expected_dim, f"Dim mismatch: {len(emb)} != {expected_dim}"
    rows = [{"id": r["id"], "embedding": emb}
            for r, emb in zip(records, embeddings)]
    for i in range(0, len(rows), batch_size):
        driver.execute_query(
            "UNWIND $rows AS row MATCH (c:Chunk {id: row.id}) SET c.embedding = row.embedding",
            rows=rows[i:i+batch_size]
        )

❌ Never create index after embeddings are already stored — always create index first. ✅ Create index → poll ONLINE → ingest embeddings.


SEARCH clause (2026.01+, preferred)
cypher
CYPHER 25
MATCH (c:Chunk)
  SEARCH c IN (
    VECTOR INDEX chunk_embedding
    FOR $queryEmbedding
    LIMIT 10
  ) SCORE AS score
RETURN c.text, score
ORDER BY score DESC

With in-index filter [2026.01+] — properties must be declared in WITH at index creation:

cypher
// Index must have been created with: WITH [c.source, c.lang, c.published_year]
CYPHER 25
MATCH (c:Chunk)
  SEARCH c IN (
    VECTOR INDEX chunk_embedding
    FOR $queryEmbedding
    WHERE c.source = $source AND c.lang = 'en' AND c.published_year >= 2024
    LIMIT 10
  ) SCORE AS score
RETURN c.text, c.source, score
ORDER BY score DESC

Filtering strategy — choose one:

StrategyWhen to useTradeoff
In-index WHERE [2026.01+]Filters on pre-declared WITH properties; known at index design timeFast, consistent latency; properties must be declared upfront
Post-filter (MATCH + procedure)Arbitrary Cypher predicates, graph traversal, OR/NOTFull flexibility; may over-fetch then discard
Pre-filter (MATCH first, then SEARCH)Small known candidate set; exact nearest-neighbor within subsetDeterministic; slow on large candidate sets

In-index WHERE hard limits [2026.01+]:

  • Property must be listed in WITH [...] at index creation — undeclared properties silently fall back to post-filtering
  • AND predicates only — no OR, NOT, string ops. IN list membership allowed [2026.06+]
  • Scalar types only: INTEGER, FLOAT, STRING, BOOLEAN, temporal types — not VECTOR/LIST/POINT
Post-filter pattern (2025.x or arbitrary predicates)
cypher
CYPHER 25
CALL db.index.vector.queryNodes('chunk_embedding', 50, $queryEmbedding)
YIELD node AS c, score
WHERE c.source = $source    // post-filter: fetch more, then filter
RETURN c.text, score
ORDER BY score DESC LIMIT 10

Relationship index procedure:

cypher
CYPHER 25
CALL db.index.vector.queryRelationships('rel_embedding', 5, $queryEmbedding)
YIELD relationship AS r, score
RETURN r.text, score

SEARCH clause hard limits (all versions):

  • Index name cannot be a parameter ($indexName not allowed — use literal string)
  • Binding variable must come from the enclosing MATCH pattern
  • Query vector cannot reference the binding variable

Step 5 — Combine with Graph Traversal (simple cases)

Vector search as entry point, then graph hop:

cypher
CYPHER 25
MATCH (c:Chunk)
  SEARCH c IN (
    VECTOR INDEX chunk_embedding
    FOR $queryEmbedding
    LIMIT 10
  ) SCORE AS score
MATCH (c)<-[:HAS_CHUNK]-(a:Article)
OPTIONAL MATCH (a)-[:MENTIONS]->(org:Organization)
RETURN c.text, a.title, score, collect(DISTINCT org.name) AS organizations
ORDER BY score DESC

For full retrieval_query pipelines, HybridCypherRetriever, or neo4j-graphrag library → delegate to neo4j-graphrag-skill.


Use hybrid search when one signal misses useful candidates: semantic vectors miss exact terms, lexical fulltext misses paraphrases, structural graph signals find topology not present in text. The common pattern is vector + fulltext, but the same approach works for several vector indexes, GDS-written embeddings, graph traversal scores, or any two+ ranked/scored sources. Load references/hybrid-search.md and apply its query shape.

Rules:

  • Run each source independently; rank each by score DESC, stable_id ASC.
  • Combine by rank, not raw scores; fulltext and vector scores are not comparable.
  • Every UNION ALL branch returns same columns: matched node + contribution.
  • Use sourceK > finalK; combine before final limiting.
  • Sum contributions per node; order final rows by wrrf DESC, stable_id ASC.
  • Add more sources with extra UNION ALL branches and new sourceWeights keys.

Embedding Provider Quick-Reference

Provider / ModelDimensionsSimilarityNotes
OpenAI text-embedding-3-small1536cosineDefault; reducible to 256–1536 via dimensions= param
OpenAI text-embedding-3-large3072cosineReducible to 256–3072
OpenAI text-embedding-ada-0021536cosineLegacy; prefer 3-small
Cohere embed-v3 (English)1024cosineUse input_type='search_document' at ingest, 'search_query' at query
Voyage voyage-3-large1024cosineHigh quality; needs voyage-ai package
Google text-embedding-004768cosineVia Vertex AI
Ollama nomic-embed-text768cosineLocal dev/testing
Ollama mxbai-embed-large1024cosineLocal; production-quality

vector.dimensions must exactly match model output — no auto-truncation.


Show full SKILL.md (648 more words)Show less

Vector Functions

Ad-hoc similarity (not for kNN search — use index for that):

cypher
MATCH (a:Chunk {id: $id1}), (b:Chunk {id: $id2})
RETURN vector.similarity.cosine(a.embedding, b.embedding) AS sim
// vector.similarity.euclidean(a, b) — same signature, 0–1 range

// vector_distance (2025.10+) — metrics: EUCLIDEAN, EUCLIDEAN_SQUARED, MANHATTAN, COSINE, DOT, HAMMING
// Returns distance (lower = more similar, inverse of similarity)
RETURN vector_distance(a.embedding, b.embedding, 'COSINE') AS dist

// vector_dimension_count (2025.10+)
RETURN vector_dimension_count(n.embedding) AS dims

// vector_norm (2025.20+) — metrics: EUCLIDEAN, MANHATTAN
RETURN vector_norm(n.embedding, 'EUCLIDEAN') AS norm

Convert LIST to typed VECTOR:

cypher
// vector(value, dimension, coordinateType)
// coordinateType: FLOAT64, FLOAT32, INTEGER8/16/32/64
WITH vector([1.0, 2.0, 3.0], 3, 'FLOAT32') AS v
RETURN vector_dimension_count(v)

Index Management

cypher
// Show all vector indexes with config
SHOW VECTOR INDEXES YIELD name, state, populationPercent,
  labelsOrTypes, properties, indexConfig
RETURN name, state, populationPercent, labelsOrTypes, properties, indexConfig;

// Drop (node data unchanged — only index structure removed)
DROP INDEX chunk_embedding IF EXISTS;

// No ALTER VECTOR INDEX — to change dimensions or similarity function:
// 1. DROP INDEX old_index IF EXISTS
// 2. CREATE VECTOR INDEX new_index ... with new OPTIONS
// 3. Re-generate all embeddings with new model
// 4. Poll until ONLINE

Common Errors

ErrorCauseFix
IllegalArgumentException: Index dimension mismatchStored embedding dim ≠ vector.dimensionsFix embed generation; drop + recreate index with correct dim
Search returns incomplete resultsIndex still POPULATINGPoll until state = 'ONLINE'
Unknown procedure db.index.vector.queryNodesNeo4j < 5.11No vector index support below 5.11; upgrade
SEARCH clause not availableNeo4j < 2026.01Use queryNodes() procedure
OR/NOT not allowed in SEARCH WHERESEARCH in-index filter restrictionMove complex predicates to outer WHERE after SEARCH
Zero results from correct queryWrong similarity function or all-zeros embeddingVerify with vector.similarity.cosine(); check embed call succeeded
Score always 1.0All-zeros or identical vectorsEmbedding generation failed; add dimension assertion before ingest
vector.quantization.enabled / .type option rejectedprovider vector-1.0 (Neo4j < 5.18)Omit quantization option or upgrade to 5.18+
BINARY quantization rejectedprovider older than vector-2026.06Upgrade to 2026.06+; SHOW VECTOR INDEXES YIELD name, indexProvider to check (vector-2026.07 adds High-Fidelity Quantized search)

Checklist

  • vector.dimensions matches embedding model output exactly
  • Vector index created before ingesting embeddings
  • Similarity function chosen explicitly (cosine for normalized, euclidean for distance-based)
  • Index polled to state = 'ONLINE' before first query
  • Dimension validated on every embedding before ingest
  • SEARCH clause on Neo4j >= 2026.01 (preferred); procedure fallback only on 2025.x (deprecated 2026.04)
  • SEARCH WHERE uses AND-only predicates with scalar types
  • Batch UNWIND pattern used for > 100 nodes
  • If model changes: drop index → recreate with new dimensions → re-generate all embeddings

In-Cypher Embedding Generation — ai.text.embed() [2025.12]

cypher
// Syntax (requires CYPHER 25)
CYPHER 25
// ai.text.embed(resource :: STRING, provider :: STRING, configuration :: MAP) :: VECTOR

Provider strings are lowercase ('openai', 'vertexai', 'bedrock-titan', 'azure-openai'). Full provider config → neo4j-genai-plugin-skill.

Full query pattern — embed at query time, search immediately (procedure fallback for 2025.x):

cypher
CYPHER 25
WITH ai.text.embed(
    "What are good open source projects",
    "openai",
    { token: $openaiKey, model: 'text-embedding-3-small' }) AS userEmbedding
CALL db.index.vector.queryNodes('chunk_embedding', 6, userEmbedding)  // deprecated 2026.04
YIELD node AS c, score
RETURN c.text, score
ORDER BY score DESC

With SEARCH clause (2026.01+):

cypher
CYPHER 25
WITH ai.text.embed("my query", "openai", { token: $openaiKey, model: 'text-embedding-3-small' }) AS userEmbedding
MATCH (c:Chunk)
  SEARCH c IN (VECTOR INDEX chunk_embedding FOR userEmbedding LIMIT 6) SCORE AS score
RETURN c.text, score
ORDER BY score DESC

Never pass API key as string literal — use $openaiKey parameter (inject via driver params) or apoc.static.get().

Rule: Use same model at ingest time and query time — embeddings from different models are not comparable.

Deprecated (still works but do not use in new code):

  • genai.vector.encode() [deprecated] → use ai.text.embed() [2025.12]
  • genai.vector.encodeBatch() [deprecated] → use CALL ai.text.embedBatch() [2025.12]
  • genai.vector.listEncodingProviders() [deprecated] → use CALL ai.text.embed.providers() [2025.12]

For full ai.text.* reference (completion, structured output, chat, tokenization) → neo4j-genai-plugin-skill.


Cypher-Based Embedding Ingestion — db.create.setNodeVectorProperty

Set vector property via Cypher (e.g. during LOAD CSV or MERGE pipeline):

cypher
LOAD CSV WITH HEADERS FROM 'https://example.com/data.csv' AS row
MERGE (q:Question {text: row.question})
WITH q, row
CALL db.create.setNodeVectorProperty(q, 'embedding', apoc.convert.fromJsonList(row.question_embedding))

apoc.convert.fromJsonList() converts "[0.1,0.2,...]" STRING to LIST<FLOAT>. Python-generated embeddings → UNWIND batch pattern (Step 3).


Similarity Function — Extended Guidance

Choose based on training loss function:

  • Check embedding model docs — models trained with cosine loss → use 'cosine'
  • Models trained with L2/Euclidean loss → use 'euclidean'
  • When docs are silent: default to 'cosine' (all major hosted APIs use it)

Common pitfall — wrong similarity function:

❌ Created index with 'euclidean' but model outputs L2-normalized vectors
   → scores are mathematically correct but rankings differ from expected cosine order
   → no error thrown; wrong results silently returned
✅ Verify: run vector.similarity.cosine(a.embedding, b.embedding) manually on known
   similar pairs — score should be > 0.9 for near-duplicate text

Sanity check query after index creation:

cypher
MATCH (c:Chunk) WITH c LIMIT 2
WITH collect(c) AS nodes
RETURN vector.similarity.cosine(nodes[0].embedding, nodes[1].embedding) AS cosine_check,
       vector.similarity.euclidean(nodes[0].embedding, nodes[1].embedding) AS euclidean_check

If both return null → embeddings not set. If cosine returns 1.0 → identical vectors (embed call failed).


Gotchas — Extended

GotchaDetailFix
Index not ONLINE at ingest timeInserting nodes before index exists is valid — index auto-populates. But querying during POPULATING returns partial resultsAlways poll state = 'ONLINE' before first query
Wrong dimensions — silent failureStored vector dim ≠ vector.dimensions → IllegalArgumentException at query time, not at ingest timeAssert len(emb) == expected_dim before every SET c.embedding
Different models at ingest vs queryNo error; cosine scores ~0.3–0.5 for clearly similar textUse same model string/version for both; store model name as node metadata
Missing model at queryai.text.embed returns null silently if provider config wrongTest encode call standalone; check CYPHER 25 RETURN ai.text.embed(...) before embedding into pipeline
Large single-transaction ingestOne transaction for 10k nodes → OOM or timeoutUse UNWIND $rows ... CALL IN TRANSACTIONS OF 500 ROWS or Python batch loop
Chunk overlap not setAdjacent chunks with no overlap → context at boundaries lost → poor recall for cross-paragraph queriesSet chunk_overlap ≥ 10% of chunk_size

References

Load on demand:

© neo4j-contrib, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in neo4j-vector-index-skill of neo4j-contrib/neo4j-skills.

  • SKILL.md
  • README.md
  • references/hybrid-search.md

Open the folder on GitHubat commit bb30e1f

Compare with similar skills

Neo4j Vector Index Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Neo4j Vector Index Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Neo4j Vector Index Skill this skillneo4j-contrib/neo4j-skills114—~5.6kAutomated safety check: NotesMIT
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
Qdrant Search Qualitygithub/awesome-copilot40k1 repos~336Automated safety check: PassMIT
Scholar RAGjoshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesCustom licence
Cortexdbliliang-cn/cortexdb274—~18kAutomated safety check: WarnMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Qdrant Search Quality

    github/awesome-copilot

    Official

    Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~336 tokens
    AI & LLM EngineeringAuto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 23 days ago
    Research & ScienceAuto-check: notes
  • Cortexdb

    liliang-cn/cortexdb

    Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.

    274 GitHub stars~18k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check: warnings
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from neo4j-contrib/neo4j-skills

All 28 skills in this repo
  • Neo4j Aura Agent Skill

    neo4j-contrib/neo4j-skills

    Manages Neo4j Aura Agents via the v2beta1 REST API — create, list, get, update, delete, and invoke Aura agents backed by an AuraDB instance.

    114 GitHub stars~4.4k tokensUpdated yesterday
    Auto-check: notes
  • Neo4j Cypher Skill

    neo4j-contrib/neo4j-skills

    Generates, optimizes, and validates Cypher 25 queries for Neo4j 2025.x and 2026.x.

    114 GitHub starsUsed in 1 repo~6.1k tokens
    Auto-check passed
  • Neo4j Aura Graph Analytics Skill

    neo4j-contrib/neo4j-skills

    Serverless Aura Graph Analytics (AGA) GDS Sessions — covers GdsSessions, AuraGraphDataScience, AuraAPICredentials, DbmsConnectionInfo, SessionMemory, getorcreate, remote graph projection with…

    114 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check: notes
  • Neo4j Getting Started Skill

    neo4j-contrib/neo4j-skills

    Orchestrates zero-to-running-app in 8 stages — prerequisites → context → provision → model → load → explore → query → build.

    114 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check: warnings
  • Neo4j Aura Provisioning Skill

    neo4j-contrib/neo4j-skills

    Provisions and manages Neo4j Aura instances via CLI (aura-cli v1.7+) or REST API.

    114 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check: notes
  • Neo4j Driver Dotnet Skill

    neo4j-contrib/neo4j-skills

    Neo4j .NET Driver v6 — IDriver lifecycle, DI registration (singleton), ExecutableQuery fluent API, ExecuteReadAsync/ExecuteWriteAsync managed transactions, IResultCursor (FetchAsync/ ToListAsync)…

    114 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Neo4j Vector Index Skill

What does Neo4j Vector Index Skill do?

Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…. Neo4j Vector Index Skill is an agent skill from neo4j-contrib/neo4j-skills.x), configure HNSW and quantization options, pick similarity function and embedding provider dimensions, and batch-update embeddings.

When should I use Neo4j Vector Index Skill?

Neo4j Vector Index Skill fits situations like: tasks involve CREATE VECTOR INDEX; vector.dimensions; cosine/euclidean search; embedding ingestion pipelines.

How do I install Neo4j Vector Index Skill in Claude Code?

Run `npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a claude-code`. Or copy the skill folder (neo4j-vector-index-skill in neo4j-contrib/neo4j-skills) into .claude/skills/neo4j-vector-index-skill in your project. Claude Code loads it when a task matches its description.

How do I install Neo4j Vector Index Skill in Codex?

Run `npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a codex`. Or copy the skill folder (neo4j-vector-index-skill in neo4j-contrib/neo4j-skills) into .agents/skills/neo4j-vector-index-skill in your project. Codex loads it when a task matches its description.

Can I use Neo4j Vector Index Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add neo4j-contrib/neo4j-skills --skill neo4j-vector-index-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/neo4j-vector-index-skill, .gemini/skills/neo4j-vector-index-skill, .github/skills/neo4j-vector-index-skill and .opencode/skills/neo4j-vector-index-skill in your project.

What does Neo4j Vector Index Skill need to run?

Going by SKILL.md and its folder, Neo4j Vector Index Skill needs credentials named NEO4J_PASSWORD. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, WebFetch. Compatibility (from SKILL.md): Neo4j >= 2025.01; SEARCH clause requires 2026.01+.

Does Neo4j Vector Index Skill access the network?

SKILL.md names 1 domain. As links in the text: neo4j.com. This is read from the text; nothing was executed.

Is Neo4j Vector Index Skill safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Neo4j Vector Index Skill use?

Neo4j Vector Index Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Neo4j Vector Index Skill use?

About 5.6k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Neo4j Vector Index Skill?

Skills that share tags, products or a category with Neo4j Vector Index Skill: Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Qdrant Search Quality (github/awesome-copilot, 40k stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars) and Cortexdb (liliang-cn/cortexdb, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Neo4j Vector Index Skill?

neo4j-contrib (a GitHub organization) maintains it in neo4j-contrib/neo4j-skills, which has 114 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 9, 2026.

Source: neo4j-contrib/neo4j-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.