Neo4j Document Import Skill
neo4j-contrib/neo4j-skills
Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.
Build and query embedded GraphRAG knowledge graphs with DuckDB-backed vector, full-text, and hybrid search.
$ npx skills add bradAGI/GraphMemory --skill graphmemory -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bradAGI/GraphMemory graphmemory --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "graphmemory" agent skill from https://github.com/bradAGI/GraphMemory/tree/main into .claude/skills/graphmemory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "graphmemory", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bradAGI/GraphMemory --skill graphmemory -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bradAGI/GraphMemory graphmemory --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "graphmemory" agent skill from https://github.com/bradAGI/GraphMemory/tree/main into .agents/skills/graphmemory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "graphmemory", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bradAGI/GraphMemory --skill graphmemory -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bradAGI/GraphMemory graphmemory --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "graphmemory" agent skill from https://github.com/bradAGI/GraphMemory/tree/main into .cursor/skills/graphmemory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "graphmemory", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bradAGI/GraphMemory --skill graphmemory -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bradAGI/GraphMemory graphmemory --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "graphmemory" agent skill from https://github.com/bradAGI/GraphMemory/tree/main into .gemini/skills/graphmemory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "graphmemory", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bradAGI/GraphMemory graphmemoryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bradAGI/GraphMemory --skill graphmemory -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "graphmemory" agent skill from https://github.com/bradAGI/GraphMemory/tree/main into .github/skills/graphmemory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "graphmemory", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bradAGI/GraphMemory --skill graphmemory -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bradAGI/GraphMemory graphmemory --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "graphmemory" agent skill from https://github.com/bradAGI/GraphMemory/tree/main into .opencode/skills/graphmemory/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "graphmemory", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
graphmemoryBuild and query embedded GraphRAG knowledge graphs with DuckDB-backed vector, full-text, and hybrid search.
Graphmemory is an agent skill from bradAGI/GraphMemory. Build and query embedded GraphRAG knowledge graphs with DuckDB-backed vector, full-text, and hybrid search. Use when the user wants to store entities and relations, run RAG over a graph, extract knowledge graphs from text with DSPy, run graph algorithms (PageRank, centrality, components), merge/upsert nodes and edges, fuzzy-dedupe an existing graph, or visualize interactively in a browser. Trigger phrases include "knowledge graph", "GraphRAG", "graph database", "hybrid search", "extract entities and relations"…
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files (for example `README.md`, `examples/dspy_example_typed_pred.py` and `examples/lexical_graph.py`).
It sits in Knowledge Management, covering Knowledge graphs and Retrieval-augmented generation. It works with DuckDB. The repository describes itself as: GraphRAG database - hybrid graph / vector db. The licence is MIT.
Read from SKILL.md and the folder at commit efecc83. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pippython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Graphmemory loads about 3.1k tokens when it runs. Until then it costs about 148 tokens; SKILL.md has 727 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from bradAGI/GraphMemory at commit efecc83, republished under its MIT licence (© bradAGI). 727 words, ~3,087 tokens.
.claude/skills/graphmemory/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.Embedded GraphRAG database built on DuckDB. Single Python package — no server, no external services. Ships vector (HNSW), full-text (BM25), hybrid search, fluent query builder, multi-hop traversal, fuzzy dedup, DSPy extraction, NetworkX algorithms, and a zero-dep D3.js visualizer.
Do not use when: user already has Neo4j/Neptune/ArangoDB, or scale is hundreds of millions of nodes — GraphMemory is DuckDB-embedded, single-writer.
pip install graphmemory
pip install graphmemory[extraction] # DSPy entity/relation extraction
pip install graphmemory[algorithms] # NetworkX algorithms| User intent | Method |
|---|---|
| Insert one node | graph.insert_node(node) |
| Bulk insert | graph.bulk_insert_nodes(nodes) |
| Insert-or-update by property | graph.merge_node(node, match_keys=["name"]) |
| Fuzzy insert-or-update | graph.merge_node(node, match_keys=["name"], similarity_threshold=0.9) |
Dedupe edges on (src, tgt, relation) | graph.merge_edge(edge) |
| Clean up existing duplicates | graph.resolve_duplicates(match_keys=["name"], similarity_threshold=0.9) |
| Pure vector kNN | graph.nearest_nodes(vector, limit) |
| Pure BM25 text | graph.search_nodes(query, limit) |
| Combined text + vector | graph.hybrid_search(query, query_vector, text_weight, vector_weight) |
| Lookup by property | graph.nodes_by_attribute("name", "Alice") |
| Direct neighbors | graph.connected_nodes(node_id) |
| Multi-hop traversal | graph.query().traverse(source_id=id, depth=2).execute() |
| Filtered query | graph.query().match(type="Person").where(role="eng").execute() |
| GraphRAG context assembly | graph.retrieve(query, query_vector, max_hops, max_tokens) |
| End-to-end Q&A | graph.ask(query, query_vector, llm_callable=fn) |
| Extract + store from text | extract_and_merge(graph, text, match_keys=["name"]) |
| Extract in parallel across chunks | extract_and_merge_parallel(graph, chunks, max_workers=8) |
| PageRank / centrality | pagerank(graph), betweenness_centrality(graph) |
| Atomic block | with graph.transaction(): ... |
| Browser visualization | graph.visualize() |
from graphmemory import GraphMemory, Node, Edge, MergeStrategy
# database=None is in-memory; pass a path for persistence.
# vector_length and distance_metric are fixed at init time.
graph = GraphMemory(
database="graph.db",
vector_length=1536, # must match your embedding model
distance_metric="cosine", # "l2" | "cosine" | "inner_product"
hnsw_ef_construction=128,
hnsw_ef_search=64,
hnsw_m=16,
auto_index=True, # HNSW auto-built on init
max_retries=3, # transient IO error retry
)alice = Node(type="Person", properties={"name": "Alice"}, vector=embed("Alice"))
bob = Node(type="Person", properties={"name": "Bob"}, vector=embed("Bob"))
graph.insert_node(alice)
graph.insert_node(bob)
graph.insert_edge(Edge(source_id=alice.id, target_id=bob.id, relation="reports_to"))
# Idempotent re-ingest on a natural key
graph.merge_node(alice, match_keys=["name"])
# Fuzzy merge — tolerates "Alice Smith" vs "alice smith"
graph.merge_node(
alice,
match_keys=["name"],
similarity_threshold=0.9, # Jaro-Winkler threshold (1.0 = exact)
vector_threshold=0.2, # optional cosine distance cap
match_type=True, # also require same `type`
strategy=MergeStrategy.UPDATE, # UPDATE | REPLACE | KEEP
)results = graph.hybrid_search(
query_text="who leads ML?",
query_vector=embed("who leads ML?"),
text_weight=0.5,
vector_weight=0.5,
limit=10,
)
for r in results:
print(r.score, r.node.properties)# Context-only (own the prompt)
result = graph.retrieve(
query=q, query_vector=qv,
max_hops=2, max_tokens=4000, search_limit=10,
)
print(result.context_text, result.token_estimate, result.seed_node_count, result.total_node_count)
# End-to-end — llm_callable signature: (system_prompt, user_prompt) -> str
def my_llm(system, user):
return openai_client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "system", "content": system}, {"role": "user", "content": user}],
).choices[0].message.content
answer = graph.ask(query=q, query_vector=qv, llm_callable=my_llm)
print(answer["answer"])Pass llm_callable=None to get retrieval-only output — useful to inspect the context before wiring an LLM.
# Filter by type + property
engineers = graph.query().match(type="Person").where(role="engineer").execute()
# Multi-hop traversal — returns TraversalResult with depth + path
two_hop = graph.query().traverse(source_id=alice.id, depth=2).execute()
# Paginate + order
page = graph.query().match(type="Person").order_by("name").limit(20).offset(40).execute()
# Return edges instead of nodes
edges = graph.query().match(type="Person").edges().execute()import dspy
from graphmemory.extraction import extract_and_merge, extract_and_merge_parallel
dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))
# Single pass
node_results, edge_results = extract_and_merge(
graph, text, match_keys=["name"], similarity_threshold=0.88,
)
# Parallel across chunks — two phases: nodes first (all chunks), then edges
# with the full node context. Saturates your RPM.
node_results, edge_results = extract_and_merge_parallel(
graph,
chunks=paragraph_chunks,
match_keys=["name"],
similarity_threshold=0.88,
max_workers=8, # match your provider's RPM headroom
on_progress=lambda phase, done, total: print(f"{phase}: {done}/{total}"),
)with graph.transaction():
graph.insert_node(a)
graph.insert_node(b)
graph.insert_edge(Edge(source_id=a.id, target_id=b.id, relation="x"))
# Exception inside the block → ROLLBACK. Clean exit → COMMIT.Extract with a loose threshold, then clean up with a tighter one. This is the pattern in examples/test_ingest.py.
# Pass 1 — during ingest, be permissive to avoid fragmenting entities
extract_and_merge_parallel(graph, chunks, similarity_threshold=0.88, max_workers=50)
# Pass 2 — after ingest, resolve residual duplicates more strictly
clusters = graph.resolve_duplicates(
match_keys=["name"],
match_type=True,
similarity_threshold=0.9,
vector_threshold=0.15,
)
for c in clusters:
print(f"Kept {c.survivor.properties['name']}, merged {len(c.merged)} dups")resolve_duplicates picks the first-seen node as survivor, reassigns all incoming/outgoing edges to it, and deletes the rest. Self-loops from the reassignment are dropped.
Pattern from examples/lexical_graph.py:
prev = None
for chunk in chunks:
node = Node(type="Chunk", properties={"text": chunk}, vector=embed(chunk))
graph.insert_node(node)
if prev is not None:
graph.insert_edge(Edge(source_id=prev.id, target_id=node.id, relation="followed_by"))
prev = noderesult = graph.retrieve(query=q, query_vector=qv, max_hops=2, max_tokens=4000)
print(result.context_text) # See exactly what the LLM would receive
# Tune max_hops / max_tokens / search_limit before wiring ask()vector_length and distance_metric are locked at init. Swapping embedding models means a new database. Valid metrics: "l2", "cosine", "inner_product".insert_node — bulk_insert_nodes skips nodes whose vectors don't match vector_length and logs a warning. Validate upstream if correctness matters.auto_index=True). Tune via hnsw_ef_construction, hnsw_ef_search, hnsw_m. Call graph.compact_index() after heavy deletes to reclaim space (also called automatically by delete_node).search_nodes/hybrid_search call after writes rebuilds it. Expect first-search latency. Force a rebuild with graph.reindex() if you want it warm before traffic.(source_id, target_id, relation). Relations are normalized (lowercased, underscored) before comparison — "Reports To" and "reports_to" collide. Edge properties are NOT part of the key.delete_node cascades edges in both directions (as source AND as target). No orphan-edge safety net.merge_node strategies — UPDATE shallow-merges dicts (incoming wins on collision), REPLACE overwrites wholesale, KEEP only inserts if new. Pick intentionally.similarity_threshold=1.0 is exact match (the default). Lower it to enable Jaro-Winkler fuzzy matching on string properties. Non-string properties always use JSON equality.match_type=True (default) requires same type for merge. Set False to merge across types — rarely what you want.resolve_duplicates is O(n²)-ish in fuzzy mode. For large graphs, narrow with match_type and a tight vector_threshold first.extraction and algorithms are optional extras. Wrap imports in try/except or check pip show before recommending code that depends on them.@with_retry (exponential backoff on transient IO errors) are built in, but don't open the same file from multiple processes for concurrent writes.cursor() returns independent cursors for concurrent reads; the main connection is RLock-guarded for writes.ask() with llm_callable=None returns retrieval only — no generation. Always use this first to validate context before paying for LLM calls.| Model | Key fields |
|---|---|
Node | id: UUID, type: str | None, properties: dict, vector: list[float] |
Edge | id, source_id, target_id, relation: str, weight: float | None |
SearchResult | node, score (higher = better for both BM25 and hybrid) |
NearestNode | node, distance (lower = closer) |
TraversalResult | node, depth, path: list[UUID] |
RetrievalContext | node, relationships: list[dict], hop_distance: int |
RetrievalResult | query, contexts, context_text, token_estimate, seed_node_count, total_node_count |
MergeResult | node, created: bool (True = inserted, False = updated) |
EdgeMergeResult | edge, created: bool |
DuplicateCluster | survivor: Node, merged: list[Node] |
All models are Pydantic. IDs auto-generate as UUIDs.
examples/openai_example.py — OpenAI embeddings, similarity search, attribute lookupexamples/lexical_graph.py — chunked Wikipedia text with SentenceTransformer, sequential followed_by edgesexamples/dspy_example_typed_pred.py — DSPy typed-predictor extractionexamples/test_ingest.py — parallel extraction (50 workers, 0.88 threshold) + post-pass resolve_duplicates at 0.90Read examples/test_ingest.py before building a real ingest pipeline — it's the template.
python3 -m pytest tests/tests.py -v296 tests cover the public API. Run them when modifying the library.
© bradAGI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 15 other files in the repository root of bradAGI/GraphMemory.
Open the folder on GitHubat commit efecc83
Graphmemory next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Graphmemory this skillbradAGI/GraphMemory | 160 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Neo4j Document Import Skillneo4j-contrib/neo4j-skills | 114 | — | ~5.4k | Automated safety check: Notes | MIT | |
| Knowledge Graph Constructionwentorai/research-plugins | 298 | 1 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Cortexdb Memory Hermesliliang-cn/cortexdb | 274 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Cortexdb Memory Openclawliliang-cn/cortexdb | 274 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Docsmint Document ManagerHiAi-gg/docsmint | 118 | — | ~584 | Automated safety check: Pass | Apache-2.0 |
neo4j-contrib/neo4j-skills
Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.
wentorai/research-plugins
Build research knowledge graphs for literature synthesis and RAG systems
liliang-cn/cortexdb
Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…
liliang-cn/cortexdb
Give a Node.js agent (such as OpenClaw) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client npm package.
HiAi-gg/docsmint
Manage and research DocsMint documents through its scoped MCP tools, including categories, folders, hybrid search, GraphRAG, rerank, and index refresh.
joshzyj/open-scholar-skill
Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.
Works with
Categories
Build and query embedded GraphRAG knowledge graphs with DuckDB-backed vector, full-text, and hybrid search. Graphmemory is an agent skill from bradAGI/GraphMemory. Build and query embedded GraphRAG knowledge graphs with DuckDB-backed vector, full-text, and hybrid search.
Graphmemory fits situations like: the user wants to store entities and relations; run RAG over a graph; extract knowledge graphs from text with DSPy; run graph algorithms (PageRank.
Run `npx skills add bradAGI/GraphMemory --skill graphmemory -a claude-code`. Or copy the skill folder (the bradAGI/GraphMemory repository) into .claude/skills/graphmemory in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bradAGI/GraphMemory --skill graphmemory -a codex`. Or copy the skill folder (the bradAGI/GraphMemory repository) into .agents/skills/graphmemory in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bradAGI/GraphMemory --skill graphmemory -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/graphmemory, .gemini/skills/graphmemory, .github/skills/graphmemory and .opencode/skills/graphmemory in your project.
Going by SKILL.md and its folder, Graphmemory needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Graphmemory is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Graphmemory: Neo4j Document Import Skill (neo4j-contrib/neo4j-skills, 114 stars), Knowledge Graph Construction (wentorai/research-plugins, 298 stars), Cortexdb Memory Hermes (liliang-cn/cortexdb, 274 stars) and Cortexdb Memory Openclaw (liliang-cn/cortexdb, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bradAGI (a GitHub user) maintains it in bradAGI/GraphMemory, which has 160 GitHub stars. The repository was last updated on April 20, 2026.
Source: bradAGI/GraphMemory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.