DB
oracle/skills
Oracle Database guidance for SQL, PL/SQL, SQLcl, ORDS, Oracle Vector SDK, administration, app development, performance, security, migrations, and agent-safe database workflows.
Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add liliang-cn/cortexdb --skill cortexdb -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install liliang-cn/cortexdb cortexdb --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cortexdb" agent skill from https://github.com/liliang-cn/cortexdb/tree/main into .claude/skills/cortexdb/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cortexdb", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add liliang-cn/cortexdb --skill cortexdb -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install liliang-cn/cortexdb cortexdb --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cortexdb" agent skill from https://github.com/liliang-cn/cortexdb/tree/main into .agents/skills/cortexdb/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cortexdb", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add liliang-cn/cortexdb --skill cortexdb -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install liliang-cn/cortexdb cortexdb --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cortexdb" agent skill from https://github.com/liliang-cn/cortexdb/tree/main into .cursor/skills/cortexdb/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cortexdb", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add liliang-cn/cortexdb --skill cortexdb -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install liliang-cn/cortexdb cortexdb --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cortexdb" agent skill from https://github.com/liliang-cn/cortexdb/tree/main into .gemini/skills/cortexdb/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cortexdb", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install liliang-cn/cortexdb cortexdbInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add liliang-cn/cortexdb --skill cortexdb -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cortexdb" agent skill from https://github.com/liliang-cn/cortexdb/tree/main into .github/skills/cortexdb/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cortexdb", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add liliang-cn/cortexdb --skill cortexdb -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install liliang-cn/cortexdb cortexdb --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cortexdb" agent skill from https://github.com/liliang-cn/cortexdb/tree/main into .opencode/skills/cortexdb/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cortexdb", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cortexdbUse CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.
Cortexdb is an agent skill from liliang-cn/cortexdb. Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling. Use when working with CortexDB, embeddings, memory, RAG, GraphRAG, knowledge graph, RDF, SPARQL, SHACL, data import, CSV/SQL dump import, memoryflow, graphflow, importflow, or MCP tools.
Its SKILL.md is about 18k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1671 other files, including scripts (for example `.agents/plugins/marketplace.json`, `.agents/product-marketing.md` and `.claude-plugin/marketplace.json`).
It sits in Knowledge Management, covering Knowledge graphs, MCP servers and Vector databases. It works with SQL and SQLite. The repository describes itself as: AI memory and a knowledge graph in one SQLite file. Pure Go: vectors, RAG, agent memory, RDF/SPARQL, Cypher, 80+ MCP tools. Works without an embedding model. The licence is MIT.
Read from SKILL.md and the folder at commit 17f8a4f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
gocargopipnpmnpxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
w3.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
CORTEXDB_GRPC_TOKENOPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cortexdb loads about 18k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 7,305 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
OPENAI_BASE_URL=http://43.167.167.6:8080/v1Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from liliang-cn/cortexdb at commit 17f8a4f, republished under its MIT licence (© liliang-cn). 7,305 words, ~18,178 tokens.
.claude/skills/cortexdb/SKILL.md (or your agent's skills folder). This skill also uses 1665 other files; get the full folder from GitHub.CortexDB is a pure-Go, single-file AI memory and knowledge graph library built on SQLite.
Use the right layer:
pkg/cortexdb
Main public DB facade: vectors, text search, knowledge, memory, KnowledgeMemory, KG, tools, MCP.
pkg/memoryflow
Agent memory workflow: transcript ingest, recall, wake-up layers, diary, promotion.
pkg/graphflow
Corpus-to-graph workflow: extraction schema, build, analyze, report, export, HTML.
pkg/importflow
External structured-data import (CSV / MySQL-PG dumps) into RAG + knowledge-graph foundations, AI-assisted mapping optional.
pkg/graph
Low-level graph engine: property graph, RDF triples/quads, SPARQL, RDFS, SHACL.
pkg/core
Storage engine (SQLite by default, PostgreSQL + pgvector by DSN), embeddings, FTS5,
vector indexes, chat/session primitives.Default recommendation:
pkg/cortexdb for application code.pkg/memoryflow for chat/session/agent memory workflows.pkg/graphflow for document/corpus-to-graph extraction and report/export workflows.pkg/graph only for low-level RDF/SPARQL/RDFS/SHACL or property graph control.Choose by the job to be done:
| Scenario | Start with | What it really does |
|---|---|---|
| Store vectors, FTS5, or run simple TopK retrieval | pkg/cortexdb Quick(); drop to pkg/core only if needed | Thinnest layer for pure retrieval without knowledge or memory semantics. |
| Document RAG or knowledge-base QA | SaveKnowledge + SearchKnowledge | Ingest chunks, builds retrieval artifacts, can attach entities and relations, then runs semantic or lexical GraphRAG retrieval. |
| User preferences, session memory, or long-term memory | SaveMemory + SearchMemory | Resolves a memory bucket from scope / user / session / namespace, stores memory, then uses semantic retrieval or lexical fallback. |
| Retrieve both memory and knowledge and assemble prompt context | db.KnowledgeMemory().Recall() or BuildContextPack() | Runs memory recall and knowledge recall, then fuses entities, chunks, and context into one packed response. |
| Full chat-assistant memory workflow | pkg/memoryflow | Turns transcripts into episodic memories, supports recall, wake-up layers, and end-of-session promotion. |
| RDF, SPARQL, RDFS, or SHACL work | Knowledge graph APIs in pkg/cortexdb | Standard graph surface for upsert, find, query, import, export, validation, inference, and explanation. |
| Build a graph from documents, code, or other corpora and analyze it | pkg/graphflow | Runs the fixed pipeline detect -> extract -> build -> analyze -> report -> export. |
| Import external CSV or MySQL/PG SQL dumps into RAG + knowledge graph | pkg/importflow New().Run() / AutoImport() | Parses a source into records, routes columns via a MappingPlan to RAG content and KG entities/relations, with optional AI mapping inference and triple extraction. |
| Typed objects, governed writes, or composable object sets | Ontology APIs in pkg/cortexdb | Activates a schema that validates every write, then exposes actions, object sets, typed tool generation and schema diffing. |
| Expose CortexDB to an agent or MCP client | db.GraphRAGTools() or db.NewMCPServer(); importflow.NewMCPServer() for import tools | Wraps the high-level APIs as tool definitions or an MCP server rather than introducing a second implementation path. |
SaveKnowledge / SearchKnowledge.SaveMemory / SearchMemory.KnowledgeMemory.memoryflow.graphflow.pkg/core + pkg/graphpkg/cortexdbKnowledgeMemory, memoryflow, graphflowGraphRAGTools, MCPmemoryflow.IngestTranscript -> memoryflow.WakeUpLayers -> memoryflow.CloseSessionSaveKnowledge -> SearchKnowledgeSaveMemory -> KnowledgeMemory.RecallUpsertKnowledgeGraph -> QueryKnowledgeGraph -> ExplainKnowledgeGraphInferencegraphflow.Pipeline.RunREADME.md, README_CN.mdpkg/cortexdb/knowledge_api.gopkg/cortexdb/memory_api.gopkg/cortexdb/graphrag.gopkg/cortexdb/brain.gopkg/memoryflow/service.gopkg/graphflow/pipeline.gopkg/importflow/importer.go, pkg/importflow/source_dump.go, pkg/importflow/mcp.gopkg/cortexdb/graphrag_tool_defs.go, pkg/cortexdb/mcp.goimport "github.com/liliang-cn/cortexdb/v2/pkg/cortexdb"db, err := cortexdb.Open(cortexdb.DefaultConfig("KnowledgeMemory.db"))
if err != nil {
return err
}
defer db.Close()
quick := db.Quick()
_, _ = quick.Add(ctx, []float32{0.1, 0.2, 0.9}, "SQLite is a single-file database.")
hits, _ := quick.Search(ctx, []float32{0.1, 0.2, 0.8}, 3)
_ = hitsThe DSN chooses the backend. Nothing else changes:
db, _ := cortexdb.Open(cortexdb.DefaultConfig("/var/lib/cortexdb/brain.db")) // SQLite (default)
db, _ := cortexdb.Open(cortexdb.DefaultConfig("postgres://u:pw@host:5432/cortex")) // PostgreSQL + pgvectorPostgresStore satisfies the same BrainStore contract as SQLiteStore:
vectors, hybrid search, memory, RDF graph, SPARQL, ontology and the tool
surface all run on either. db.Dialect().Kind() reports which.core.RegisterStore / core.OpenBrainStore) is compile-time,
not a plugin system — storage is the hot path, and a process boundary would
add an IPC round trip per search and make transactions impossible.unicode61 → tsvector @@ plainto_tsquery('simple') and the CJK trigram companion → LIKE + pg_trgm
(pkg/core/fts_postgres.go carries the table); pgvector refuses to index past
2000 dimensions, so a wider model falls back to an exact scan and says so;
pgvector itself is optional and a missing extension degrades to the in-Go scan.CORTEXDB_TEST_POSTGRES="postgres://..." go test ./... turns on 104 tests
across pkg/core, pkg/graph, pkg/cortexdb, pkg/agentmem; without it the
run prints 59 skips naming what is uncovered. Most are parity tests — one body,
both databases, same answer required — because the failure mode is silent.cortexdb.QuerySource (Name() + Search()) registered with
WithQuerySource, asked for as QueryPrefetch{Kind: QueryPrefetchSource, Source: "<name>"}. A source returns ids and scores only — content is
hydrated from the brain, ids the brain lacks are dropped, candidates still
pass the filter and the normal fusion, and a source error fails the query
rather than silently shrinking the result set. db.QuerySources() lists the
registered ones. Adapters live outside pkg/ (see
examples/17_query_source, Meilisearch over plain net/http).graph_nodes/graph_edges in the same database and the same transaction as
the vectors, with SPARQL/RDFS/SHACL implemented over that SQL. pkg/sqldialect
is dialect adaptation (placeholders, BLOB→BYTEA, error text, JSON accessors),
not a query builder — it does not reach Cypher. Export with
knowledge_graph_export instead and keep CortexDB as the system of record._, _ = db.SaveKnowledge(ctx, cortexdb.KnowledgeSaveRequest{
KnowledgeID: "apollo-plan",
Title: "Apollo launch plan",
Content: "Alice owns Apollo. Apollo ships on Friday.",
ChunkSize: 24,
Entities: []cortexdb.ToolEntityInput{
{Name: "Alice", Type: "person", ChunkIDs: []string{"chunk:apollo-plan:000"}},
{Name: "Apollo", Type: "project", ChunkIDs: []string{"chunk:apollo-plan:000"}},
},
Relations: []cortexdb.ToolRelationInput{
{From: "Alice", To: "Apollo", Type: "owns"},
},
})
resp, _ := db.SearchKnowledge(ctx, cortexdb.KnowledgeSearchRequest{
Query: "Who owns Apollo?",
Keywords: []string{"Apollo", "Alice", "owns"},
RetrievalMode: cortexdb.RetrievalModeLexical,
TopK: 3,
})
_ = resp.Context
_, _ = db.SaveMemory(ctx, cortexdb.MemorySaveRequest{
MemoryID: "style",
UserID: "user-1",
Scope: cortexdb.MemoryScopeUser,
Namespace: "assistant",
Content: "User prefers concise status updates.",
})No-embedder mode is supported. Use lexical retrieval plus LLM-planned Keywords, AlternateQueries, EntityNames, and RetrievalMode.
Retrieval mode ppr runs Personalized PageRank (HippoRAG 2 style) from the entities a question names, rank-fused with the first stage; tune it with ppr: {fusion, damping, passage_seed_weight, edge_type_weights}. Without an embedder, auto takes that walk whenever a knowledge search names an entity specific enough to start from (an entity more than 10% of passages mention, like a conversation's speaker, is not) and stays lexical otherwise — recall@5 through the public API, lexical → auto: 2WikiMultiHopQA 0.657 → 0.807, MuSiQue 0.457 → 0.542, LoCoMo 0.493 → 0.495, LongMemEval 0.861 → 0.862. SaveKnowledge builds the entity graph without an embedder too, and a title that reads as a name links every chunk of its document. With an embedder auto stays hybrid; pass retrieval_mode: "ppr" for multi-hop questions.
Chinese (and other CJK) questions work in lexical mode as written. A sentence such as 微信机器人为什么从来不主动给我发提醒 is cut at function words (的, 了, 为什么, 那个, …), broken into character bigrams, matched by substring and ranked with BM25, then merged beside the word-index results — for memory, SearchKnowledge, SearchTextOnly and the search_text tool alike, on both backends, and inside the authorization gate. A row is returned only when it matches more than one word of the question (two separate stretches, or one of four characters or more, or the whole of a one-word question) and at least 30% of its content characters, so a question the store knows nothing about still returns nothing.
High-level APIs live in pkg/cortexdb:
UpsertKnowledgeGraphFindKnowledgeGraphDeleteKnowledgeGraphImportKnowledgeGraphExportKnowledgeGraphQueryKnowledgeGraphValidateKnowledgeGraphSHACLRefreshKnowledgeGraphInferenceSummarizeKnowledgeGraphInferenceExplainKnowledgeGraphInferenceExplainKnowledgeGraphInferenceMatch_, _ = db.UpsertKnowledgeGraph(ctx, cortexdb.KnowledgeGraphUpsertRequest{
Triples: []cortexdb.KnowledgeGraphTriple{
{
Subject: graph.NewIRI("https://example.com/alice"),
Predicate: graph.NewIRI(graph.RDFType),
Object: graph.NewIRI("https://example.com/Person"),
},
},
})
result, _ := db.QueryKnowledgeGraph(ctx, cortexdb.KnowledgeGraphQueryRequest{
Query: `SELECT ?o WHERE { <https://example.com/alice> ?p ?o . }`,
})
_ = resultSPARQL is SPARQL 1.1 and 1.2, measured against the W3C test suites: SELECT (DISTINCT, REDUCED, (expr AS ?v)), ASK, CONSTRUCT (GRAPH blocks in templates produce quads), DESCRIBE; FROM / FROM NAMED; update forms (INSERT/DELETE DATA, DELETE WHERE, DELETE…INSERT…WHERE, WITH, USING, USING NAMED, and ADD/COPY/MOVE/CLEAR/DROP/CREATE); GRAPH, OPTIONAL, UNION, MINUS, VALUES, BIND, FILTER, EXISTS, NOT EXISTS, subqueries; every property path (^p, p/q, p|q, p+, p*, p?, !(p|^q), grouped); the SPARQL 1.1 function library (term tests, strings with character positions, numerics, dates, hashes) and XSD casts (xsd:integer(?x)); aggregates with DISTINCT, GROUP BY, HAVING; ORDER BY by value on expressions and aliases. A per-row type error drops the row in FILTER and leaves the variable unbound in BIND. Without FROM the default graph is the unnamed graph plus the property-graph projection. SERVICE answers only through a handler the embedding application sets (GraphStore.SetSPARQLServiceHandler); LOAD is refused — the store never fetches. Results can be written as SPARQL JSON/XML/CSV/TSV (SPARQLResult.WriteResults), and RDF/XML can be imported.
RDF 1.2 and SPARQL 1.2: triple terms <<( s p o )>> (object position only), reified triples << s p o ~ r >>, {| |} annotations and base-directed literals ("x"@ar--rtl) in N-Triples / N-Quads / Turtle / TriG, and SPARQL triple-term patterns with TRIPLE, isTRIPLE, SUBJECT, PREDICATE, OBJECT, LANGDIR, hasLANG, hasLANGDIR, STRLANGDIR. Use them to say things about a fact — its source, its confidence — and query that back:
INSERT DATA { _:r rdf:reifies <<( :alice :worksFor :acme )>> ; :source <doc1> ; :confidence 0.9 }JSON-LD export refuses triple terms rather than flattening them.
Cypher. graph_cypher_query / db.QueryCypher runs a read-only openCypher / GQL subset over the property graph: MATCH, OPTIONAL MATCH, WITH, UNWIND, RETURN, UNION; labels are node types, relationship types are edge types; variable-length paths *m..n capped at 6 hops; aggregates, about 45 functions, $params. n.name falls back to title. Write clauses are refused by design, as are CALL, shortestPath, pattern predicates and temporal functions. Call graph_schema first.
MATCH (p:project)-[:depends_on]->(c {name: 'CortexDB'}) RETURN p.nameThe property graph is readable as RDF. Everything extraction, upsert_entities and upsert_relations write also answers SPARQL, inference and SHACL, as read-only triples in graph <urn:cortexdb:graph:property> — nothing is copied, so nothing goes stale:
| Property graph | Triple |
|---|---|
node X | cxn:X (id percent-encoded exactly: entity:abc → cxn:entity%3Aabc) |
node type T | cxn:X a cxt:T |
edge of type R | cxn:A cxr:R cxn:B |
scalar property k | cxn:X cxp:k "value" (JSON numbers and booleans keep their XSD type) |
name, else title | cxn:X rdfs:label "…" |
SELECT ?who WHERE { ?x cxr:depends_on ?y . ?y rdfs:label "CortexDB" . ?x rdfs:label ?who }Call graph_schema first to learn which types and relations exist. Non-ASCII ids need the full <urn:cortexdb:node:…> form. Deleting or inserting projected triples is refused; change the property graph through its own APIs. GraphStore.SetPropertyGraphProjection(false) turns the projection off; exports leave it out.
Import and export speak N-Triples, N-Quads, Turtle, TriG and JSON-LD 1.1 (jsonld, also accepted as json-ld). JSON-LD import never fetches a remote @context: schema.org's is answered from memory, any other URL is refused with an error naming it, so inline the context instead.
Inference is semi-naive materialization over RDFS (rdfs:subClassOf, rdfs:subPropertyOf, rdfs:domain, rdfs:range) and OWL 2 RL (owl:inverseOf, owl:SymmetricProperty, owl:TransitiveProperty, owl:equivalentClass, owl:equivalentProperty, owl:sameAs, and the class expressions owl:someValuesFrom, owl:allValuesFrom, owl:hasValue, owl:intersectionOf, owl:unionOf, owl:oneOf, max cardinality). Every inferred triple records its rule and supports, so knowledge_graph_infer_explain traces it back to explicit triples. Declared over the projection, the OWL rules fix what extraction gets wrong: cxt:host owl:equivalentClass cxt:Host unifies spellings, cxr:depends_on owl:inverseOf cxr:depended_on_by answers the reverse question, cxn:entity%3Anode_e owl:sameAs cxn:entity%3Asds_e merges one machine stored under two names. A sameAs class larger than MaxSameAsClassSize (default 32) is reported in OversizedSameAsClasses, never half-materialized. OWL 2 RL keys and chains too: owl:FunctionalProperty / owl:InverseFunctionalProperty / owl:hasKey derive sameAs (the same e-mail is the same person), and owl:propertyChainAxiom (an RDF list, length ≥ 2) derives the chain. Contradictions — owl:disjointWith / owl:AllDisjointClasses, owl:propertyDisjointWith / owl:AllDisjointProperties, two different values of a functional property, owl:differentFrom / owl:AllDifferent against a derived sameAs, irreflexive and asymmetric properties, negative property assertions, owl:complementOf, owl:Nothing, a max cardinality of 0 — are reported in inconsistencies / inconsistency_count with the conflicting triple ids, never resolved; a sameAs class contradicted by differentFrom is not materialized.
Inference stays current by itself (on by default; WithAutoInference(false) turns it off). Every committed change is read from the change feed and applied with delete-and-rederive, asynchronously, so writers pay nothing; db.WaitForInference(ctx) waits until your own writes' consequences are in. With no RDFS/OWL axioms declared it is dormant and costs nothing. It never touches triples a SHACL rule produced; re-run the rules after a manual full refresh, which clears them.
Change feed. Every committed write to nodes, edges, triples, memories, knowledge and ontology schemas is appended to change_log in the same transaction — in commit order, exactly once, nothing from a rolled-back transaction. Read it by cursor with changes_since / db.Changes(ctx, after, limit) (resume from the last seq you saw; pruned says the cursor fell behind retention, so rebuild from current state), or db.SubscribeChanges in-process. Retention defaults to 7 days or 500k events.
_, _ = db.QueryKnowledgeGraph(ctx, cortexdb.KnowledgeGraphQueryRequest{
Query: `INSERT DATA { cxr:depends_on owl:inverseOf cxr:depended_on_by }`,
})
_, _ = db.RefreshKnowledgeGraphInference(ctx, cortexdb.KnowledgeGraphInferenceRefreshRequest{})Incremental refresh:
refresh, _ := db.RefreshKnowledgeGraphInference(ctx, cortexdb.KnowledgeGraphInferenceRefreshRequest{
Mode: cortexdb.KnowledgeGraphInferenceRefreshModeIncremental,
Triples: []cortexdb.KnowledgeGraphTriple{
{
Subject: graph.NewIRI("https://example.com/Employee"),
Predicate: graph.NewIRI("http://www.w3.org/2000/01/rdf-schema#subClassOf"),
Object: graph.NewIRI("https://example.com/Person"),
},
},
})
_ = refreshSHACL supports targets sh:targetClass (with subclasses), sh:targetNode, sh:targetSubjectsOf, sh:targetObjectsOf; sh:property with any SHACL property path (sequence, sh:alternativePath, sh:inversePath, sh:zeroOrMorePath, sh:oneOrMorePath, sh:zeroOrOnePath); implicit class targets (a shape that is also an rdfs:Class); sh:class, sh:datatype, sh:nodeKind, sh:minCount, sh:maxCount, sh:min/maxInclusive, sh:min/maxExclusive, sh:minLength, sh:maxLength, sh:pattern + sh:flags, sh:languageIn, sh:uniqueLang, sh:in, sh:hasValue, sh:equals, sh:disjoint, sh:node, sh:not, sh:and, sh:or, sh:xone, sh:closed + sh:ignoredProperties, sh:lessThan, sh:lessThanOrEquals, sh:qualifiedValueShape + sh:qualifiedMinCount / sh:qualifiedMaxCount / sh:qualifiedValueShapesDisjoint, sh:deactivated, sh:severity, sh:message, on node and property shapes alike — all of SHACL Core (the W3C Core suite passes 98/98). SHACL-SPARQL too: sh:sparql SELECT constraints and SPARQL-based constraint components (sh:ConstraintComponent with sh:parameter and ASK/SELECT validators, $PATH), with $this, $value, $currentShape and parameters pre-bound. Results carry the constraint component IRI. Recursive shapes are refused with an error; any result makes conforms false, whatever its severity. Over the projection it is a quality gate for extracted graphs, e.g. every cxr:runs_on must point at an sh:class cxt:host. SHACL-AF rules: knowledge_graph_shacl_rules / db.ApplyKnowledgeGraphSHACLRules runs sh:TripleRule — with sh:condition, sh:order, sh:deactivated and node expressions sh:this, constants, sh:path (incl. sh:inversePath), sh:filterShape, sh:intersection, sh:union — to a fixpoint. Results are explainable inferred triples named shacl_triple_rule:*, and the shapes passed are the whole rule set: a rule left out is retracted on the next run.
report, _ := db.ValidateKnowledgeGraphSHACL(ctx, cortexdb.KnowledgeGraphSHACLValidateRequest{
Shapes: []cortexdb.KnowledgeGraphTriple{
{Subject: graph.NewIRI("https://example.com/PersonShape"), Predicate: graph.NewIRI(graph.RDFType), Object: graph.NewIRI(graph.SHACLNodeShape)},
{Subject: graph.NewIRI("https://example.com/PersonShape"), Predicate: graph.NewIRI(graph.SHACLTargetClass), Object: graph.NewIRI("https://example.com/Person")},
},
})
_ = reportdb.Graph() reports the graph's own vocabulary, so an ontology or drift check
never has to query graph_nodes / graph_edges directly. Raw SQL against those
tables is pinned to one schema and one dialect and fails silently when either
moves; these go through the same dialect-aware path as the rest of the package
and are covered by the SQLite/PostgreSQL parity suite.
nodeTypes, _ := db.Graph().NodeTypeCounts(ctx) // map[string]int, untyped under ""
edgeTypes, _ := db.Graph().EdgeTypeCounts(ctx)
// Which type shapes the edges really have — how a declared relation is wrong.
shapes, _ := db.Graph().EdgeShapes(ctx, "uses")
// []EdgeShape{{EdgeType, FromType, ToType, Count}}
// Which nodes those edges run through — what is wrong.
pairs, _ := db.Graph().EdgeEndpointPairs(ctx, "uses")
// []EdgeEndpointPair{{EdgeType, From, To EdgeEndpoint{ID, NodeType, Content}, Count}}Both edge queries take optional edge types, matched exactly as stored; passing none reports the whole graph.
fact_provenance and the knowledge contract say how a fact is known. The
ledger says why an action was taken: which facts it rested on, who took it,
what it replaced, and what else was decided the same way.
A decision is a graph record, not a side table. It is a node decision:<id>
typed Decision, whose content is the note and whose properties are kind,
actor, at, verdict, subject plus the knowledge contract keys — so
contract_tally counts a decision without being told about decisions, and
expand_graph, render_graph_html, the live view and SPARQL all see it. Three
edge types explain it: based_on (a premise), about (the subject) and
supersedes (the decision it replaces).
rec, err := db.RecordDecision(ctx, cortexdb.DecisionRecordRequest{
Kind: cortexdb.DecisionKindReview, // load | review | action | assert (open vocabulary)
Actor: "liliang",
Note: "Held the ledger-svc release: riskd's rule source is unverified.",
Verdict: "hold",
Subject: "svc:ledger",
Premises: []string{"svc:riskd", "fact:ledger-depends-on-riskd"}, // node ids and edge ids
})
chain, _ := db.DecisionChain(ctx, rec.ID, 0) // premises with each one's grade and source
prior, _ := db.Precedents(ctx, cortexdb.PrecedentsQuery{Kind: "review", Subject: "svc:ledger"})
mine, _ := db.DecisionsBy(ctx, "liliang", 20)
_, _, _ = chain, prior, errIt fails closed and it fails before writing anything: an empty actor, an empty
note, a premise or subject that does not exist, a supersedes target that is not
a decision, or metadata ValidateContract refuses. By default a decision is
stamped _grade=verified, _producer=human, _by=<actor> — a named actor
signed it — and an agent recording its own decision says so with Producer and
Grade.
A premise that is a fact is an edge, and graph_edges declares a foreign
key on to_node_id, enforced on both backends. So the based_on edge is
anchored on the fact's subject node and names the fact in its own
premise_edge_id property; DecisionChain resolves it back and reports the
edge's grade, not the anchor node's.
Tools (in-process and MCP), and the mirroring cortexdb.v1.DecisionService:
decision_record — writes (read-only keys are refused it)decision_chain — readsdecision_precedents — readsAn agent's own record of a task: one AgentRun node per run and one node per
step — LLMCall, ToolCall, Retrieval, DecisionPoint, Validation or a
kind of your own, which is also the node type and Cypher label — joined by
HAS_STEP (run → step), TRIGGERED (a step → a step that consumed its output)
and SPAWNED (a step → a sub-agent's run). Steps need no vector.
run, _ := db.StartRun(ctx, cortexdb.RunStart{Task: "checkout p99 alert", Agent: "devops-agent"})
plan, _ := db.RecordStep(ctx, cortexdb.StepStart{RunID: run.ID, Kind: cortexdb.StepKindLLMCall, Name: "plan"},
cortexdb.StepEnd{Output: "check deps", Tokens: 640, CostUSD: 0.004})
// Two-phase: begin before the work (written as running, with its TRIGGERED
// edges), end after. A crash still leaves what was attempted.
st, _ := db.BeginStep(ctx, cortexdb.StepStart{RunID: run.ID, Kind: cortexdb.StepKindToolCall, Name: "restart_pool", Parents: []string{plan.ID}})
_, _ = db.EndStep(ctx, st.ID, cortexdb.StepEnd{Output: "pool recycled"}) // latency measured from begin
_, _ = db.FinishRun(ctx, run.ID, cortexdb.RunEnd{Outcome: "resolved"})
sum, _ := db.SummarizeRun(ctx, run.ID) // tokens, cost, open/failed steps, critical path, least confident step
up, _ := db.StepLineage(ctx, st.ID, cortexdb.LineageUpstream, 0) // everything the action depended on
then, _ := db.ReplayRun(ctx, run.ID, someInstant) // the run as it stood then, from history
_, _, _ = sum, up, thenEvery write fails before writing anything when the run is not running, a
parent does not exist, the kind cannot be a label, or an attribute reuses a
name the graph writes itself. A step ends once; a run finishes once. Steps are
ordinary graph records, so Cypher, RecordDecision premises, expand_graph
and the change feed all read them. Keep many runs in a file of their own if
step nodes should not appear in a knowledge graph's schema.
Tools (in-process and MCP): execution_run_start, execution_step_begin,
execution_step_end, execution_step_record, execution_run_finish — write;
execution_run_get, execution_runs_list, execution_step_lineage,
execution_run_replay — read.
apply_inference composes two hops. That is one rule shape; the engine behind
it takes any Horn clause over graph edges — any number of premises, variables
bound across them — forward-chained to a bounded fixpoint. apply_inference
keeps its name and behaviour and now runs through that engine, so there is one
derivation path and one explanation for both.
// Written the way a person writes it. Keywords are case-insensitive; a term
// starting with ? is a variable.
_, _ = db.SaveRules(ctx, cortexdb.RulesSaveRequest{Rules: []cortexdb.RuleDefinition{{
ID: "mortality",
Text: "IF instance_of(?x, ?c) AND subclass_of(?c, ?d) THEN instance_of(?x, ?d)",
Confidence: 0.9,
}}})
// Dry run first; with neither rule_ids nor rules, every enabled rule runs.
plan, _ := db.ApplyRules(ctx, cortexdb.RulesApplyRequest{DryRun: true})
applied, _ := db.ApplyRules(ctx, cortexdb.RulesApplyRequest{})
_ = plan
// Why is that edge there?
why, _ := db.ExplainInference(ctx, cortexdb.InferenceExplainRequest{EdgeID: applied.CreatedEdgeIDs[0]})
// why.Explanation names the rule and its text; why.Trace is the premise chain,
// flattened in preorder.Literal terms. A term starting with ? is a variable. Anything else is a
literal, resolved against the stored nodes in this order: an exact node id
first; otherwise Type:Name, matching every node whose node_type equals
Type and whose name (the name property, or the node content) equals Name,
both case-insensitively; otherwise a bare name matched the same way across all
types. A literal matching several nodes matches all of them; one matching none
is returned in unresolved_terms, so a rule that can never fire says so.
Derived edges carry inferred=true, provenance=rule, rule_id,
rule_text, the exact support_edge_ids, and confidence = the minimum
premise confidence times the rule's own. Re-running derives nothing new: an edge
already there from the same rule is reported in unchanged_edge_ids, not
rewritten.
Caps. Chaining stops at 16 rounds and 50 000 derived edges by default
(max_iterations, max_derived). Hitting either is an error —
graph.ErrRuleCapExceeded — and nothing is written, because a half-built
closure in the graph is worse than none.
Rules persist in their own kg_rules table, not as graph nodes: a rule is
configuration for the engine, not a fact about the world the graph describes,
and a rule: node would show up in find_nodes, expand_graph and every
export. The table is created the first time a rule is saved.
Tools: rules_save, rules_list, rules_delete, rules_apply (with
dry_run), inference_explain. Engine: pkg/graph/rules*.go; facade:
pkg/cortexdb/rules*.go.
Retrieval answers "what is relevant to this query". Three APIs answer questions about the collection instead, and all three work identically on SQLite and PostgreSQL.
A threshold instead of a K. RangeSearchVector / RangeSearchText return
every row within a distance of the query, and nothing outside it. Top-K always
returns K rows however weak the last ones are; a range query returns what
clears the bar and nothing when nothing does. That is the honest answer for
"find all the near-duplicates of this", "does the store hold anything like this
at all", and any threshold-driven decision.
near, err := db.RangeSearchText(ctx, "the backup vault is offline",
cortexdb.VectorRangeOptions{Radius: 0.15})Radius is a distance in the store's metric, not a similarity: under the
default cosine it is 1 - similarity, so roughly 0.1 for near-identical text
and 0.5 for loosely related. MaxResults 0 means every match. Tool:
search_vector_range, which caps at 100 by default and reports truncated
when the radius reached further than the cap — a trimmed range answer is
otherwise indistinguishable from a complete one, which would make it an
unlabelled top-K.
Counting. Aggregate runs count, sum, avg, min, max and group_by over a
metadata field, optionally filtered and scoped to a collection. A group_by is
also how to ask what values a field takes and how often. Tool:
aggregate_metadata.
byKind, err := db.Aggregate(ctx, core.AggregationRequest{
Type: core.AggregationGroupBy, GroupBy: []string{"kind"}, OrderBy: "count",
})Metadata names in these requests are column references, not values, and cannot
be bound as parameters — so they are checked against a conservative identifier
shape and refused otherwise, rather than escaped. A name inside a JSON path may
carry dots and hyphens (user.id, content-type); one that becomes bare SQL —
GroupBy, which is aliased by its own text, plus OrderBy and the Having
keys — may not.
Aggregating the vectors. VectorAggregate reduces the vectors themselves:
| Kind | Returns | Use it for |
|---|---|---|
centroid | a computed point | the group's average direction; dragged by outliers |
geometric_median | a computed point | the same idea, robust — it minimises the sum of distances rather than of squared distances, so one far member barely moves it |
medoid | a stored record | "which of these near-duplicates is the canonical one", "give me the one record that represents this cluster" |
The medoid is the one worth reaching for: it names a row that exists, with its
text, so the answer can be read and traced rather than only measured. It
answers the consolidation question arithmetically, with no model in the loop.
Tool: representative_records — medoid only, because a synthetic point can only
reach a model as several hundred floats it can do nothing with.
Facets. SearchWithFacets filters a vector search by typed, nestable
metadata facets and can return the per-value counts beside the hits. It is
program-only on purpose: the filter language is a lot of schema for a model to
get right, and the facet question a model actually has — what values, how many
— is aggregate_metadata with a group_by.
Retrieval says what is relevant to a query; these say what the graph is. All work on both backends.
Its observed schema. GraphSchema reports which node and edge types exist
and how many of each, which node-type pairs every edge type actually connects,
and which property keys each node type carries. GraphPropertyValues reports
what values one key actually takes. Call both before writing a traversal or a
SPARQL query against a graph you have not inspected: a query naming a type the
graph does not have still runs, matches nothing, and reads as an answer — and
stored values are routinely codes, so a filter written from a key's name
(color == "black" against rows that say BLK) returns nothing and looks like
a fact. ontology_get answers the declared schema, which is opt-in and
usually absent on graphs built by extraction; this is measured from the rows.
The response carries a rendered text field, because this is read in a prompt.
Tools: graph_schema, graph_property_values.
Its shape and its centres. GraphPageRank ranks nodes by structural
importance — what a knowledge base is about, rather than what it says most
often — with each node's label and type beside the score. GraphStatistics
reports node and edge counts, average degree, density and connected components;
a component count far above 1 means entities were written and never linked.
GraphPredictEdges lists nodes one node is not connected to but arguably should
be, which is two findings in one shape: a missing fact, or one entity stored
twice. Read tied_at_top first — several unrelated nodes tied at exactly 1.0
means this graph's vectors cannot tell them apart, which is what short lexical
hashes do to short entity names. Tools: rank_graph_nodes, graph_statistics,
predict_graph_edges.
rank_graph_nodes reads a cached ranking. Scores live in a side table
(graph_pagerank_scores) with the time they were computed; a call on an
unchanged graph is an indexed LIMIT read (on a 2000-node graph: 0.11ms p50
against 16ms recomputing), and a change to the graph's nodes or edges is
detected in the database — on SQLite by rowid maxima plus triggers on deletes
and endpoint rewrites, so writes from another process count; on PostgreSQL by
a topology aggregate — and triggers a recompute. The answer carries cached,
computed_at, stale, stale_reason; refresh forces a recompute,
allow_stale serves an out-of-date ranking without paying for one,
max_age_seconds bounds its age. From Go: GraphPageRankOptions.Cache
(PageRankCacheAuto / Refresh / AllowStale; the zero value computes fresh
as before), RefreshPageRankCache, InvalidatePageRankCache, and
StartPageRankRefresher(ctx, interval, …) to keep it warm on a timer.
Whole-corpus questions (GraphRAG global search). global_search answers
what retrieval cannot — "what are the main themes", "what does this brain
cover" — by map-reducing over community reports. The caller chooses it; no
query is routed to it by its wording. build_community_hierarchy detects the
entity graph's communities with Leiden — every community connected at every level, which Louvain does not guarantee; HierarchyOptions.Algorithm selects Louvain — and keeps every level (0 = finest,
each level up merges the one below), then writes a report per community
bottom-up: level 0 from its entities and relations, higher levels from their
children's reports. global_search takes a level (default: the coarsest,
fewest reports, cheapest map) and builds the hierarchy on first use. With a
model (graphflow.JSONGenerator; in the MCP binary CORTEXDB_LLM_*) reports
and answer are the model's; without one reports are assembled
deterministically and global_search returns the lexically most relevant
reports unsynthesised (mode: "no_model"). Go: graphflow.BuildCommunityHierarchy,
graphflow.GlobalSearch(…, GlobalSearchOptions{Level: &l}),
graph.GraphStore.HierarchicalCommunities. Registered by
graphflow.AddGlobalSearchMCPTools in cortexdb-mcp-stdio (local mode) and in
graphflow.NewMCPServer.
Disambiguating a name against the graph. DisambiguateMentions resolves an
ambiguous name using the other names given with it: it finds the shortest paths
between this mention's candidates and the others', renders each as a sentence,
and ranks by that support. Steps one to four are deterministic and live here;
the choice is the caller's, so a caller with no model takes the top rank and one
with a model reads the evidence — and either way can say afterwards why. A
mention nothing connects to is unresolved, never the nearest string match, and
a bound that stopped the search is reported apart from an absence.
res, err := db.DisambiguateMentions(ctx, []string{"Claude", "MCP", "OpenClaw"},
cortexdb.DisambiguationOptions{NodeTypes: []string{"entity", "project", "protocol", "tool"}})Pass NodeTypes on any brain that keeps documents and chunks as graph nodes.
Without it, one document title holding two of the mentions matches both by
substring and scores a perfect 1.0 for each — a coincidence of two words that
ties with the real entity and resolves nothing. Tool: disambiguate_mentions.
Resolving duplicate entities: link first, merge later.
graphflow.ResolveEntities scores each candidate pair on name similarity,
shared neighbours and type agreement (plus the model's grouping, if an LLM is
given). Only a score of 0.95 or more merges; 0.85–0.95 writes a possiblySame
edge carrying that evidence, graded held, which contract_needs_attention
lists. Grade the edge verified and the next pass merges it; refused keeps
the pair apart. Spelling alone never merges two differently spelled names, and
two entities with the same name and no neighbour in common (two main.go
files, Mercury the planet and the element) are linked, not merged.
ResolveOptions{Mode: graphflow.ResolveMergeAll} is the old unscored merge.
Why a fact ended. Every row that leaves the live graph — superseded,
retracted, merged, removed with its document — lands in
graph_node_history/graph_edge_history with reason, superseded_by and
producer. Read it with db.NodeHistory / db.EdgeHistory, or from the
invalidation on each retracted or changed row of graph_diff. A caller that
knows more than the default says so with graph.WithInvalidation(ctx, …).
Path retrieval for multi-hop questions. Chunk search ranks passages by how
much they look like the question; for "how is Alice Chen connected to Kestrel?"
the passage stating the middle link ("Borealis Labs built Kestrel") does not
look like it, and one that merely shares its words does. SearchPaths walks
the graph between the entities the question names — bounded breadth-first
search, at most MaxDepth edges (default 3), scored decay^(hops-1) times the
product of relation weights — and returns each chain as edges that cite the
chunk they were extracted from (chunk_ids / source_chunk_id on the edge).
With one seed it returns the paths leaving it, which is the shape of "the city
of the company Alice works for". Paths never run through chunk or document
nodes unless IncludeBookkeeping is set: a shared chunk is a co-mention, not a
relation.
resp, err := db.SearchPaths(ctx, cortexdb.ToolSearchPathsRequest{
EntityNames: []string{"Alice Chen", "Kestrel"},
RelationPolicies: graph.RelationPolicies{"co_occurs_with": {Weight: 0.2, MaxDepth: 1}},
WithText: true,
})
// resp.Paths[i].Edges[j].ChunkIDs — the evidence, link by link; resp.ChunkIDs — all of it, in path orderGraphRAGQueryOptions.ReturnPaths (return_paths on search_graphrag_lexical)
adds the same chains to a GraphRAG result, seeded by the plan's entity names or,
without any, by the entities the top three chunks mention; chunks and context
are unchanged. Tool: search_paths.
Per-relation-type weight and depth. graph.RelationPolicies maps a relation
type to {Weight, MaxDepth} — weight multiplies a path's score (0 = 1.0,
negative = never follow), max depth is the farthest hop from the start at which
the type is still followed — with "*" for every unlisted type. It is accepted
by SearchPaths, expand_graph (relation_policies; the response then carries
a score per node and limit keeps the best), Neighbors/ScoredNeighbors
(TraversalOptions.Relations) and HybridSearch (GraphFilter.Relations,
where GraphScore becomes the relation-weight product over distance+1).
Unset, every one of them behaves exactly as before.
GraphRAGQueryOptions.ChunkWindow (and chunk_window on the search tools)
returns each hit's neighbouring chunks in the same document alongside it. Off by
default, and the neighbours are context rather than hits: never scored, never
reranked, never displacing a match, and tagged +context in the assembled text
so a reader can tell what matched from what came along.
It is for the failure chunking guarantees. A real retrieval over a book chapter
returned a hit whose first word was ance. — importance, cut at a chunk
boundary — and the question, about two steps of a process, was only answerable
at ChunkWindow: 2, when both steps were finally in the context. Overlapping
windows merge into one span, and the existing MaxContextChars budget is
respected: hits are charged first and never refused.
pkg/cortexdb models a Palantir-style ontology on the same file. One schema is
active at a time and validates every write.
_, err := db.SaveOntologySchema(ctx, cortexdb.OntologySaveRequest{
Schema: cortexdb.OntologySchema{
SchemaID: "aviation",
InterfaceTypes: []cortexdb.OntologyInterfaceType{{APIName: "Facility"}},
ObjectTypes: []cortexdb.OntologyObjectType{{
APIName: "Airport",
PrimaryKey: "iataCode", // mandatory
TitleProperty: "facilityName",
Implements: []string{"Facility"},
Properties: []cortexdb.OntologyProperty{
{APIName: "iataCode", DataType: cortexdb.OntologyDataType{Kind: cortexdb.OntologyDataString}, Required: true},
{APIName: "facilityName", DataType: cortexdb.OntologyDataType{Kind: cortexdb.OntologyDataString}, Required: true, Searchable: true},
},
}},
LinkTypes: []cortexdb.OntologyLinkType{{
APIName: "flightDeparture",
A: cortexdb.OntologyLinkSide{APIName: "departures", ObjectTypeAPIName: "Airport", Cardinality: cortexdb.OntologyCardinalityMany},
B: cortexdb.OntologyLinkSide{APIName: "origin", ObjectTypeAPIName: "Flight", Cardinality: cortexdb.OntologyCardinalityOne, ForeignKeyProperty: "originIata"},
}},
},
Activate: true,
})Key points:
primary_key. Node identity becomes entity:<objectType>:<primaryKey>; with no active schema the older name-derived IDs still apply.ONE / MANY); only a ONE side may carry a foreign_key_property.Facility returns every implementor. Multiple inheritance is allowed; cycles are rejected.shared_properties are declared once and reused by name; storage keeps them unexpanded, so read paths that care about types expand them first.Object sets compose retrieval (db.ResolveObjectSetObjects): kinds base,
interface_base, static, reference, filter, search_around, union,
intersect, subtract; predicates eq/lt/lte/gt/gte/in/is_null/
contains/starts_with/contains_all_terms/contains_any_term/
nearest_neighbors plus and/or/not. Vector, full-text and graph
traversal are peers in one expression. At most 3 chained search_around hops.
Action types are governed, audited writes (db.ApplyAction): typed parameters,
edit rules, submission criteria, validate_only / return_edits (mutually
exclusive). strict_actions: true on the schema closes the generic upsert
tools so actions are the only write path.
Two schema-lifecycle helpers:
tools, _ := db.GenerateOntologyTools(ctx, cortexdb.OntologyToolGenOptions{IncludeObjectTypes: true})
diff, _ := db.DiffOntologySchema(ctx, cortexdb.OntologyDiffRequest{SchemaID: "aviation", Candidate: candidate})GenerateOntologyTools emits typed tool definitions from the active schema —
one per action type, optionally one list tool per object type. The result is
capped (32 by default) and is not registered with NewMCPServer: every
extra tool is paid for in the agent's context window on every request, so
exposing them is an explicit caller decision.
DiffOntologySchema reports what a candidate would change and flags the
breaking classes: removed object/link type, removed property, changed property
data type, property became required, new required property, changed primary
key, retargeted link side, cardinality tightened MANY -> ONE.
Writes MERGE: an upsert updates the properties it names and leaves the rest
alone, so naming an object a document merely mentions no longer erases what a
fuller write established. There is no way to remove a property by omitting it;
use delete_entities.
The extractor's untyped output is bookkeeping and is exempt from schema
validation, so an active strict ontology no longer blocks SaveKnowledge with
an embedder. A schema that itself declares an object type named entity keeps
its own validation. The exemption keys on the type name alone, so a CALLER-
supplied extractor can write a bare name-keyed node under strict enforcement;
the upsert_entities and SaveKnowledge entity doors are unaffected.
A property declared vectorized is embedded on write when an embedder is
configured: the object's node vector becomes the embedding of its vectorized
text (name first, then each vectorized property), so the nearest_neighbors
object-set predicate answers by meaning. With no embedder the node keeps the
lexical hash vector and the write still goes through. A search hit names the
objects its chunks mention whatever their object type, in auto mode as well as
graph mode; an explicit lexical mode or disable_graph still leaves the graph
untouched. Hybrid results carry each retriever's score and rank beside the
fused score, which is a reciprocal-rank value that is right to order by and
wrong to read.
Known limitation: modify_object does not rewrite a node's stored display
title. Foundry's
function runtime, branches/proposals, dynamic row-level security and backing
datasources are deliberately not modelled.
Use pkg/memoryflow for agent memory workflows:
flow, _ := memoryflow.New(db, planner, extractor)
_, _ = flow.IngestTranscript(ctx, memoryflow.IngestTranscriptRequest{
Transcript: memoryflow.Transcript{
SessionID: "session-1",
UserID: "user-1",
Source: "chat",
Turns: []memoryflow.TranscriptTurn{
{Role: "user", Content: "Apollo ships on Friday."},
{Role: "assistant", Content: "Captured."},
},
},
Scope: cortexdb.MemoryScopeSession,
Namespace: "assistant",
})
layers, _ := flow.WakeUpLayers(ctx, memoryflow.WakeUpLayersRequest{
Identity: "You are the Apollo project assistant.",
Recall: memoryflow.RecallRequest{
Query: "startup context",
SessionID: "session-1",
Scope: cortexdb.MemoryScopeSession,
Namespace: "assistant",
},
})
_ = layersLLM-dependent interfaces:
QueryPlannerSessionExtractorPromotionPolicyOptional Hindsight recall strategy plugin:
flow, _ := memoryflow.New(
db,
planner,
extractor,
memoryflow.WithRecallStrategy(hindsight.NewStrategy(db, hindsight.StrategyOptions{
BankID: "apollo-agent",
EntityNames: []string{"Apollo"},
Keywords: []string{"deadline"},
UseKG: true,
})),
)Use pkg/graphflow for corpus-to-graph workflows:
extraction := graphflow.ExtractionResult{ /* nodes + edges */ }
_, _ = graphflow.Build(ctx, db, []graphflow.ExtractionResult{extraction}, graphflow.BuildOptions{})
analysis, _ := graphflow.Analyze(ctx, db, graphflow.AnalyzeRequest{TopN: 10})
report, _ := graphflow.RenderReport(ctx, analysis)
_, _ = graphflow.Export(ctx, db, graphflow.ExportRequest{OutputDir: "graphflow-out", Analysis: analysis, Report: report})
_, _ = graphflow.ExportHTML(ctx, db, graphflow.ExportRequest{OutputDir: "graphflow-out", Analysis: analysis})LLM extraction uses only this interface:
type JSONGenerator interface {
GenerateJSON(ctx context.Context, systemPrompt string, userPrompt string) ([]byte, error)
}The example examples/05_graphflow uses github.com/openai/openai-go/v3 with JSON Schema structured output:
OPENAI_API_KEY=...
OPENAI_BASE_URL=http://43.167.167.6:8080/v1
OPENAI_MODEL=gpt-5.4Use pkg/importflow to import external structured data (CSV, MySQL/PostgreSQL dumps)
into RAG + knowledge-graph foundations in one pass:
src, _ := importflow.NewSQLDumpSource(strings.NewReader(dump), importflow.DumpOptions{Dialect: importflow.DialectMySQL})
defer src.Close()
// or importflow.NewCSVSource(reader, importflow.CSVOptions{Table: "people"})
im := importflow.New(db)
plan := importflow.MappingPlan{Tables: map[string]importflow.TablePlan{
"people": {
RAG: &importflow.RAGPlan{ContentTmpl: "{name}: {bio}", IDColumn: "id"},
KG: &importflow.KGPlan{Entities: []importflow.EntityMap{{Ref: "p", Type: "Person", IDTmpl: "{id}", LabelTmpl: "{name}"}}},
},
}}
rep, _ := im.Run(ctx, src, plan) // rep.RowsRead/ChunksIndexed/TriplesCreated/UnparsedStatementsAI is optional and injected via interfaces — WithMappingInferer (infers the plan from
the schema), WithTextRefiner, WithExtractor (graphflow triple extraction). With an
inferer, im.AutoImport(ctx, src, importflow.Goal{BuildRAG: true, BuildKG: true}) does
plan + run in one step. RAG works with no embedder (lexical FTS5). The mapping inferer
reuses the same graphflow.JSONGenerator interface.
importflow.MappingFromDDL(ddl, importflow.DDLMappingOptions{}) turns a CREATE TABLE
DDL subset (Postgres/MySQL) into a reviewable MappingPlan deterministically, no LLM —
tables→entity classes, PK→entity id, FK→relations, columns→RAG content + props (no
ALTER/views). importflow.ParseDDL returns the raw parsed tables. Review the plan, then
feed it to im.Run / the connector to build the graph.
importflow.MappingFromDDLWithLLM(ctx, ddl, gen, opts) is the LLM-enhanced sibling:
it computes that deterministic baseline, then has the graphflow.JSONGenerator refine it
(semantic relation names, implicit relations, free-text→text_extract, junction-table
collapse), with graceful fallback to the baseline on any LLM error and a deterministic
executability guard (a refined table is adopted only when its relations bind to declared
entities; dropped tables are unioned back from the baseline). It returns
(plan, tables, llmUsed, err).
pkg/connector is a privacy gate in front of ImportFlow: connect to a live
PostgreSQL/MySQL DB (or wrap any importflow.Source), introspect schema, classify PII
(rule-based + optional LLM), and apply a human-signed MaskingPlan before any bulk
data moves (schema-first, data-second):
src, _ := connector.NewPostgresSource(dsn, connector.SourceOptions{}) // or NewCSVSource, NewMySQLSource, ...
plan, _ := connector.BuildMaskingPlan(ctx, src, connector.NewRuleClassifier(), connector.PlanOptions{ScanTextColumns: true})
plan.Sign("you", time.Now()) // Run refuses an unsigned plan
vault, _ := connector.OpenSQLiteVault("tenant.vault.db")
d, _ := connector.NewDesensitizer(plan, connector.DesensitizerOptions{Tenant: "acme", KeyProvider: kp, Vault: vault})
rep, _ := importflow.New(db).Run(ctx, connector.Desensitized(src, d), mapping)Default-deny: a column not covered by the signed plan is dropped, never silently passed
through (the desensitizer fails closed). Per-column actions are
drop/redact/mask/generalize/hash/pseudonymize, plus in-place free-text PII
redaction. Pseudonyms are reversible via a separate AES-256-GCM token vault keyed per
tenant; the RAG/LLM path never reads the vault and connector.Unmask is the only reverse
path. Quasi-identifier re-identification risk is reduced, not zero.
The desensitizer is an importflow.Source decorator, so the same masked records feed
both RAG and the knowledge graph — pseudonymized join keys become deterministic tokens,
so KG entity IRIs and bought/etc. edges survive while the original PII never enters the
graph. See examples/09_connector (RAG) and the e2e_test.go KG case.
Agent surface: connector.NewToolbox(db, ToolboxOptions{Vault, KeyProvider, Tenant}) →
connector.NewMCPServer(tb, opts) / connector.RunMCPStdio(ctx, tb, opts), or
connector.RegisterMCPTools(server, tb) to ride an existing MCP server (e.g. importflow's).
The cmd/cortexdb-connector-mcp binary runs the four tools over stdio (config via
CORTEXDB_PATH, CONNECTOR_VAULT_PATH, CONNECTOR_TENANT, CONNECTOR_KEY_FILE).
A Watcher keeps a knowledge base continuously in sync with a live DB through
the same privacy gate — it consumes row-level ChangeEvents and applies
idempotent upserts (and hard deletes) to RAG + KG. Three change sources feed
it; true CDC (hard-delete capture, continuous streaming) is available for both
Postgres and MySQL.
Route A (polling, NewPollingChangeSource): polls
WHERE <cursor> > <watermark> per table on demand, is DB-agnostic
(PG/MySQL/Neon), resumes from a checkpoint in the knowledge DB. Library API
only — w.Run(ctx) does one pass and the caller schedules it (loop / cron).
Needs a monotonic cursor column (e.g. updated_at) and cannot see hard
deletes.
src, _ := connector.NewPollingChangeSource("postgres", dsn, connector.PollingOptions{
Tables: []connector.TableCursor{{Table: "orders", CursorColumn: "updated_at", KeyColumns: []string{"id"}}},
})
w, _ := connector.NewWatcher(db, src, connector.WatcherOptions{
SourceKey: "orders", Desensitizer: d, Checkpoint: cp,
Mapping: importflow.MappingPlan{Tables: map[string]importflow.TablePlan{
"orders": {RAG: &importflow.RAGPlan{ContentTmpl: "{name}", IDColumn: "id"}},
}},
})
_ = w.Run(ctx) // schedule on an interval; resumes from the checkpointRoute B-PG (Postgres logical replication, NewPostgresCDCSource): a true
CDC source over pgoutput that captures hard deletes and streams
continuously — w.Run(ctx) blocks until ctx is cancelled and resumes by LSN
from the checkpoint (Checkpoint.Position). Prerequisites: wal_level=logical,
a publication (CREATE PUBLICATION cdc_pub FOR TABLE ...), and a replication
slot (auto-created); default REPLICA IDENTITY (PK) carries the delete key.
src, _ := connector.NewPostgresCDCSource(dsn, connector.PostgresCDCOptions{
Publication: "cdc_pub", Slot: "cdc_slot",
Tables: map[string][]string{"orders": {"id"}},
})
w, _ := connector.NewWatcher(db, src, connector.WatcherOptions{
SourceKey: "orders-cdc", Desensitizer: d, Checkpoint: cp,
Mapping: importflow.MappingPlan{Tables: map[string]importflow.TablePlan{
"orders": {RAG: &importflow.RAGPlan{ContentTmpl: "{name}", IDColumn: "id"}},
}},
})
go w.Run(ctx) // streams insert/update/delete continuously; resumes by LSNRoute B-MySQL (MySQL binlog, NewMySQLBinlogSource): a true CDC source
reading the ROW-format binlog (via go-mysql canal). Like Route B-PG it
captures hard deletes and streams continuously — w.Run(ctx) blocks
until ctx is cancelled and resumes by binlog position (Checkpoint.Position).
Prerequisites: MySQL binlog_format=ROW, binlog_row_image=FULL (defaults in
MySQL 8), a user with REPLICATION SLAVE, REPLICATION CLIENT, and a unique
ServerID.
src, _ := connector.NewMySQLBinlogSource(dsn, connector.MySQLBinlogOptions{
ServerID: 1101, Tables: map[string][]string{"orders": {"id"}},
})
w, _ := connector.NewWatcher(db, src, connector.WatcherOptions{
SourceKey: "orders-binlog", Desensitizer: d, Checkpoint: cp,
Mapping: importflow.MappingPlan{Tables: map[string]importflow.TablePlan{
"orders": {RAG: &importflow.RAGPlan{ContentTmpl: "{name}", IDColumn: "id"}},
}},
})
go w.Run(ctx) // streams insert/update/delete; resumes by binlog positionPrecondition (all routes): every RAG table's MappingPlan must key on the
primary key (RAGPlan.IDColumn) so updates/deletes hit the right chunk —
NewWatcher errors otherwise. Privacy is unchanged: every streamed row still
passes the signed MaskingPlan, raw PII never enters RAG/KG, and pseudonymized
keys become the same deterministic tokens so KG edges survive.
In-process tool calls:
tools := db.GraphRAGTools()
defs := tools.Definitions()
resp, err := tools.Call(ctx, "knowledge_graph_query", payload)
_, _, _ = defs, resp, errMCP server:
server := db.NewMCPServer(cortexdb.MCPServerOptions{})
_ = serverImportant tools:
ingest_document, search_text, expand_graph, build_contextknowledge_save, knowledge_search, memory_save, memory_searchknowledge_graph_upsert, knowledge_graph_query, knowledge_graph_shacl_validate, knowledge_graph_infer_refreshknowledge_memory_recall, knowledge_memory_build_context_pack, knowledge_memory_reflect, knowledge_memory_consolidateontology_save, ontology_get, ontology_list, ontology_delete, ontology_diff, ontology_action_list, ontology_action_apply, object_set_resolveapply_inference, rules_save, rules_list, rules_delete, rules_apply, inference_explaindecision_record, decision_chain, decision_precedentsexecution_run_start, execution_step_begin, execution_step_end, execution_step_record, execution_run_finish, execution_run_get, execution_runs_list, execution_step_lineage, execution_run_replayaggregate_metadata, representative_records, search_vector_rangegraph_schema, graph_property_values, graph_statistics, graph_healthverify_claims, fact_provenance, uncited_factsrank_graph_nodes, predict_graph_edgesdisambiguate_mentionssearch_paths (and return_paths on search_graphrag_lexical, relation_policies on expand_graph)Separate workflow toolboxes:
memoryflow_ingest_transcript, memoryflow_recall, memoryflow_wake_up_layers, memoryflow_prepare_replygraphflow_build, graphflow_analyze, graphflow_report, graphflow_export, graphflow_run, global_search, build_community_hierarchyBy default every agent opens its own SQLite file. Point them at one central
cortexdb-grpc instead and Claude Code, Codex, OpenClaw and agents in other VMs
read and write the same memory and knowledge graph.
On the host that owns the database:
CORTEXDB_PATH=$HOME/.cortexdb/cortexdb.db \
CORTEXDB_GRPC_ADDR=10.0.0.5:47821 \
CORTEXDB_GRPC_TOKEN=<token> cortexdb-grpcOn every client:
export CORTEXDB_REMOTE="10.0.0.5:47821"
export CORTEXDB_GRPC_TOKEN="<the same token>"That is the whole change. The MCP server then opens no local database: it
discovers the tool surface from the server at startup and proxies every call, so
all tools — current and future — work identically. The UserPromptSubmit
auto-recall hook follows the same remote, so injected memories come from the
same brain the tools write to. So do --memory-html, --export-memory,
--memory-usage and --sync-memory: all four route through one shared
remote-fetch path that walks the brain a memory_list_all page at a time
rather than asking for everything in one call, so none of them stop working
once a brain no longer fits in one gRPC message.
Transport is plaintext by design — run it over loopback, a trusted LAN, or
Tailscale. The token is the access control: anyone holding it has full
read/write access. Embedder and LLM settings live on the server, not the
clients. --graph-html, --recall-eval and --capture-session follow the
remote too. Some one-shot modes are local-only by design — --learn-path,
--reembed-memories and --graph-cleanup among them; the latter two need
direct database access, so they run where the database lives.
The graph view is also an MCP tool, render_graph_html. It is the one tool that
is not proxied to the shared brain: the graph is read remotely, but the HTML
is rendered and written where the MCP server runs, because the caller needs the
file on its own filesystem to open or attach it — a server-side render would
land it on the brain's host, out of reach of whatever asked. Set
CORTEXDB_VIEW_DIR to choose where renders go.
serve_graph_3d serves the live 3D view, and the view can be asked, not only
looked at. Its Explore bar finds a node by name anywhere in the store (not only
in the drawn core), asks the brain a question and lights up what the answer
names, and runs read-only Cypher; any node can be expanded to its neighbours,
including ones outside the core, and the inspector works on a shared brain too.
Everything it runs is on a fixed list of read-only tools (liveview.ExploreTools,
held against the catalogue's Mutates by a test); SPARQL is not on it, because
knowledge_graph_query also runs updates. A view opens wherever its link says —
?focus=ID|name&hops=N, ?find=, ?ask=, ?cypher=, ?q=, ?type=,
?edge=, ?explore=0, ?theme=space|ember|mono, ?mode=light|dark|auto — and embedders can drive it with postMessage
(cortexdb:focus, cortexdb:find, cortexdb:ask, cortexdb:cypher). On a
phone the inspector is a bottom sheet and the other panels start folded. The same server draws the brain as a library (library_url in the result, or view=library): each memory is a book on the shelf of the project it names, shelves group into wings by their projects' relations, an entity is an index card listing every book that mentions it, and a book opens on a lectern with its text, index and see-also.
import_agent_memory brings in the memory a person already has, so a new brain does not start empty. By default it imports Claude Code's memory notes (~/.claude/projects/*/memory/*.md), CLAUDE.md / AGENTS.md, and Codex's own memories (~/.codex/memories_1.sqlite and the notes in ~/.codex/memories). With sources: ["claude_sessions", "codex_sessions"] it also distils past transcripts into memories, the way a session is captured when it ends. Each transcript is a model call (CORTEXDB_LLM_*), so one call does max_sessions (default 5) and says how many are left; since limits the window (30d, 2026-09-01), and ~/.cortexdb/imported-sessions.json keeps a transcript from being read twice. Like render_graph_html it runs where the MCP server runs, because that is where the files are, and writes to whichever brain the server uses, local or shared. dry_run counts first; ids are stable, so running it again refreshes rather than duplicates. From a shell: cortexdb-mcp --import-agent-memory [--sessions] [--since 30d] [--max N] [--dry-run].
A capture also retires what it made untrue. Each fact a session yields is looked up in the brain, and one more model call is asked which existing memories it contradicts — the same thing with a different current value, not merely the same topic. Those are saved as superseded by the new memory (supersedes): kept, exported, stamped superseded_by / superseded_at, and no longer recalled as current, so a wrong call is undone by clearing one key. The later date wins: a fact from a session imported months late does not retire newer memories, and is dropped if one already contradicts it. At most five memories are retired per pass, and a failed check saves everything and retires nothing. The SessionEnd capture, the cortexdb-live mod and import_agent_memory's session import all do this.
Both bulk listings behind these views are paged: memory_list_all (default
limit 500) and graph_list_all (default limit 2000) each take limit and
cursor. memory_list_all always pairs truncated with next_cursor —
never one without the other — so a caller resumes exactly where it left off
instead of restarting the walk. graph_list_all has two modes, and only one of
them makes that same pairing: with no cursor and no order it keeps the
most-connected core, which is what makes a large graph renderable, and sets
truncated alone — degree ranking has no stable page boundary to resume from,
so there is no next_cursor to give. order: "id" — or simply supplying a
cursor — switches to a stable, resumable walk of the whole graph instead,
where truncated and next_cursor pair the same way memory_list_all's do; a
page's edges may reference nodes that land on a later page, so the subgraph is
complete only once the walk finishes. Cursors are opaque and tagged with the
listing that produced them, so handing a memory_list_all cursor to
graph_list_all, or the reverse, fails loudly rather than silently returning
the wrong page.
Native host-memory adapters live under plugins/:
plugins/openclaw-cortexdb-memory registers OpenClaw's exclusive memory
capability and CortexDB recall/store/delete tools.plugins/hermes-cortexdb-memory registers Hermes Agent's MemoryProvider
with automatic prefetch and completed-turn synchronization.Both adapters call the existing gRPC ToolsService and prefer
knowledge_memory_recall; do not implement a separate retrieval or storage path
inside an agent plugin. The skills/cortexdb-memory-* directories remain the
lighter explicit-tool integration.
importflow_plan, importflow_run, importflow_ddl_plan (deterministic CREATE TABLE DDL → reviewable knowledge-graph MappingPlan, no LLM), importflow_ddl_plan_ai (LLM-refined DDL→KG plan; returns refined plan + deterministic baseline + llm_used; needs an LLMInferer) (via importflow.NewToolbox(im) or importflow.NewMCPServer(im, opts) / importflow.RunMCPStdio(ctx, im, opts))connector_introspect, connector_plan, connector_run, connector_unmask — privacy gate over a live DB/source feeding RAG + KG. Wire via connector.NewMCPServer(tb, opts) / connector.RunMCPStdio(ctx, tb, opts), connector.RegisterMCPTools(server, tb) to ride another server, or run cmd/cortexdb-connector-mcp.The full facade is also served over gRPC for non-Go callers:
cmd/cortexdb-grpc (env/flags: CORTEXDB_PATH,
CORTEXDB_GRPC_ADDR default 127.0.0.1:47821, CORTEXDB_GRPC_TOKEN enables
bearer auth, OPENAI_BASE_URL/OPENAI_API_KEY/CORTEXDB_EMBED_MODEL/
CORTEXDB_EMBED_DIM wire an OpenAI-compatible embedder; unset → lexical mode).
Flags override the environment, which overrides the defaults. -health
probes a server already running at -addr and exits non-zero when it is not
serving, so the binary is its own liveness check; a wildcard listen address is
probed on loopback. Running it as a service: deploy/ has the systemd unit,
the container image (whose HEALTHCHECK is cortexdb-grpc -health) and the
compose file, plus backup/upgrade notes.proto/cortexdb/v1/ — services KnowledgeService,
MemoryService, KnowledgeGraphService, GraphRagService, ToolsService
(generic CallTool(name, json) over the same toolbox as MCP), AdminService.pkg/rpc/v1, conversion-only handlers in
pkg/rpcserver (no business logic there). Regenerate via scripts/gen-proto.sh.cortexdb-client, mirroring the same
knowledge/memory/graph/graphrag/tools/admin sub-clients + bearer token):clients/rust/ (crates.io): cargo add cortexdb-client;
CortexClient::builder(endpoint).token(t).connect(); optional
managed-server feature downloads/spawns the sidecar with an auto token
(Sidecar::ensure().spawn(db)). Generated code committed (src/gen/).clients/python/ (PyPI): pip install cortexdb-client;
CortexClient.connect(endpoint, token=t). Target for Python agents (Hermes).clients/node/ (npm): npm install cortexdb-client;
CortexClient.connect(endpoint, { token }), promise-based. Target for Node
agents (OpenClaw); protos loaded at runtime, no build step.skills/cortexdb-memory-hermes (Python) and
skills/cortexdb-memory-openclaw (Node) are agentskills.io-format skills that
wire CortexDB in as agent memory + KG. Published to ClawHub
(clawhub skill install cortexdb-agent-memory) and installable from git /
skills.sh (hermes skills install git:liliang-cn/cortexdb@main,
npx skills add liliang-cn/cortexdb).pkg/semantic-router is optional. Use it before CortexDB tools when you need intent routing.
No-embedder lexical router:
router, _ := semanticrouter.NewLexicalRouter(semanticrouter.WithSparseThreshold(0.1))
_ = router.Add(&semanticrouter.SparseRoute{Name: "memory_save", Utterances: []string{"remember this", "save to memory"}})
route, _ := router.Route(ctx, "please remember this")
_ = route.RouteNameThe examples are architecture-oriented:
go run ./examples/01_core
go run ./examples/02_rag
go run ./examples/03_memoryflow
go run ./examples/04_knowledge_graph
go run ./examples/05_graphflow
go run ./examples/06_tools_mcp
go run ./examples/07_importflow
go run ./examples/08_self_knowledge_graph # docs -> graphflow -> KG of this project; qa_test.go answers from graph edges
go run ./examples/16_ontology # typed object types, actions, object sets, typed tools, schema diffUse examples/05_graphflow to verify OpenAI-compatible LLM graph extraction with structured output.
When changing CortexDB, run:
go build ./...
go test ./...© liliang-cn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1,665 other files (scripts) in the repository root of liliang-cn/cortexdb.
Open the folder on GitHubat commit 17f8a4f
Cortexdb next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cortexdb this skillliliang-cn/cortexdb | 274 | — | ~18k | Automated safety check: Warn | MIT | |
| DBoracle/skills | 873 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | |
| Google Cloud Solution RAG Enterprise Search Gke Sqldbgoogle/skills | 21k | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| Enterprise AIoracle/skills | 873 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | |
| Neo4j Vector Index Skillneo4j-contrib/neo4j-skills | 114 | — | ~5.6k | Automated safety check: Notes | MIT | |
| Neo4j Graphrag Skillneo4j-contrib/neo4j-skills | 114 | — | ~4.2k | Automated safety check: Notes | MIT |
oracle/skills
Oracle Database guidance for SQL, PL/SQL, SQLcl, ORDS, Oracle Vector SDK, administration, app development, performance, security, migrations, and agent-safe database workflows.
Discovers requirements, and generates architectural, design, and deployment guidance for a retrieval-augmented generation (RAG)-capable enterprise search system in Google Cloud.
oracle/skills
Oracle Enterprise AI guidance for building, deploying, securing, estimating cost for, and integrating AI models, agents, RAG, Responses API workflows, custom or imported models, fine-tuning, model…
neo4j-contrib/neo4j-skills
Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…
neo4j-contrib/neo4j-skills
Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).
HiAi-gg/docsmint
Manage and research DocsMint documents through its scoped MCP tools, including categories, folders, hybrid search, GraphRAG, rerank, and index refresh.
liliang-cn/cortexdb
Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…
liliang-cn/cortexdb
Give a Node.js agent (such as OpenClaw) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client npm package.
liliang-cn/cortexdb
Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, and MCP/tool calling.
Categories
Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling. Cortexdb is an agent skill from liliang-cn/cortexdb. Use CortexDB for local-first AI memory, vector search, RAG, knowledge graphs, SPARQL/RDFS/SHACL, corpus-to-graph workflows, external structured-data import (CSV / SQL dumps), and MCP/tool calling.
Cortexdb fits situations like: working with CortexDB; knowledge graph; CSV/SQL dump import.
Run `npx skills add liliang-cn/cortexdb --skill cortexdb -a claude-code`. Or copy the skill folder (the liliang-cn/cortexdb repository) into .claude/skills/cortexdb in your project. Claude Code loads it when a task matches its description.
Run `npx skills add liliang-cn/cortexdb --skill cortexdb -a codex`. Or copy the skill folder (the liliang-cn/cortexdb repository) into .agents/skills/cortexdb in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add liliang-cn/cortexdb --skill cortexdb -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cortexdb, .gemini/skills/cortexdb, .github/skills/cortexdb and .opencode/skills/cortexdb in your project.
Going by SKILL.md and its folder, Cortexdb needs the command-line tools its instructions call (go, cargo, pip, npm and npx) and credentials named CORTEXDB_GRPC_TOKEN and OPENAI_API_KEY.
SKILL.md names 1 domain. In commands or code: w3.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): links to a raw public ip address. Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Cortexdb is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 18k tokens (SKILL.md is roughly 73k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cortexdb: DB (oracle/skills, 873 stars), Google Cloud Solution RAG Enterprise Search Gke Sqldb (google/skills, 21k stars), Enterprise AI (oracle/skills, 873 stars) and Neo4j Vector Index Skill (neo4j-contrib/neo4j-skills, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
liliang-cn (a GitHub user) maintains it in liliang-cn/cortexdb, which has 274 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.
Source: liliang-cn/cortexdb on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.