Agent skill

Cognee Custom Graph Models

by topoteretes in topoteretes/cognee

Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Cognee Custom Graph Models

skills CLI
$ npx skills add topoteretes/cognee --skill cognee-custom-graph-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install topoteretes/cognee cognee-custom-graph-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cognee-custom-graph-models .claude/skills/cognee-custom-graph-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cognee-custom-graph-models
GitHub stars
32k
Token cost
~2.5k tokens
SKILL.md length
1,033 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes.

  • Defining custom node and edge types for a cognee knowledge graph
  • SKILL.md covers Use it, Pitfalls, How it works and Extending it
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Fixing the same person appearing as duplicated nodes across runs

What it does

By default cognee extracts a generic KnowledgeGraph of entities and relationships. This skill shows how to pass your own model with graph_model= so the LLM fills your node and edge types instead. Every node type subclasses DataPoint: scalar and string fields become node properties, while a field holding a DataPoint or a list of them becomes edges named after that field.

The metadata key controls behavior. identity_fields derive the node id so the same entity from two chunks or two runs merges into one node, index_fields create a vector collection per field so recall can find the node, and transparent drops a root container. Without identity_fields every node gets a random id and duplicates pile up. The skill advises writing metadata explicitly, because the Dedup annotation shortcut works but Embeddable currently does not index.

It also covers typed Edge fields, FromIdentity references, building a model from a JSON schema, and debugging duplicated nodes, missing edges and InvalidReferenceTypeError.

When your agent uses it

  • Defining custom node and edge types for a cognee knowledge graph
  • Fixing the same person appearing as duplicated nodes across runs
  • Making nodes findable by recall through index fields
  • Debugging missing edges or an InvalidReferenceTypeError

Example prompts

  • “Write DataPoint classes for Person and Company with a works_at edge for cognee.”
  • “My cognee graph repeats the same person in every chunk, so set up identity fields so they merge.”
  • “Build a cognee graph_model from the JSON schema in ./schemas/invoice.json.”

Requirements

  • cognee installed in a Python project
  • An LLM configured for cognee to extract the graph

What it can do on your machine

Read from SKILL.md and the folder at commit 0ec7a9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cognee Custom Graph Models loads about 2.5k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 1,033 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from topoteretes/cognee at commit 0ec7a9f, republished under its Apache-2.0 licence (© topoteretes). 1,033 words, ~2,451 tokens.

Download SKILL.mdSave it as .claude/skills/cognee-custom-graph-models/SKILL.md (or your agent's skills folder).
name
cognee-custom-graph-models
description
Use when defining the shape of cognee's knowledge graph with graph_model= — writing DataPoint node classes, choosing identity and index fields so nodes merge and are searchable, declaring typed Edge fields and FromIdentity references, building a model from a JSON schema, or debugging duplicated nodes, missing edges, or InvalidReferenceTypeError.

Custom graph models

By default cognee extracts a generic KnowledgeGraph of entities and relationships. Pass your own model with graph_model= and the LLM fills your node and edge types instead.

python
from typing import Annotated, Literal
import cognee
from cognee.low_level import DataPoint, Edge, FromIdentity


class Role(DataPoint):
    name: str
    metadata: dict = {"index_fields": ["name"], "identity_fields": ["name"]}


class Person(DataPoint):
    name: str
    is_a: Annotated[Role, FromIdentity()] | None = None  # reference by name
    reports_to: list[Edge["Person", "Person"]] = []  # edge owned by Person
    metadata: dict = {"index_fields": ["name"], "identity_fields": ["name"]}


class PeopleGraph(DataPoint):  # the root the LLM fills
    people: list[Person]
    friends_with: list[Edge[Person, Person]] = []
    family: list[Edge[Person, Person, Literal["married_to", "sibling_of"]]] = []
    metadata: dict = {"index_fields": [], "transparent": True}


await cognee.remember(text, graph_model=PeopleGraph, custom_prompt="Extract every person...")

Full example: examples/guides/custom_graph_model.py.

Use it

Nodes: DataPoint classes

Every node type subclasses DataPoint (from cognee.low_level import DataPoint). Its fields become:

  • Node properties: scalars, strings, dicts, and anything that is not a DataPoint.
  • Edges named after the field: a field holding a DataPoint or a list of DataPoints. members: list[Person] becomes members edges.

dict[str, Person], sets and plain tuples are stored as properties, not edges.

Identity and search: metadata
KeyWhat it does
identity_fieldsThe node id is derived from these field values (normalized: lowercased, spaces to _, apostrophes removed). The same entity from two chunks or two runs becomes one node.
index_fieldsEach field gets a vector collection named <ClassName>_<field>, so recall can find the node.
transparentThe node is not stored; its children take its place. Use it for a root container like PeopleGraph.

Without identity_fields every node gets a random id, so the same person is duplicated in every chunk and every run. Set it on every node type that represents a real-world entity.

Write metadata explicitly, as in the examples above. There is also an annotation shortcut (from cognee.infrastructure.engine import Dedup, Embeddable; name: Annotated[str, Embeddable(), Dedup()]), but today only half of it works:

  • Dedup() works: ids are derived from the marked fields.
  • Embeddable() does not index. The markers update the class-level default, but each instance still carries {"index_fields": []}, and indexing reads the instance, so no vector collection is created and recall cannot find the node.

Markers are also ignored entirely when the class declares metadata itself.

Typed edges: list[Edge[Source, Target, Name]]

The LLM answers edges as flat rows of identity strings (source, target), and cognee resolves them to the extracted nodes. The third parameter controls the relationship name:

DeclarationRelationship name
list[Edge[Person, Person]]The field name (friends_with)
list[Edge[Person, Person, Literal["a", "b"]]]The LLM picks one value
list[Edge[Person, Person, str]]Free-form from the LLM, normalized
  • Where to declare: on the root model for relationships with no obvious owner, or on the owning node. On the owner, endpoints of the owner's own type must be strings (Edge["Person", "Person"]), because the class is not defined yet inside its own body.
  • Always a list: Edge[...] or Edge[...] | None on its own raises.
  • Both endpoint types need exactly one identity_fields entry.
References: Annotated[Target, FromIdentity()]

Instead of a nested object, the LLM answers the identity string of a node (is_a: "engineer"), and cognee links to that node. Supported spellings: Target, Target | None, list[Target], list[Target] | None. Anything else raises InvalidReferenceTypeError. The target needs exactly one identity field, and its other required fields need defaults.

Edge values you build by hand

Edge(source=..., target=..., relationship_type=..., weight=..., properties={...}). An omitted source falls back to the node declaring the field; on a parametrized field that node must be the declared Source type, or it raises. On a root container, always pass source=. The tuple form (Edge(weight=0.8), target_node) attaches edge properties to a plain DataPoint field (examples/guides/custom_data_models.py).

From JSON instead of Python
  • cognee.low_level.graph_model_from_spec(spec): a small entity/relation spec (names, fields, one/many relations) compiled to DataPoint classes, with identity and index on name by default. Example: examples/guides/graph_model_from_json.py.
  • cognee.low_level.graph_schema_to_graph_model(json_schema): a JSON Schema (needs a top-level title; only internal # refs).
  • HTTP: POST /api/v1/remember takes a graph_model form field (JSON schema string); POST /api/v1/cognify takes a graphModel JSON object. POST /api/v1/llm/infer-schema proposes a schema from sample text.

Neither JSON path can express typed Edge fields or FromIdentity; use Python classes for those.

Show full SKILL.md (446 more words)Show less

Pitfalls

  • Duplicated nodes → missing identity_fields.
  • Node never shows up in recall → no index_fields, or the recalling process never imported the model class (graph completion searches the collections of DataPoint classes loaded in that process).
  • Edge rows silently missing → an unresolved row is dropped with a warning. Endpoints resolve by exact type against nodes extracted from the same chunk: a subclass instance does not match where its base is declared, and a node from another chunk or an earlier run is not a candidate.
  • InvalidReferenceTypeError, "declares Edge in a shape the LLM extraction cannot fill" → an unsupported FromIdentity or Edge spelling. These are raised when the model is converted during extraction, not at class definition, so they appear mid-pipeline.
  • String endpoints resolve only to the owning model itself or a module-level class. A string naming another class defined inside a function raises InvalidReferenceTypeError, so define models at module level.
  • Field names that collide with DataPoint's own fields (id, type, version, metadata, created_at, belongs_to_set, …) are stripped from what the LLM sees. Rename them.
  • A subclass that overrides metadata replaces the parent's entirely. Dropping identity_fields only logs a warning. A subclass also hashes ids under its own class name.
  • Write a custom_prompt. Without one the generic knowledge-graph prompt is used; your schema reaches the LLM only as structured output.
  • Not with GLiNER. extractor="gliner_demo" (alias "gliner") raises with a custom graph_model, and so does the default GRAPH_EXTRACTOR=auto when no LLM key is configured (it resolves to gliner_demo).
  • Remote mode drops it. After cognee.serve(url), remember() and cognify() do not forward graph_model; the server builds a generic graph.
  • A custom model skips the generic path's ontology resolution, per-graph node dedup, and functional_relationships. Summaries still run.

How it works

extract_content_graph converts a DataPoint model into a plain Pydantic model for the LLM: infrastructure fields and metadata are stripped, typed edge fields become row lists (FriendsWithEdge with source/target strings), and FromIdentity fields become strings. The answer is converted back into DataPoint instances with ids from identity_fields, edge rows are resolved against the nodes in that answer, and the result is attached to the chunk (chunk.contains) and stored by add_data_points. Ownership is recorded per document, so forget(data_id=...) removes a custom-model document's nodes while shared nodes survive.

  • LLM boundary, both directions: cognee/shared/llm_graph_model.py
  • DataPoint, metadata, ids: cognee/infrastructure/engine/models/DataPoint.py
  • Markers: cognee/infrastructure/engine/models/FieldAnnotations.py
  • Edge: cognee/infrastructure/engine/models/Edge.py
  • Property vs edge decision: cognee/modules/graph/utils/field_edges.py
  • Custom-model branch of extraction: cognee/tasks/graph/extract_graph_from_data.py
  • JSON schema / spec paths: cognee/shared/graph_model_utils.py, cognee/modules/graph_models/
  • Vector collections: cognee/tasks/storage/index_data_points.py

Extending it

  • Tests for the LLM round trip: cognee/tests/unit/modules/graph/test_content_graph_to_data_point.py. Edge typing: cognee/tests/unit/interfaces/graph/test_typed_edge_model.py, test_typed_edges_graph.py. Identity: cognee/tests/unit/infrastructure/engine/test_identity_fields.py. Property vs edge: cognee/tests/unit/modules/graph/test_field_edges.py.
  • Deletion of custom-model nodes: cognee/tests/test_delete_custom_graph.py.
  • A new Edge or FromIdentity spelling must be handled in both directions in llm_graph_model.py, and rejected with InvalidReferenceTypeError when unsupported, never silently accepted.

© topoteretes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/cognee-custom-graph-models of topoteretes/cognee.

Open the folder on GitHubat commit 0ec7a9f

Compare with similar skills

Cognee Custom Graph Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cognee Custom Graph Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cognee Custom Graph Models this skilltopoteretes/cognee32k—~2.5kAutomated safety check: PassApache-2.0
Cortexdb Memory Hermesliliang-cn/cortexdb274—~1.7kAutomated safety check: PassMIT
Neo4j Graphrag Skillneo4j-contrib/neo4j-skills114—~4.2kAutomated safety check: NotesMIT
Hermes Memory Providersmnemosyne-oss/mnemosyne3.4k—~1.8kAutomated safety check: PassMIT
Neo4j Genai Plugin Skillneo4j-contrib/neo4j-skills114—~3kAutomated safety check: NotesMIT
Cortexdb Memory Openclawliliang-cn/cortexdb274—~1.6kAutomated safety check: PassMIT

Similar skills

  • Cortexdb Memory Hermes

    liliang-cn/cortexdb

    Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…

    274 GitHub stars~1.7k tokensUpdated 2 days ago
    Knowledge ManagementAuto-check passed
  • Neo4j Graphrag Skill

    neo4j-contrib/neo4j-skills

    Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).

    114 GitHub stars~4.2k tokensUpdated today
    Knowledge ManagementAuto-check: notes
  • Hermes Memory Providers

    mnemosyne-oss/mnemosyne

    Install and configure Mnemosyne as a Hermes Agent memory provider — local SQLite with vector search, episodic consolidation, and temporal knowledge graphs.

    3.4k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Neo4j Genai Plugin Skill

    neo4j-contrib/neo4j-skills

    Use Neo4j GenAI Plugin ai.text. An agent skill from neo4j-contrib/neo4j-skills.

    114 GitHub stars~3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Cortexdb Memory Openclaw

    liliang-cn/cortexdb

    Give a Node.js agent (such as OpenClaw) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client npm package.

    274 GitHub stars~1.6k tokensUpdated 2 days ago
    Knowledge ManagementAuto-check passed
  • Compact Memory Implementation

    simbajigege/book2skills

    A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.

    183 GitHub stars~2.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from topoteretes/cognee

All 19 skills in this repo
  • Cognee CLI Memory Commands

    topoteretes/cognee

    Drives cognee from the terminal with remember, recall, forget and improve memory commands, dataset and config management and database migrations.

    32k GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • Cognee Community Packages

    topoteretes/cognee

    Guide to using and contributing cognee community packages: database adapters, data-source connectors, custom tasks and retrievers, and Keywords AI observability.

    32k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Cognee Custom Pipelines

    topoteretes/cognee

    Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph.

    32k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Cognee Docker Setup

    topoteretes/cognee

    Runs the Cognee AI memory platform in Docker, from a one-file prebuilt image to a full compose stack with UI, MCP server, Postgres and Neo4j.

    32k GitHub stars~901 tokensUpdated today
    Auto-check: notes
  • Cognee Forget

    topoteretes/cognee

    Removes data from cognee memory with forget(), finding the right dataset and document first and choosing between one document, a dataset or only the graph and vector memory.

    32k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Explains how cognee stores session memory by session_id and bridges it into the permanent graph with improve(), including the stages, results and settings.

    32k GitHub stars~3.2k tokensUpdated today
    Auto-check passed

Works with

Questions about Cognee Custom Graph Models

What does Cognee Custom Graph Models do?

Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes. By default cognee extracts a generic KnowledgeGraph of entities and relationships. This skill shows how to pass your own model with graph_model= so the LLM fills your node and edge types instead.

When should I use Cognee Custom Graph Models?

Cognee Custom Graph Models fits situations like: defining custom node and edge types for a cognee knowledge graph; fixing the same person appearing as duplicated nodes across runs; making nodes findable by recall through index fields; debugging missing edges or an InvalidReferenceTypeError.

How do I install Cognee Custom Graph Models in Claude Code?

Run `npx skills add topoteretes/cognee --skill cognee-custom-graph-models -a claude-code`. Or copy the skill folder (.agents/skills/cognee-custom-graph-models in topoteretes/cognee) into .claude/skills/cognee-custom-graph-models in your project. Claude Code loads it when a task matches its description.

How do I install Cognee Custom Graph Models in Codex?

Run `npx skills add topoteretes/cognee --skill cognee-custom-graph-models -a codex`. Or copy the skill folder (.agents/skills/cognee-custom-graph-models in topoteretes/cognee) into .agents/skills/cognee-custom-graph-models in your project. Codex loads it when a task matches its description.

Can I use Cognee Custom Graph Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add topoteretes/cognee --skill cognee-custom-graph-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cognee-custom-graph-models, .gemini/skills/cognee-custom-graph-models, .github/skills/cognee-custom-graph-models and .opencode/skills/cognee-custom-graph-models in your project.

What does Cognee Custom Graph Models need to run?

SKILL.md names no scripts, command-line tools or credentials: Cognee Custom Graph Models is instructions for the agent only. Our summary lists: cognee installed in a Python project; An LLM configured for cognee to extract the graph.

Does Cognee Custom Graph Models access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cognee Custom Graph Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cognee Custom Graph Models use?

Cognee Custom Graph Models is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cognee Custom Graph Models use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cognee Custom Graph Models?

Skills that share tags, products or a category with Cognee Custom Graph Models: Cortexdb Memory Hermes (liliang-cn/cortexdb, 274 stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Hermes Memory Providers (mnemosyne-oss/mnemosyne, 3.4k stars) and Neo4j Genai Plugin Skill (neo4j-contrib/neo4j-skills, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cognee Custom Graph Models?

topoteretes (a GitHub organization) maintains it in topoteretes/cognee, which has 31,919 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: topoteretes/cognee on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.