Agent skill

Cognee Custom Pipelines

by topoteretes in topoteretes/cognee

Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Cognee Custom Pipelines

skills CLI
$ npx skills add topoteretes/cognee --skill cognee-custom-pipelines -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install topoteretes/cognee cognee-custom-pipelines --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cognee-custom-pipelines .claude/skills/cognee-custom-pipelines && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cognee-custom-pipelines
GitHub stars
32k
Token cost
~2.8k tokens
SKILL.md length
985 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph.

  • Writing a custom extraction or enrichment task for cognee
  • SKILL.md covers Use it, Pitfalls, How it works and Extending it
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Chaining tasks with run_custom_pipeline over a dataset

What it does

In cognee everything runs as a pipeline of tasks, each a plain Python function whose output feeds the next. The skill explains when remember is enough and when to build your own pipeline: custom extraction, custom node types or post-processing over the graph. It compares three runners: run_custom_pipeline for the normal case with permissions, a per-dataset lock, run records and status; the full orchestrator it builds on; and a lightweight run_pipeline for quick chains with no permissions, locks or run rows.

Tasks can be async, generators or plain functions, take extra arguments after the pipeline data, accept a ctx parameter carrying user, data item, dataset and run identifiers, and be flagged with needs_llm set to false so LLM-free pipelines skip the connection check. Further topics include add_data_points for storing DataPoints, memify for enrichment over the existing graph, checking run status, and how data flows between tasks through batch size and per-document processing.

When your agent uses it

  • Writing a custom extraction or enrichment task for cognee
  • Chaining tasks with run_custom_pipeline over a dataset
  • Storing your own DataPoint types in the graph
  • Debugging how data moves between tasks, including batch size and Drop

Example prompts

  • “Write a cognee task that extracts action items and stores them as DataPoints.”
  • “Run a custom pipeline over the existing documents in my dataset in the background.”
  • “Why does my second task receive a batch instead of a single item?”

Requirements

  • Python with the cognee package

What it can do on your machine

Read from SKILL.md and the folder at commit 0ec7a9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cognee Custom Pipelines loads about 2.8k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 985 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from topoteretes/cognee at commit 0ec7a9f, republished under its Apache-2.0 licence (© topoteretes). 985 words, ~2,818 tokens.

Download SKILL.mdSave it as .claude/skills/cognee-custom-pipelines/SKILL.md (or your agent's skills folder).
name
cognee-custom-pipelines
description
Use when building your own cognee processing — writing custom tasks, chaining them into a pipeline with run_custom_pipeline or the lightweight run_pipeline (from cognee.pipelines import run_pipeline), storing custom DataPoints with add_data_points, running custom extraction/enrichment over the existing graph with memify, checking pipeline run status, or debugging how data flows between tasks (batch_size, data_per_batch, ctx, Drop, enriches).

Custom tasks and pipelines

Everything cognee does runs as a pipeline: an ordered list of tasks, each a plain Python function whose output feeds the next one. remember() is the right tool for ordinary ingestion. Build a pipeline when you need processing cognee does not ship: your own extraction, your own node types, or a post-processing step over the graph.

python
import cognee
from cognee.modules.pipelines import Task
from cognee.tasks.storage import add_data_points
from cognee.low_level import DataPoint


class Person(DataPoint):
    name: str
    metadata: dict = {"index_fields": ["name"], "identity_fields": ["name"]}


async def extract_people(data_items: list) -> list[Person]:
    people = []
    for item in data_items:  # always a list, see below
        text = item if isinstance(item, str) else ""
        people += [Person(name=n.strip()) for n in text.split(",") if n.strip()]
    return people


result = await cognee.run_custom_pipeline(
    tasks=[
        Task(extract_people, needs_llm=False),
        Task(add_data_points, needs_llm=False),  # store in graph + vector DBs
    ],
    data=["Ada Lovelace, Alan Turing"],
    dataset="people",
)

Use it

Pick the runner

There are three, and two share the name run_pipeline:

RunnerImportUse it for
cognee.run_custom_pipeline(...)cogneeThe normal choice: runs your tasks against a dataset with permissions, a per-dataset lock, run records, and status
Full orchestrator run_pipeline(tasks=..., data=..., datasets=...)cognee.modules.pipelinesWhat run_custom_pipeline and cognify call; yields PipelineRunInfo
Lightweight run_pipeline([...], data=...)from cognee.pipelines import run_pipeline (after import cognee, the attribute cognee.pipelines.run_pipeline is the orchestrator)Quick chains of task() specs with no permissions, locks, run rows, or migrations; returns the last step's outputs

cognee.run_custom_pipeline(tasks, data=None, dataset="main_dataset", user=None, incremental_loading=False, data_per_batch=20, run_in_background=False, pipeline_name="custom_pipeline", data_cache=False, ...) returns {dataset_id: PipelineRunInfo} (the started run when run_in_background=True). With data=None it runs over the dataset's existing documents (Data rows).

Write a task
python
from cognee.modules.pipelines import Task
from cognee.modules.pipelines.models import PipelineContext
from cognee.modules.pipelines.tasks.task import task_summary
from cognee.pipelines import Drop


@task_summary("Tagged {n} chunk(s)")
async def tag_chunks(chunks: list, ctx: PipelineContext = None, label: str = "x"):
    for chunk in chunks:
        chunk.metadata["label"] = label
    return chunks  # or yield per item; return/yield Drop to discard


tag = Task(tag_chunks, label="reviewed", batch_size=10, needs_llm=False)
  • A task is an async def, a generator, an async generator, or a plain def. Extra Task(fn, *args, **kwargs) arguments are passed after the pipeline data.
  • needs_llm=False on tasks that never call an LLM lets an LLM-free pipeline skip the LLM connection check.
  • ctx (injected by the parameter name ctx) carries user, data_item, dataset, pipeline_run_id, pipeline_name, and extras.
  • task.with_config(batch_size=..., **kwargs) returns a modified copy.
How data flows
  • Each document runs the whole chain on its own with run_custom_pipeline or the orchestrator, and the first task receives it as a one-element list ([data_item]), not the bare item. The lightweight run_pipeline passes data to the first task unchanged.
  • data_per_batch (default 20) is how many documents run at the same time. It is a concurrency limit, not a batch size.
  • batch_size belongs to the consumer. A task's batch_size decides how the previous task's generator output is grouped before it is passed in. Generator tasks always hand over lists; a coroutine or function hands over its single return value.
  • Streaming: each upstream result goes down the chain immediately, so a downstream task can run many times per document.
  • enriches=True: if the task returns None, its input is passed on unchanged (coroutines and functions only, not generators).
  • Drop: returning or yielding it removes that item from the stream.
  • Every DataPoint passing through is stamped automatically with where it came from (source_pipeline, source_task, source_user, …).
Store results

add_data_points(data_points, custom_edges=None, embed_triplets=False, graph_only=False) writes a list of DataPoints to the graph and indexes their index_fields in the vector DB. It returns the same list, so it can sit mid-chain. Give every node type identity_fields so repeated runs merge instead of duplicating (see the cognee-custom-graph-models skill).

Work on the existing graph: memify
python
await cognee.memify(
    extraction_tasks=["extract_subgraph_chunks"],  # names or Task objects
    enrichment_tasks=[Task(my_enrichment, needs_llm=False)],
    dataset="people",
    node_name=["AI"],  # optional subgraph filter
)

With no data, memify passes the graph (or the node_type / node_name subgraph) to the first task. Registered task names: extract_subgraph, extract_subgraph_chunks, get_triplet_datapoints, extract_user_sessions, cognify_session, extract_agent_trace_feedbacks, cognify_agent_trace_feedback, apply_feedback_weights, detect_entity_duplicates, merge_entity_duplicates, index_data_points. improve() also forwards extraction_tasks / enrichment_tasks to memify, but only inside its enrichment stage. With custom tasks that stage skips the TRIPLET_EMBEDDING gate and the has-the-graph-changed check, so they run on every improve (unless the stage is disabled, the lock is held, or the fatal persist_session_qa stage errors and stops the run first).

Check status
python
status = await cognee.datasets.get_status([dataset_id], pipeline_names=["custom_pipeline"])

Without pipeline_names it reports only cognify_pipeline. It returns {str(dataset_id): PipelineRunStatus} ({str(dataset_id): {pipeline_name: status}} for several pipeline_names): DATASET_PROCESSING_STARTED, _COMPLETED, or _ERRORED (_INITIATED exists only on legacy rows). The value run_custom_pipeline returns per dataset is a PipelineRunInfo instead, whose class names the outcome: PipelineRunCompleted, PipelineRunAlreadyCompleted, PipelineRunErrored, and so on.

Show full SKILL.md (395 more words)Show less

Pitfalls

  • Wrong run_pipeline. The one imported via from cognee.pipelines import run_pipeline wants task() specs called (extract(), not extract) and raises TypeError otherwise; the orchestrator in cognee.modules.pipelines wants Task objects and raises WrongTaskTypeError otherwise.
  • String task names only work in memify. run_custom_pipeline accepts only Task objects despite its type hint.
  • Some callables are rejected by Task (ValueError: Unsupported task type): bound methods and functools.partials of a plain (non-generator) sync function, and callable objects (instances with __call__). Generator and async variants, plain functions, and lambdas work. When in doubt, wrap it in a plain def / async def.
  • run_custom_pipeline does not run database migrations. On an existing database, run await cognee.run_migrations() (or any remember() first).
  • Keep pipeline_name="custom_pipeline" unless you add your name to WRITE_PIPELINE_NAMES in cognee/modules/improve/graph_changes.py. Otherwise improve() does not notice your graph writes and may skip enrichment as "already completed".
  • memify defaults. An omitted or empty task list is replaced by the defaults: index_data_points enrichment, plus get_triplet_datapoints extraction only when TRIPLET_EMBEDDING=true (off by default). memify uses only the first dataset it resolves.
  • Nodes duplicate on every run when a DataPoint has no identity_fields (or Dedup() fields). examples/guides/custom_data_models.py and examples/guides/custom_tasks_and_pipelines.py have this bug; don't copy it.

How it works

run_custom_pipeline → orchestrator run_pipeline (checks write permission, takes the per-dataset lock, records a PipelineRun) → run_tasks (a semaphore of data_per_batch, one chain per document) → run_tasks_base (streams each task's output into the next, batching by the consumer's batch_size, injecting ctx, stamping provenance).

  • Package overview and the runner semantics: cognee/modules/pipelines/__init__.py
  • Task, task(), TaskSpec, BoundTask, @task_summary: cognee/modules/pipelines/tasks/task.py
  • Orchestrator: cognee/modules/pipelines/operations/pipeline.py; execution: run_tasks.py, run_tasks_base.py, run_tasks_data_item.py
  • Lightweight runner: cognee/modules/pipelines/operations/run_pipeline.py, exported from cognee/pipelines/
  • Context: cognee/modules/pipelines/models/PipelineContext.py
  • run_custom_pipeline: cognee/modules/run_custom_pipeline/run_custom_pipeline.py
  • memify: cognee/modules/memify/memify.py, cognee/memify_pipelines/memify_task_registry.py, memify_default_tasks.py
  • Storage: cognee/tasks/storage/add_data_points.py
  • Index of all shipped tasks: cognee/tasks/README.md

Examples:

  • examples/demos/custom_pipelines/custom_pipeline_single_object_example.py: the best reference. It runs over added documents, does LLM extraction into typed DataPoints, then recalls. Its models declare identity with Dedup() (the Annotated alternative to metadata["identity_fields"]).
  • examples/demos/custom_pipelines/organizational_hierarchy/: low-level run_tasks, no LLM, dedup via identity_fields, status polling.
  • examples/demos/custom_pipelines/custom_cognify_pipeline_example.py: rebuilds add + cognify from the default task list.
  • examples/demos/custom_pipelines/memify_coding_agent_rule_extraction_example.py: memify with a custom enrichment task.

Extending it

  • A new shipped task: put it in the cognee/tasks/ subpackage for its stage, export it from that package's __init__.py, follow the template in cognee/tasks/README.md, and add a unit test under cognee/tests/unit/tasks/.
  • A new memify task name: register it in cognee/memify_pipelines/memify_task_registry.py.
  • A new write pipeline name: add it to WRITE_PIPELINE_NAMES.
  • Pipeline tests: cognee/tests/unit/modules/pipelines/ (runner semantics, context, provenance, rollback) and cognee/tests/unit/pipelines/ (the lightweight API).

© topoteretes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/cognee-custom-pipelines of topoteretes/cognee.

Open the folder on GitHubat commit 0ec7a9f

Compare with similar skills

Cognee Custom Pipelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cognee Custom Pipelines compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cognee Custom Pipelines this skilltopoteretes/cognee32k—~2.8kAutomated safety check: PassApache-2.0
Cortexdb Memory Hermesliliang-cn/cortexdb274—~1.7kAutomated safety check: PassMIT
Neo4j Graphrag Skillneo4j-contrib/neo4j-skills114—~4.2kAutomated safety check: NotesMIT
Hermes Memory Providersmnemosyne-oss/mnemosyne3.4k—~1.8kAutomated safety check: PassMIT
Compact Memory Implementationsimbajigege/book2skills183—~2.5kAutomated safety check: PassMIT
Deep Agents Corelangchain-ai/langchain-skills1.3k—~3.1kAutomated safety check: PassMIT

Similar skills

  • Cortexdb Memory Hermes

    liliang-cn/cortexdb

    Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…

    274 GitHub stars~1.7k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check passed
  • Neo4j Graphrag Skill

    neo4j-contrib/neo4j-skills

    Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).

    114 GitHub stars~4.2k tokensUpdated yesterday
    Knowledge ManagementAuto-check: notes
  • Hermes Memory Providers

    mnemosyne-oss/mnemosyne

    Install and configure Mnemosyne as a Hermes Agent memory provider — local SQLite with vector search, episodic consolidation, and temporal knowledge graphs.

    3.4k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Compact Memory Implementation

    simbajigege/book2skills

    A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.

    183 GitHub stars~2.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Deep Agents Core

    langchain-ai/langchain-skills

    Official

    Explains how to build agents with the Deep Agents framework: create_deep_agent, the built-in middleware, the harness, SKILL.md format and configuration options.

    1.3k GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Adds, searches, lists, updates and deletes memories on the Mem0 platform from the terminal with the mem0 command, including a JSON mode built for agents.

    67k GitHub stars~2k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check: notes

More from topoteretes/cognee

All 19 skills in this repo
  • Cognee CLI Memory Commands

    topoteretes/cognee

    Drives cognee from the terminal with remember, recall, forget and improve memory commands, dataset and config management and database migrations.

    32k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Cognee Community Packages

    topoteretes/cognee

    Guide to using and contributing cognee community packages: database adapters, data-source connectors, custom tasks and retrievers, and Keywords AI observability.

    32k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Cognee Custom Graph Models

    topoteretes/cognee

    Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes.

    32k GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Cognee Docker Setup

    topoteretes/cognee

    Runs the Cognee AI memory platform in Docker, from a one-file prebuilt image to a full compose stack with UI, MCP server, Postgres and Neo4j.

    32k GitHub stars~901 tokensUpdated yesterday
    Auto-check: notes
  • Cognee Forget

    topoteretes/cognee

    Removes data from cognee memory with forget(), finding the right dataset and document first and choosing between one document, a dataset or only the graph and vector memory.

    32k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Explains how cognee stores session memory by session_id and bridges it into the permanent graph with improve(), including the stages, results and settings.

    32k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Cognee Custom Pipelines

What does Cognee Custom Pipelines do?

Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph. In cognee everything runs as a pipeline of tasks, each a plain Python function whose output feeds the next. The skill explains when remember is enough and when to build your own pipeline: custom extraction, custom node types or post-processing over the graph.

When should I use Cognee Custom Pipelines?

Cognee Custom Pipelines fits situations like: writing a custom extraction or enrichment task for cognee; chaining tasks with run_custom_pipeline over a dataset; storing your own DataPoint types in the graph; debugging how data moves between tasks, including batch size and Drop.

How do I install Cognee Custom Pipelines in Claude Code?

Run `npx skills add topoteretes/cognee --skill cognee-custom-pipelines -a claude-code`. Or copy the skill folder (.agents/skills/cognee-custom-pipelines in topoteretes/cognee) into .claude/skills/cognee-custom-pipelines in your project. Claude Code loads it when a task matches its description.

How do I install Cognee Custom Pipelines in Codex?

Run `npx skills add topoteretes/cognee --skill cognee-custom-pipelines -a codex`. Or copy the skill folder (.agents/skills/cognee-custom-pipelines in topoteretes/cognee) into .agents/skills/cognee-custom-pipelines in your project. Codex loads it when a task matches its description.

Can I use Cognee Custom Pipelines in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add topoteretes/cognee --skill cognee-custom-pipelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cognee-custom-pipelines, .gemini/skills/cognee-custom-pipelines, .github/skills/cognee-custom-pipelines and .opencode/skills/cognee-custom-pipelines in your project.

What does Cognee Custom Pipelines need to run?

SKILL.md names no scripts, command-line tools or credentials: Cognee Custom Pipelines is instructions for the agent only. Our summary lists: Python with the cognee package.

Does Cognee Custom Pipelines access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cognee Custom Pipelines safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cognee Custom Pipelines use?

Cognee Custom Pipelines is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cognee Custom Pipelines use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cognee Custom Pipelines?

Skills that share tags, products or a category with Cognee Custom Pipelines: Cortexdb Memory Hermes (liliang-cn/cortexdb, 274 stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Hermes Memory Providers (mnemosyne-oss/mnemosyne, 3.4k stars) and Compact Memory Implementation (simbajigege/book2skills, 183 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cognee Custom Pipelines?

topoteretes (a GitHub organization) maintains it in topoteretes/cognee, which has 31,919 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: topoteretes/cognee on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.