Agent skill

Cognee Data Ingestion

by topoteretes in topoteretes/cognee

Explains how to load text, files, folders, URLs, repositories and databases into cognee memory with remember(), including datasets, extractors and tags.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Cognee Data Ingestion

skills CLI
$ npx skills add topoteretes/cognee --skill cognee-ingestion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install topoteretes/cognee cognee-ingestion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cognee-ingestion .claude/skills/cognee-ingestion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cognee-ingestion
GitHub stars
32k
Token cost
~2.6k tokens
SKILL.md length
1,087 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Explains how to load text, files, folders, URLs, repositories and databases into cognee memory with remember(), including datasets, extractors and tags.

  • Loading documents, folders or URLs into cognee memory
  • SKILL.md covers Use it, Pitfalls, How it works and Extending it
  • Calls pip; needs LLM_API_KEY
  • Indexing a code repository or a SQL database through remember()

What it does

This skill covers cognee's remember() call, the single ingestion entry point that stores data, builds a knowledge graph from it and enriches the graph. It explains what remember() accepts: strings, lists of strings, file paths (absolute, file:// or s3://), http and https URLs, binary streams, folders, code repositories and GitHub or GitLab URLs, SQL connection strings, dlt sources, CSV files and SKILL.md playbooks.

It then explains the options that shape the result: dataset_name or dataset_id for the target dataset and its permissions, node_set tags that recall can filter on later, session_id for writing to the fast session cache, and an extractor choice between an LLM and GLiNER, picked automatically from whether an API key is configured. Code files take a separate deterministic route with no LLM calls, searchable with SearchType.CODE. The skill is meant for questions about loaders, ontologies, chunking, dry-run cost estimates and temporal graphs, and for diagnosing a remember() call that rejects a keyword argument.

When your agent uses it

  • Loading documents, folders or URLs into cognee memory
  • Indexing a code repository or a SQL database through remember()
  • Choosing between the LLM and GLiNER graph extractors
  • Tagging data with node sets or splitting it across datasets
  • Debugging a remember() call that raises on a keyword argument

Example prompts

  • “Ingest everything under ./docs into a cognee dataset called product_docs.”
  • “Add our GitHub repository to cognee so I can search the code graph.”
  • “Store these meeting notes in cognee with the node set Sales so recall can filter on it later.”
  • “My cognee.remember call fails on an unexpected keyword argument, so work out which option is wrong.”

Requirements

  • Python with the cognee package
  • An LLM_API_KEY unless using the GLiNER extractor
  • git on PATH when ingesting code repositories

What it can do on your machine

Read from SKILL.md and the folder at commit 0ec7a9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cognee Data Ingestion loads about 2.6k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,087 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:88
    `.env`. `annotate` (default) only enriches; `strict` drops entities that

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from topoteretes/cognee at commit 0ec7a9f, republished under its Apache-2.0 licence (© topoteretes). 1,087 words, ~2,625 tokens.

Download SKILL.mdSave it as .claude/skills/cognee-ingestion/SKILL.md (or your agent's skills folder).
name
cognee-ingestion
description
Use when putting data into cognee memory with remember() — choosing inputs (text, files, folders, URLs, repos, databases), datasets and node_sets, loaders, ontologies, the graph extractor (LLM or GLiNER), chunking, dry-run cost estimates, or when remember() raises on a keyword argument.

Ingest data with remember()

remember() is cognee's ingestion API. One call stores the data, builds the knowledge graph, and enriches it. Use it for all ingestion; every option in this skill is a remember() argument unless it says otherwise.

python
import cognee

result = await cognee.remember("Einstein was born in Ulm.")  # text
result = await cognee.remember(
    ["./notes.md", "./report.pdf"],  # files
    dataset_name="research",
)
print(result.status, result.dataset_id)  # "completed", UUID

All cognee functions are async. Without dataset_name data goes to main_dataset. Needs LLM_API_KEY unless you use the GLiNER extractor (below).

Use it

Inputs

data accepts a string, a list of strings, file paths (absolute, file://, s3://), http(s) URLs, binary streams, or a list mixing them.

  • URLs are fetched and scraped (needs ALLOW_HTTP_REQUESTS=true, the default).
  • Folders are ingested file by file. A folder that looks like a code project, or a GitHub/GitLab URL, becomes one code repository (needs git on PATH).
  • Code files (.py, .ts, .go, …) go down the code-graph route: a deterministic graph, no LLM calls, searchable only with SearchType.CODE. To index a whole repository explicitly, pass content_type="code".
  • Databases and dlt sources: a SQL connection string, a dlt DltResource / DltSource, or a CSV. dlt is a core dependency, so no extra is needed (cognee[dlt] is an empty compatibility extra). Options: primary_key (default "id"), write_disposition ("replace" default, or "append"), query, max_rows_per_table.
  • Skill playbooks (SKILL.md files): content_type="skills"; ingests into the target dataset (default main_dataset), so pass dataset_name to keep skills in their own dataset.
Where the data goes
ArgumentWhat it does
dataset_name / dataset_idTarget dataset. dataset_id wins. A dataset is the unit of permissions and isolation.
node_set=["AI", "FinTech"]Tags the data so recall can filter to it later with recall(..., node_name=["AI"]).
session_id="chat_1"Writes to the fast session cache instead of the graph; improve() bridges it into the graph in the background. See the cognee-improve-sessions skill. Requires CACHING=true.
How the graph is built
ArgumentWhat it does
extractor"llm" or "gliner_demo" (alias "gliner"). Default is GRAPH_EXTRACTOR=auto: the LLM when an API key is configured, otherwise GLiNER.
graph_model=MyModelExtract into your own DataPoint model instead of the generic KnowledgeGraph. See the cognee-custom-graph-models skill.
custom_promptReplaces the entity-extraction prompt (ignored by GLiNER).
config={"ontology_config": {...}}Ground entities in an OWL ontology (below).
chunk_size, chunkerMax tokens per chunk (default: derived from the embedding and LLM limits) and the chunker class (default TextChunker).
preferred_loadersChoose a loader per file type (below).
self_improvementDefault True: runs improve() after the graph is built. Its outcome is on result.improve / result.improve_error; a failed improve never fails the remember.
run_in_background=TrueReturns immediately with status="running"; await result to wait.
Ontologies
python
from cognee.modules.ontology.rdf_xml.RDFLibOntologyResolver import RDFLibOntologyResolver

config = {
    "ontology_config": {
        "ontology_resolver": RDFLibOntologyResolver(ontology_file="./my.owl"),
        # "ontology_mode": "strict",   # drop entities with no ontology match
    }
}
await cognee.remember(texts, config=config)

Or set ONTOLOGY_FILE_PATH (plus ONTOLOGY_MODE, MATCHING_STRATEGY) in .env. annotate (default) only enriches; strict drops entities that match no ontology class or individual. It prunes only the graph, chunk text is still stored. Strict mode with an empty or missing ontology file is a hard error. Over HTTP, upload the ontology to /api/v1/ontologies and pass its ontology_key to POST /api/v1/remember. Example: examples/guides/ontology_quickstart.py.

Loaders

Each file is claimed by the first loader that accepts it. Default order: code, text, pypdf, image, audio, video, dlt_csv, csv, unstructured, advanced_pdf, docling. Names: text_loader, code_loader, csv_loader, dlt_csv_loader, pypdf_loader, image_loader, audio_loader, video_loader, unstructured_loader, advanced_pdf_loader, docling_loader, beautiful_soup_loader.

python
# Treat a code file as a plain document (chunking + LLM extraction):
await cognee.remember("./script.py", preferred_loaders={"text_loader": {}})

Office formats (DOCX, PPTX, …) need the docs (unstructured) or docling extra. A preferred loader that is not installed is skipped with only an info log, so check the extra is installed when a file comes out wrong.

Check the cost first

dry_run=True returns a token and cost estimate without ingesting anything or calling the LLM. It excludes the calls improve() makes. Not supported with GLiNER, sessions, or a remote instance.

dry_run="presort" on a folder returns a PresortReport (junk, duplicates, version candidates, possible personal data, proposed dataset groups). Apply it with await cognee.remember(report), or pass auto_apply=True.

Without an LLM: GLiNER

extractor="gliner" builds the graph and summaries with a local GLiNER2 model, with no LLM call (embeddings still run). Install pip install "cognee[gliner]"; the model (about 750 MB) downloads on first use. It cannot be combined with a custom graph_model, dry_run, session_id, or a remote instance.

For production: the open-source GLiNER extractor is a demo. cognee's enterprise GLiNER extraction is more accurate and covers more labels. The same goes for the Postgres graph adapter (postgres_demo). Contact social@cognee.ai.

Show full SKILL.md (417 more words)Show less

Pitfalls

  • Unknown keyword arguments raise. remember() forwards kwargs through a fixed allow-list and raises TypeError: Unexpected keyword arguments for anything else. These real options are not on it yet:

    OptionWorkaround through remember()
    ontology_file_pathconfig={"ontology_config": ...} or ONTOLOGY_FILE_PATH (above)
    functional_relationships, chunk_attachmentNone yet. Only cognee.cognify() accepts them.
    extraction_rulesPass it through the loader: preferred_loaders={"beautiful_soup_loader": {"extraction_rules": {...}}} (works in remember() and add()). Needs the scraping extra: without it the loader is not registered and the rules are silently ignored
    tavily_config, soup_crawler_configNot honoured by add() or remember(); only the cognee/tasks/web_scraper tasks use them
    column_value_columns (dlt)None yet. Only cognee.add() accepts it.

    If a user needs one with no workaround, say so plainly: the option exists on the lower-level add() / cognify() but not on remember() yet.

  • Changed files raise DocumentUpdateRequiredError. Re-remembering the same path (or the same filename for an upload) with different content is an update, not a new document. Use cognee.update(data_id=..., data=..., dataset_id=...), which re-extracts only the changed chunks and keeps the document's id. Identical content is a no-op.

  • content_type is strict. Only None, "skills", or "code". "code" rejects session_id and needs repository paths or git URLs; "skills" ingests into the target dataset like any other call (default main_dataset); pass dataset_name to keep skills in their own dataset.

  • Session mode needs CACHING=true, and extractor cannot be combined with session_id.

  • Remote mode. After cognee.serve(url), calls go to the server: extractor and session_ids raise, and other options the client does not forward (including graph_model, node_set, and ontology config) are dropped without an error.

  • Every remember runs improve() unless self_improvement=False or IMPROVE_AUTO_ENABLED=false. In scripts, call await cognee.wait_for_background_tasks() before exiting.

How it works

remember(data) runs add() (store raw data and create Data rows), then cognify() (classify documents, chunk, extract the graph and summaries, store in graph and vector DBs), then improve(). remember(data, session_id=...) writes to the session cache instead.

  • Entry point and kwarg routing: cognee/api/v1/remember/remember.py (RememberKwargs, _ADD_ONLY / _COGNIFY_ONLY / _SHARED)
  • Storage: cognee/api/v1/add/add.py, cognee/tasks/ingestion/ingest_data.py
  • Graph build: cognee/api/v1/cognify/cognify.py, cognee/tasks/graph/extract_graph_from_data.py, cognee/tasks/storage/add_data_points.py
  • Extractor choice: cognee/modules/cognify/config.py:resolve_extractor; GLiNER package: cognee/tasks/graph/gliner_demo/
  • Ontologies: cognee/modules/ontology/
  • Loaders: cognee/infrastructure/loaders/ (supported_loaders.py, LoaderEngine.py)
  • dlt: cognee/tasks/ingestion/resolve_dlt_sources.py

Examples in examples/guides/: simple_cognee_example.py, nodeset_grouping_example.py, ontology_quickstart.py, gliner_demo_llm_free_cognify.py, no_llm_remember_recall.py, temporal_recall.py, presort_downloads.py, web_url_content_ingestion_example.py, code_graph_example.py.

Extending it

  • New remember() option: add it to RememberKwargs and to the matching routing set in remember.py. An option on add()/cognify() that is not in a routing set raises TypeError from remember().
  • New loader: implement LoaderInterface (cognee/infrastructure/loaders/LoaderInterface.py), register it in supported_loaders.py (extras-gated loaders go under external/), and add it to the priority list in LoaderEngine.py if it should run by default.
  • New cognify task: see the cognee-custom-pipelines skill and cognee/tasks/README.md.

© topoteretes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/cognee-ingestion of topoteretes/cognee.

Open the folder on GitHubat commit 0ec7a9f

Compare with similar skills

Cognee Data Ingestion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cognee Data Ingestion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cognee Data Ingestion this skilltopoteretes/cognee32k—~2.6kAutomated safety check: NotesApache-2.0
Cortexdb Memory Hermesliliang-cn/cortexdb274—~1.7kAutomated safety check: PassMIT
Neo4j Graphrag Skillneo4j-contrib/neo4j-skills114—~4.2kAutomated safety check: NotesMIT
Hermes Memory Providersmnemosyne-oss/mnemosyne3.4k—~1.8kAutomated safety check: PassMIT
Neo4j Genai Plugin Skillneo4j-contrib/neo4j-skills114—~3kAutomated safety check: NotesMIT
Cortexdb Memory Openclawliliang-cn/cortexdb274—~1.6kAutomated safety check: PassMIT

Similar skills

  • Cortexdb Memory Hermes

    liliang-cn/cortexdb

    Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…

    274 GitHub stars~1.7k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check passed
  • Neo4j Graphrag Skill

    neo4j-contrib/neo4j-skills

    Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).

    114 GitHub stars~4.2k tokensUpdated yesterday
    Knowledge ManagementAuto-check: notes
  • Hermes Memory Providers

    mnemosyne-oss/mnemosyne

    Install and configure Mnemosyne as a Hermes Agent memory provider — local SQLite with vector search, episodic consolidation, and temporal knowledge graphs.

    3.4k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Neo4j Genai Plugin Skill

    neo4j-contrib/neo4j-skills

    Use Neo4j GenAI Plugin ai.text. An agent skill from neo4j-contrib/neo4j-skills.

    114 GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Cortexdb Memory Openclaw

    liliang-cn/cortexdb

    Give a Node.js agent (such as OpenClaw) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client npm package.

    274 GitHub stars~1.6k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check passed
  • Compact Memory Implementation

    simbajigege/book2skills

    A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.

    183 GitHub stars~2.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from topoteretes/cognee

All 19 skills in this repo
  • Cognee CLI Memory Commands

    topoteretes/cognee

    Drives cognee from the terminal with remember, recall, forget and improve memory commands, dataset and config management and database migrations.

    32k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Cognee Community Packages

    topoteretes/cognee

    Guide to using and contributing cognee community packages: database adapters, data-source connectors, custom tasks and retrievers, and Keywords AI observability.

    32k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Cognee Custom Graph Models

    topoteretes/cognee

    Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes.

    32k GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Cognee Custom Pipelines

    topoteretes/cognee

    Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph.

    32k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Cognee Docker Setup

    topoteretes/cognee

    Runs the Cognee AI memory platform in Docker, from a one-file prebuilt image to a full compose stack with UI, MCP server, Postgres and Neo4j.

    32k GitHub stars~901 tokensUpdated yesterday
    Auto-check: notes
  • Cognee Forget

    topoteretes/cognee

    Removes data from cognee memory with forget(), finding the right dataset and document first and choosing between one document, a dataset or only the graph and vector memory.

    32k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Cognee Data Ingestion

What does Cognee Data Ingestion do?

Explains how to load text, files, folders, URLs, repositories and databases into cognee memory with remember(), including datasets, extractors and tags. This skill covers cognee's remember() call, the single ingestion entry point that stores data, builds a knowledge graph from it and enriches the graph.md playbooks.

When should I use Cognee Data Ingestion?

Cognee Data Ingestion fits situations like: loading documents, folders or URLs into cognee memory; indexing a code repository or a SQL database through remember(); choosing between the LLM and GLiNER graph extractors; tagging data with node sets or splitting it across datasets.

How do I install Cognee Data Ingestion in Claude Code?

Run `npx skills add topoteretes/cognee --skill cognee-ingestion -a claude-code`. Or copy the skill folder (.agents/skills/cognee-ingestion in topoteretes/cognee) into .claude/skills/cognee-ingestion in your project. Claude Code loads it when a task matches its description.

How do I install Cognee Data Ingestion in Codex?

Run `npx skills add topoteretes/cognee --skill cognee-ingestion -a codex`. Or copy the skill folder (.agents/skills/cognee-ingestion in topoteretes/cognee) into .agents/skills/cognee-ingestion in your project. Codex loads it when a task matches its description.

Can I use Cognee Data Ingestion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add topoteretes/cognee --skill cognee-ingestion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cognee-ingestion, .gemini/skills/cognee-ingestion, .github/skills/cognee-ingestion and .opencode/skills/cognee-ingestion in your project.

What does Cognee Data Ingestion need to run?

Going by SKILL.md and its folder, Cognee Data Ingestion needs the command-line tools its instructions call (pip) and credentials named LLM_API_KEY. Our summary lists: Python with the cognee package; An LLM_API_KEY unless using the GLiNER extractor; git on PATH when ingesting code repositories.

Does Cognee Data Ingestion access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Cognee Data Ingestion safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Cognee Data Ingestion use?

Cognee Data Ingestion is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cognee Data Ingestion use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cognee Data Ingestion?

Skills that share tags, products or a category with Cognee Data Ingestion: Cortexdb Memory Hermes (liliang-cn/cortexdb, 274 stars), Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Hermes Memory Providers (mnemosyne-oss/mnemosyne, 3.4k stars) and Neo4j Genai Plugin Skill (neo4j-contrib/neo4j-skills, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cognee Data Ingestion?

topoteretes (a GitHub organization) maintains it in topoteretes/cognee, which has 31,919 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: topoteretes/cognee on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.