Agent skill

Cognee Performance Tuning

by topoteretes in topoteretes/cognee

Speeds up or throttles cognee ingestion: estimate cost with a dry run, tune batching and chunk size, set LLM rate limits and run work in the background.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Cognee Performance Tuning

skills CLI
$ npx skills add topoteretes/cognee --skill cognee-performance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install topoteretes/cognee cognee-performance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cognee-performance .claude/skills/cognee-performance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cognee-performance
GitHub stars
32k
Token cost
~2.5k tokens
SKILL.md length
1,167 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Speeds up or throttles cognee ingestion: estimate cost with a dry run, tune batching and chunk size, set LLM rate limits and run work in the background.

  • Works in 8 steps: Estimate before you ingest → Ingestion knobs → LLM throttling and models → …
  • Cognee ingestion is slow or the LLM provider keeps returning rate-limit errors
  • SKILL.md covers Use it, Pitfalls, How it works and Extending it
  • Calls gunicorn

What it does

The skill starts from one premise: the LLM is the bottleneck. Building the graph costs two LLM calls per chunk, one for graph extraction and one for summarization, so most tuning questions come down to how many of those calls run, how fast, and on which model. Database and embedding work is small in comparison.

Step one is an estimate before ingesting: calling remember on a folder with dry_run set to True returns per-stage token counts and approximate cost without making LLM calls. It covers graph extraction and summarization only, not improve(), embeddings or contradiction detection, and is unavailable with GLiNER, temporal_cognify or a remote instance. Ingestion settings are then explained: data_per_batch (default 20, a concurrency limit), chunks_per_batch (default 2000), chunk_size, run_in_background and self_improvement.

For throttling, it lists environment variables such as LLM_RATE_LIMIT_ENABLED, LLM_RATE_LIMIT_REQUESTS and LLM_RATE_LIMIT_INTERVAL, plus AUTO_RATE_LIMIT, which switches the limiter on for 15 minutes after a 429, 503, 529 or timeout. The description also mentions per-stage models, lower read latency and planning how to scale a deployment.

When your agent uses it

  • Cognee ingestion is slow or the LLM provider keeps returning rate-limit errors
  • Estimating the token cost of a large document set before running remember
  • Running cognee work in the background or limiting how many requests run at once

Example prompts

  • “Estimate what ingesting ./docs into cognee will cost before I run it.”
  • “Our cognify run keeps hitting 429 errors, so turn on rate limiting and lower the concurrency.”
  • “Make cognee ingestion faster for thousands of small markdown files.”

Requirements

  • A Python project that uses cognee
  • An LLM provider configured for cognee

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Estimate before you ingest
  2. Ingestion knobs
  3. LLM throttling and models
  4. Embeddings
  5. Skip the LLM for extraction
  6. Read latency
  7. Several datasets at once
  8. Scaling a deployment

What it can do on your machine

Read from SKILL.md and the folder at commit 0ec7a9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gunicorn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cognee Performance Tuning loads about 2.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 1,167 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from topoteretes/cognee at commit 0ec7a9f, republished under its Apache-2.0 licence (© topoteretes). 1,167 words, ~2,504 tokens.

Download SKILL.mdSave it as .claude/skills/cognee-performance/SKILL.md (or your agent's skills folder).
name
cognee-performance
description
Use when cognee is slow, expensive, or hitting rate limits — speeding up or throttling ingestion (remember/cognify batching, chunk size, LLM and embedding rate limits, per-stage models), cutting read latency, estimating cost before ingesting, running work in the background, or planning how to scale a deployment.

Tune cognee's performance

The LLM is the bottleneck. Building the graph costs two LLM calls per chunk (graph extraction and summarization), and almost every other tuning question comes down to how many of those calls run, how fast, and on which model. Database and embedding work is small next to that.

Use it

1. Estimate before you ingest
python
estimate = await cognee.remember("./docs", dry_run=True)
print(estimate)  # per-stage token counts and approximate cost; no LLM calls

The estimate covers graph extraction and summarization only, not improve(), embeddings, or contradiction detection. Not available with GLiNER or a remote instance.

2. Ingestion knobs

All are remember() arguments (and cognify() ones):

KnobDefaultWhat it controls
data_per_batch20How many documents are processed at the same time (a concurrency limit, not a batch). Lower it to reduce load; raise it for many small files.
chunks_per_batch2000 (env CHUNKS_PER_BATCH)How many chunks each extraction/storage call receives. Lower it to spread load and fail smaller. The DLT pipeline defaults to 100.
chunk_sizemin(embedding max tokens, LLM max tokens / 2); about 8191 tokens with default OpenAI modelsMax tokens per chunk. Fewer, bigger chunks mean fewer LLM calls; smaller chunks mean finer-grained graphs.
run_in_backgroundFalseReturn immediately (status="running"); await result later.
self_improvementTrueimprove() after the graph is built. False (or IMPROVE_AUTO_ENABLED=false) skips that extra work.

Worst-case concurrent LLM calls are about data_per_batch × min(chunks per document, chunks_per_batch) × 2, and there is no separate concurrency cap. The rate limiter below is the brake.

Several datasets in one cognify(datasets=[...]) call run one after another, not in parallel.

3. LLM throttling and models
Env varDefaultNotes
LLM_RATE_LIMIT_ENABLEDfalseTurn on a requests-per-interval limit
LLM_RATE_LIMIT_REQUESTS / LLM_RATE_LIMIT_INTERVAL60 / 60 s10 requests for local providers (Ollama, llama.cpp, LM Studio)
AUTO_RATE_LIMITtrueOn a 429/503/529 or timeout, switches the limiter on for 15 minutes (extended while errors continue)
LLM_EXTRACTION_MODEL / _PROVIDER / _ENDPOINT / _API_KEY / _API_VERSIONthe main modelUse a cheaper or faster model just for graph extraction
LLM_SUMMARIZATION_*, LLM_QUERY_*the main modelSame, for summaries and for answering queries
LLM_MAX_COMPLETION_TOKENS16384Also caps the default chunk size (half of it)
LLM_ARGS—Extra litellm arguments merged into every call, e.g. a request timeout

Structured-output calls retry with exponential backoff for at least two attempts and four minutes before failing; auth, not-found, and quota errors are not retried.

The rate limiter is built from the main model's settings. A stage routed to a local model does not get the local default of 10 requests; set LLM_RATE_LIMIT_REQUESTS yourself.

4. Embeddings
Env varDefault
EMBEDDING_BATCH_SIZE36 texts per request
EMBEDDING_MAX_CONCURRENT_DATA_POINTS150 (so 150 / 36 = 4 concurrent requests)
EMBEDDING_RATE_LIMIT_ENABLED / _REQUESTS / _INTERVALoff / 60 / 60 s

Embedding calls retry transient errors, rate-limit 429s included, with exponential jitter for up to 128 s (budget-exhaustion errors are terminal, not retried), but AUTO_RATE_LIMIT does not apply to them; enable EMBEDDING_RATE_LIMIT_ENABLED yourself for sustained load. Without any credentials cognee embeds locally with fastembed (BAAI/bge-small-en-v1.5), which runs on the CPU and blocks while it works.

5. Skip the LLM for extraction

remember(data, extractor="gliner") builds the graph and summaries with a local GLiNER model: no LLM calls, embeddings still run. The runtime installs on first use (GLINER_AUTO_INSTALL=false to disable; then pip install "cognee[gliner]"). See the cognee-ingestion skill for its limits.

6. Read latency
  • AUTO_FEEDBACK=false is the biggest win for chat-style use: by default every answered turn (the default session is used when no session_id is passed) makes one extra LLM call to analyze feedback. Keep CACHING=true so session memory still works.
  • Pick a type without an LLM when you need passages, not an answer: CHUNKS, CHUNKS_LEXICAL, SUMMARIES, CODE, SKILLS. Completion types make one LLM call; GRAPH_COMPLETION_COT, _DECOMPOSITION, _CONTEXT_EXTENSION, GRAPH_SUMMARY_COMPLETION and TEMPORAL make several; NATURAL_LANGUAGE makes one and retries (up to 3 attempts) only on an empty or failed query; FEELING_LUCKY adds one to pick the type.
  • Narrow the search: pass datasets=[...] (one search runs per dataset) and a smaller top_k (per dataset, default 15; HYBRID caps each lane at 10, so only values below 10 shrink its context).
  • SESSION_SEARCH_MODE=concurrent (default) overlaps the feedback analysis with the answer; sequential runs them back to back.
Show full SKILL.md (506 more words)Show less
7. Several datasets at once

With access control on (the default), each dataset has its own databases. DATASET_QUEUE_MAX_CONCURRENT (default 6, from DATABASE_MAX_LRU_CACHE_SIZE) caps how many datasets are processed at once in one process, and SUBPROCESS_IDLE_TTL_SECONDS (600) keeps idle database workers warm. DATASET_QUEUE_ENABLED (default true) enforces that cap, releases subprocess engines when a dataset's last scope exits (they stay warm for SUBPROCESS_IDLE_TTL_SECONDS and close at once only when it is 0) and pins in-use engines against eviction. Setting it false removes the cap rather than disabling parallelism, and risks file-lock leaks and engine eviction under parallel load, so keep it on.

8. Scaling a deployment

Open-source cognee scales vertically, in one process:

  • The API server runs one worker (gunicorn -w 1), and the improve and session locks, caches and semaphores are in-process only. More workers or replicas against the same data are not coordinated (the Helm chart README says to validate before scaling past one replica).
  • The default embedded databases (SQLite, LanceDB, Ladybug) live on local disk. For production, use Postgres + PGVector and a graph-native backend such as Neo4j.
  • distributed/deploy/ holds one-click deploy templates (Modal, Fly, Railway, Render, Daytona), not a distributed runner.

For production scale: horizontal scaling (distributed ingestion across workers), cognee's enterprise GLiNER extraction, and the production Postgres graph adapter are part of cognee's proprietary offering. Contact social@cognee.ai.

Pitfalls

  • CHUNK_SIZE, CHUNK_OVERLAP, CHUNK_STRATEGY (in .env.template) and cognee.config.set_chunk_size() etc. have no effect on ingestion. Pass chunk_size= to remember() instead.
  • LLM_RATE_LIMIT_TOKENS and EMBEDDING_RATE_LIMIT_TOKENS do nothing; only request-count limits are enforced.
  • FALLBACK_MODEL is used only when the main model rejects content on policy grounds, not when it is slow, down, or rate-limited.
  • data_per_batch=None crashes. Omit the argument to get the default.
  • A background cognify() is not tracked by cognee.wait_for_background_tasks() (a background remember() or improve() is). Await its pipeline status yourself.
  • Big default chunks with a small local embedder. The default chunk size follows the configured embedding limit (8191), even for fastembed's 512-token model. Pass a smaller chunk_size for local embedding models.
  • Benchmark with CACHING=true: with it off you measure cognee without its memory layer.

How it works

remember() → add() → cognify(): every document of the dataset is scheduled at once, data_per_batch of them run concurrently, and each runs the task chain (classify, chunk, extract graph + summarize, store). Tasks receive their input in batches of chunks_per_batch. extract_graph_and_summarize runs extraction and summarization together, each gathering one LLM call per chunk in the batch. Every LLM call passes through one process-wide rate limiter.

  • Cognify task list and defaults: cognee/api/v1/cognify/cognify.py
  • Concurrency (data_per_batch semaphore): cognee/modules/pipelines/operations/run_tasks.py
  • Task batching: cognee/modules/pipelines/tasks/task.py, run_tasks_base.py
  • LLM calls per chunk: cognee/tasks/graph/extract_graph_and_summarize.py
  • Chunk-size default: cognee/infrastructure/llm/utils.py:get_max_chunk_tokens
  • LLM config, rate limits, stage routing: cognee/infrastructure/llm/config.py, cognee/shared/rate_limiting.py, cognee/infrastructure/llm/overload_policy.py, cognee/infrastructure/llm/retry_config.py
  • Embedding config: cognee/infrastructure/databases/vector/embeddings/config.py
  • Dataset queue: cognee/infrastructure/databases/dataset_queue/queue.py
  • Cost estimate: cognee/modules/cognify/estimator.py

Extending it

  • Benchmarks: cognee/tests/performance/ (batch_add_cognify_test.py, locust_performance_analysis.py, and statistics_percentile/ for p50–p99 runs with mocked or replayed LLM calls). Measure before and after a change.
  • Retrieval quality and cost: cognee/eval_framework/ (python -m cognee.eval_framework; Modal runs cost money).
  • A new pipeline task that calls an LLM should keep needs_llm=True and go through LLMGateway, so it shares the rate limiter and retry policy.

© topoteretes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/cognee-performance of topoteretes/cognee.

Open the folder on GitHubat commit 0ec7a9f

Compare with similar skills

Cognee Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cognee Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cognee Performance Tuning this skilltopoteretes/cognee32k—~2.5kAutomated safety check: PassApache-2.0
Claude APIloulanyue/awesome-claude-notes2722 repos~2.1kAutomated safety check: PassMIT
Claude API Developmentwarpdotdev/warp65k3 repos~8.2kAutomated safety check: PassApache-2.0
Analyzing Claude Code Sessionsamd/gaia1.6k—~2.3kAutomated safety check: PassMIT
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT
Compact Memory Implementationsimbajigege/book2skills183—~2.5kAutomated safety check: PassMIT

Similar skills

  • Claude API

    loulanyue/awesome-claude-notes

    Anthropic Claude API patterns for Python and TypeScript. An agent skill from loulanyue/awesome-claude-notes.

    272 GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Claude API Development

    warpdotdev/warp

    Guides building, debugging and tuning apps on the Claude API and Anthropic SDK, including prompt caching, and migrating code between Claude model versions.

    65k GitHub starsUsed in 3 repos~8.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 4 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Compact Memory Implementation

    simbajigege/book2skills

    A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.

    183 GitHub stars~2.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Speculative Decoding

    Orchestra-Research/AI-Research-SKILLs

    Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.

    13k GitHub starsUsed in 2 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed

More from topoteretes/cognee

All 19 skills in this repo
  • Cognee CLI Memory Commands

    topoteretes/cognee

    Drives cognee from the terminal with remember, recall, forget and improve memory commands, dataset and config management and database migrations.

    32k GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • Cognee Community Packages

    topoteretes/cognee

    Guide to using and contributing cognee community packages: database adapters, data-source connectors, custom tasks and retrievers, and Keywords AI observability.

    32k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Cognee Custom Graph Models

    topoteretes/cognee

    Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes.

    32k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Cognee Custom Pipelines

    topoteretes/cognee

    Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph.

    32k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Cognee Docker Setup

    topoteretes/cognee

    Runs the Cognee AI memory platform in Docker, from a one-file prebuilt image to a full compose stack with UI, MCP server, Postgres and Neo4j.

    32k GitHub stars~901 tokensUpdated today
    Auto-check: notes
  • Cognee Forget

    topoteretes/cognee

    Removes data from cognee memory with forget(), finding the right dataset and document first and choosing between one document, a dataset or only the graph and vector memory.

    32k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Cognee Performance Tuning

What does Cognee Performance Tuning do?

Speeds up or throttles cognee ingestion: estimate cost with a dry run, tune batching and chunk size, set LLM rate limits and run work in the background. The skill starts from one premise: the LLM is the bottleneck. Building the graph costs two LLM calls per chunk, one for graph extraction and one for summarization, so most tuning questions come down to how many of those calls run, how fast, and on which model.

When should I use Cognee Performance Tuning?

Cognee Performance Tuning fits situations like: cognee ingestion is slow or the LLM provider keeps returning rate-limit errors; estimating the token cost of a large document set before running remember; running cognee work in the background or limiting how many requests run at once.

How do I install Cognee Performance Tuning in Claude Code?

Run `npx skills add topoteretes/cognee --skill cognee-performance -a claude-code`. Or copy the skill folder (.agents/skills/cognee-performance in topoteretes/cognee) into .claude/skills/cognee-performance in your project. Claude Code loads it when a task matches its description.

How do I install Cognee Performance Tuning in Codex?

Run `npx skills add topoteretes/cognee --skill cognee-performance -a codex`. Or copy the skill folder (.agents/skills/cognee-performance in topoteretes/cognee) into .agents/skills/cognee-performance in your project. Codex loads it when a task matches its description.

Can I use Cognee Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add topoteretes/cognee --skill cognee-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cognee-performance, .gemini/skills/cognee-performance, .github/skills/cognee-performance and .opencode/skills/cognee-performance in your project.

What does Cognee Performance Tuning need to run?

Going by SKILL.md and its folder, Cognee Performance Tuning needs the command-line tools its instructions call (gunicorn). Our summary lists: A Python project that uses cognee; An LLM provider configured for cognee.

Does Cognee Performance Tuning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cognee Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cognee Performance Tuning use?

Cognee Performance Tuning is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cognee Performance Tuning use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cognee Performance Tuning?

Skills that share tags, products or a category with Cognee Performance Tuning: Claude API (loulanyue/awesome-claude-notes, 272 stars), Claude API Development (warpdotdev/warp, 65k stars), Analyzing Claude Code Sessions (amd/gaia, 1.6k stars) and TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cognee Performance Tuning?

topoteretes (a GitHub organization) maintains it in topoteretes/cognee, which has 31,919 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: topoteretes/cognee on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.