Claude API
loulanyue/awesome-claude-notes
Anthropic Claude API patterns for Python and TypeScript. An agent skill from loulanyue/awesome-claude-notes.
Speeds up or throttles cognee ingestion: estimate cost with a dry run, tune batching and chunk size, set LLM rate limits and run work in the background.
$ npx skills add topoteretes/cognee --skill cognee-performance -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install topoteretes/cognee cognee-performance --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/cognee-performance .claude/skills/cognee-performance && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cognee-performance" agent skill from https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performance into .claude/skills/cognee-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cognee-performance", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performanceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add topoteretes/cognee --skill cognee-performance -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install topoteretes/cognee cognee-performance --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/cognee-performance .agents/skills/cognee-performance && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cognee-performance" agent skill from https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performance into .agents/skills/cognee-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cognee-performance", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add topoteretes/cognee --skill cognee-performance -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install topoteretes/cognee cognee-performance --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/cognee-performance .cursor/skills/cognee-performance && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cognee-performance" agent skill from https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performance into .cursor/skills/cognee-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cognee-performance", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/topoteretes/cognee.git --path .agents/skills/cognee-performance--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add topoteretes/cognee --skill cognee-performance -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install topoteretes/cognee cognee-performance --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/cognee-performance .gemini/skills/cognee-performance && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cognee-performance" agent skill from https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performance into .gemini/skills/cognee-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cognee-performance", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install topoteretes/cognee cognee-performanceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add topoteretes/cognee --skill cognee-performance -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/cognee-performance .github/skills/cognee-performance && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cognee-performance" agent skill from https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performance into .github/skills/cognee-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cognee-performance", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add topoteretes/cognee --skill cognee-performance -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install topoteretes/cognee cognee-performance --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/topoteretes/cognee.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/cognee-performance .opencode/skills/cognee-performance && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cognee-performance" agent skill from https://github.com/topoteretes/cognee/tree/main/.agents/skills/cognee-performance into .opencode/skills/cognee-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cognee-performance", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cognee-performanceSpeeds up or throttles cognee ingestion: estimate cost with a dry run, tune batching and chunk size, set LLM rate limits and run work in the background.
The skill starts from one premise: the LLM is the bottleneck. Building the graph costs two LLM calls per chunk, one for graph extraction and one for summarization, so most tuning questions come down to how many of those calls run, how fast, and on which model. Database and embedding work is small in comparison.
Step one is an estimate before ingesting: calling remember on a folder with dry_run set to True returns per-stage token counts and approximate cost without making LLM calls. It covers graph extraction and summarization only, not improve(), embeddings or contradiction detection, and is unavailable with GLiNER, temporal_cognify or a remote instance. Ingestion settings are then explained: data_per_batch (default 20, a concurrency limit), chunks_per_batch (default 2000), chunk_size, run_in_background and self_improvement.
For throttling, it lists environment variables such as LLM_RATE_LIMIT_ENABLED, LLM_RATE_LIMIT_REQUESTS and LLM_RATE_LIMIT_INTERVAL, plus AUTO_RATE_LIMIT, which switches the limiter on for 15 minutes after a 429, 503, 529 or timeout. The description also mentions per-stage models, lower read latency and planning how to scale a deployment.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0ec7a9f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gunicornFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cognee Performance Tuning loads about 2.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 1,167 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from topoteretes/cognee at commit 0ec7a9f, republished under its Apache-2.0 licence (© topoteretes). 1,167 words, ~2,504 tokens.
.claude/skills/cognee-performance/SKILL.md (or your agent's skills folder).The LLM is the bottleneck. Building the graph costs two LLM calls per chunk (graph extraction and summarization), and almost every other tuning question comes down to how many of those calls run, how fast, and on which model. Database and embedding work is small next to that.
estimate = await cognee.remember("./docs", dry_run=True)
print(estimate) # per-stage token counts and approximate cost; no LLM callsThe estimate covers graph extraction and summarization only, not
improve(), embeddings, or contradiction detection. Not available with
GLiNER or a remote instance.
All are remember() arguments (and cognify() ones):
| Knob | Default | What it controls |
|---|---|---|
data_per_batch | 20 | How many documents are processed at the same time (a concurrency limit, not a batch). Lower it to reduce load; raise it for many small files. |
chunks_per_batch | 2000 (env CHUNKS_PER_BATCH) | How many chunks each extraction/storage call receives. Lower it to spread load and fail smaller. The DLT pipeline defaults to 100. |
chunk_size | min(embedding max tokens, LLM max tokens / 2); about 8191 tokens with default OpenAI models | Max tokens per chunk. Fewer, bigger chunks mean fewer LLM calls; smaller chunks mean finer-grained graphs. |
run_in_background | False | Return immediately (status="running"); await result later. |
self_improvement | True | improve() after the graph is built. False (or IMPROVE_AUTO_ENABLED=false) skips that extra work. |
Worst-case concurrent LLM calls are about
data_per_batch × min(chunks per document, chunks_per_batch) × 2, and there
is no separate concurrency cap. The rate limiter below is the brake.
Several datasets in one cognify(datasets=[...]) call run one after
another, not in parallel.
| Env var | Default | Notes |
|---|---|---|
LLM_RATE_LIMIT_ENABLED | false | Turn on a requests-per-interval limit |
LLM_RATE_LIMIT_REQUESTS / LLM_RATE_LIMIT_INTERVAL | 60 / 60 s | 10 requests for local providers (Ollama, llama.cpp, LM Studio) |
AUTO_RATE_LIMIT | true | On a 429/503/529 or timeout, switches the limiter on for 15 minutes (extended while errors continue) |
LLM_EXTRACTION_MODEL / _PROVIDER / _ENDPOINT / _API_KEY / _API_VERSION | the main model | Use a cheaper or faster model just for graph extraction |
LLM_SUMMARIZATION_*, LLM_QUERY_* | the main model | Same, for summaries and for answering queries |
LLM_MAX_COMPLETION_TOKENS | 16384 | Also caps the default chunk size (half of it) |
LLM_ARGS | — | Extra litellm arguments merged into every call, e.g. a request timeout |
Structured-output calls retry with exponential backoff for at least two attempts and four minutes before failing; auth, not-found, and quota errors are not retried.
The rate limiter is built from the main model's settings. A stage routed to
a local model does not get the local default of 10 requests; set
LLM_RATE_LIMIT_REQUESTS yourself.
| Env var | Default |
|---|---|
EMBEDDING_BATCH_SIZE | 36 texts per request |
EMBEDDING_MAX_CONCURRENT_DATA_POINTS | 150 (so 150 / 36 = 4 concurrent requests) |
EMBEDDING_RATE_LIMIT_ENABLED / _REQUESTS / _INTERVAL | off / 60 / 60 s |
Embedding calls retry transient errors, rate-limit 429s included, with
exponential jitter for up to 128 s (budget-exhaustion errors are terminal,
not retried), but AUTO_RATE_LIMIT does not apply to them; enable
EMBEDDING_RATE_LIMIT_ENABLED yourself for sustained load. Without any
credentials cognee embeds locally with fastembed (BAAI/bge-small-en-v1.5),
which runs on the CPU and blocks while it works.
remember(data, extractor="gliner") builds the graph and summaries with a
local GLiNER model: no LLM calls, embeddings still run. The runtime installs
on first use (GLINER_AUTO_INSTALL=false to disable; then pip install "cognee[gliner]"). See the cognee-ingestion skill for its limits.
AUTO_FEEDBACK=false is the biggest win for chat-style use: by default
every answered turn (the default session is used when no session_id is
passed) makes one extra LLM call to analyze feedback. Keep CACHING=true so session memory still works.CHUNKS, CHUNKS_LEXICAL, SUMMARIES, CODE, SKILLS. Completion types
make one LLM call; GRAPH_COMPLETION_COT, _DECOMPOSITION,
_CONTEXT_EXTENSION, GRAPH_SUMMARY_COMPLETION and TEMPORAL make
several; NATURAL_LANGUAGE makes one and retries (up to 3 attempts) only
on an empty or failed query; FEELING_LUCKY adds one to pick the type.datasets=[...] (one search runs per dataset)
and a smaller top_k (per dataset, default 15; HYBRID caps each lane at
10, so only values below 10 shrink its context).SESSION_SEARCH_MODE=concurrent (default) overlaps the feedback analysis
with the answer; sequential runs them back to back.With access control on (the default), each dataset has its own databases.
DATASET_QUEUE_MAX_CONCURRENT (default 6, from DATABASE_MAX_LRU_CACHE_SIZE)
caps how many datasets are processed at once in one process, and
SUBPROCESS_IDLE_TTL_SECONDS (600) keeps idle database workers warm.
DATASET_QUEUE_ENABLED (default true) enforces that cap, releases
subprocess engines when a dataset's last scope exits (they stay warm for
SUBPROCESS_IDLE_TTL_SECONDS and close at once only when it is 0) and pins
in-use engines against eviction.
Setting it false removes the cap rather than disabling parallelism, and
risks file-lock leaks and engine eviction under parallel load, so keep it
on.
Open-source cognee scales vertically, in one process:
gunicorn -w 1), and the improve and
session locks, caches and semaphores are in-process only. More workers or
replicas against the same data are not coordinated (the Helm chart README
says to validate before scaling past one replica).distributed/deploy/ holds one-click deploy templates (Modal, Fly,
Railway, Render, Daytona), not a distributed runner.For production scale: horizontal scaling (distributed ingestion across workers), cognee's enterprise GLiNER extraction, and the production Postgres graph adapter are part of cognee's proprietary offering. Contact social@cognee.ai.
CHUNK_SIZE, CHUNK_OVERLAP, CHUNK_STRATEGY (in .env.template) and
cognee.config.set_chunk_size() etc. have no effect on ingestion. Pass
chunk_size= to remember() instead.LLM_RATE_LIMIT_TOKENS and EMBEDDING_RATE_LIMIT_TOKENS do nothing;
only request-count limits are enforced.FALLBACK_MODEL is used only when the main model rejects content on
policy grounds, not when it is slow, down, or rate-limited.data_per_batch=None crashes. Omit the argument to get the default.cognify() is not tracked by
cognee.wait_for_background_tasks() (a background remember() or
improve() is). Await its pipeline status yourself.chunk_size for local embedding models.CACHING=true: with it off you measure cognee without its
memory layer.remember() → add() → cognify(): every document of the dataset is
scheduled at once, data_per_batch of them run concurrently, and each runs
the task chain (classify, chunk, extract graph + summarize, store). Tasks
receive their input in batches of chunks_per_batch.
extract_graph_and_summarize runs extraction and summarization together,
each gathering one LLM call per chunk in the batch. Every LLM call passes
through one process-wide rate limiter.
cognee/api/v1/cognify/cognify.pydata_per_batch semaphore): cognee/modules/pipelines/operations/run_tasks.pycognee/modules/pipelines/tasks/task.py, run_tasks_base.pycognee/tasks/graph/extract_graph_and_summarize.pycognee/infrastructure/llm/utils.py:get_max_chunk_tokenscognee/infrastructure/llm/config.py,
cognee/shared/rate_limiting.py, cognee/infrastructure/llm/overload_policy.py,
cognee/infrastructure/llm/retry_config.pycognee/infrastructure/databases/vector/embeddings/config.pycognee/infrastructure/databases/dataset_queue/queue.pycognee/modules/cognify/estimator.pycognee/tests/performance/ (batch_add_cognify_test.py,
locust_performance_analysis.py, and statistics_percentile/ for p50–p99
runs with mocked or replayed LLM calls). Measure before and after a change.cognee/eval_framework/ (python -m cognee.eval_framework; Modal runs cost money).needs_llm=True and go
through LLMGateway, so it shares the rate limiter and retry policy.© topoteretes, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/cognee-performance of topoteretes/cognee.
Open the folder on GitHubat commit 0ec7a9f
Cognee Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cognee Performance Tuning this skilltopoteretes/cognee | 32k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Claude APIloulanyue/awesome-claude-notes | 272 | 2 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Claude API Developmentwarpdotdev/warp | 65k | 3 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 | |
| Analyzing Claude Code Sessionsamd/gaia | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT | |
| TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Compact Memory Implementationsimbajigege/book2skills | 183 | — | ~2.5k | Automated safety check: Pass | MIT |
loulanyue/awesome-claude-notes
Anthropic Claude API patterns for Python and TypeScript. An agent skill from loulanyue/awesome-claude-notes.
warpdotdev/warp
Guides building, debugging and tuning apps on the Claude API and Anthropic SDK, including prompt caching, and migrating code between Claude model versions.
amd/gaia
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
Orchestra-Research/AI-Research-SKILLs
Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.
simbajigege/book2skills
A developer guide to adding compact memory to an agent: when to trigger compaction, how to fork a compactor sub-agent, what the summary holds, and how to restore it.
Orchestra-Research/AI-Research-SKILLs
Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.
topoteretes/cognee
Drives cognee from the terminal with remember, recall, forget and improve memory commands, dataset and config management and database migrations.
topoteretes/cognee
Guide to using and contributing cognee community packages: database adapters, data-source connectors, custom tasks and retrievers, and Keywords AI observability.
topoteretes/cognee
Defines the shape of cognee's knowledge graph with graph_model: DataPoint node classes, identity and index fields, typed edges and fixes for duplicated nodes.
topoteretes/cognee
Shows how to write custom cognee tasks, chain them into pipelines, store custom DataPoints and run enrichment over the existing graph.
topoteretes/cognee
Runs the Cognee AI memory platform in Docker, from a one-file prebuilt image to a full compose stack with UI, MCP server, Postgres and Neo4j.
topoteretes/cognee
Removes data from cognee memory with forget(), finding the right dataset and document first and choosing between one document, a dataset or only the graph and vector memory.
Works with
Categories
Speeds up or throttles cognee ingestion: estimate cost with a dry run, tune batching and chunk size, set LLM rate limits and run work in the background. The skill starts from one premise: the LLM is the bottleneck. Building the graph costs two LLM calls per chunk, one for graph extraction and one for summarization, so most tuning questions come down to how many of those calls run, how fast, and on which model.
Cognee Performance Tuning fits situations like: cognee ingestion is slow or the LLM provider keeps returning rate-limit errors; estimating the token cost of a large document set before running remember; running cognee work in the background or limiting how many requests run at once.
Run `npx skills add topoteretes/cognee --skill cognee-performance -a claude-code`. Or copy the skill folder (.agents/skills/cognee-performance in topoteretes/cognee) into .claude/skills/cognee-performance in your project. Claude Code loads it when a task matches its description.
Run `npx skills add topoteretes/cognee --skill cognee-performance -a codex`. Or copy the skill folder (.agents/skills/cognee-performance in topoteretes/cognee) into .agents/skills/cognee-performance in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add topoteretes/cognee --skill cognee-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cognee-performance, .gemini/skills/cognee-performance, .github/skills/cognee-performance and .opencode/skills/cognee-performance in your project.
Going by SKILL.md and its folder, Cognee Performance Tuning needs the command-line tools its instructions call (gunicorn). Our summary lists: A Python project that uses cognee; An LLM provider configured for cognee.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cognee Performance Tuning is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cognee Performance Tuning: Claude API (loulanyue/awesome-claude-notes, 272 stars), Claude API Development (warpdotdev/warp, 65k stars), Analyzing Claude Code Sessions (amd/gaia, 1.6k stars) and TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
topoteretes (a GitHub organization) maintains it in topoteretes/cognee, which has 31,919 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.
Source: topoteretes/cognee on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.