Agent skill

Kapso

by Leeroo-AI in Leeroo-AI/kapso

Operate Kapso (PyPI leeroo-kapso), the long-running agents that optimize AI and Data systems, from a coding session — verify the install with kapso doctor, launch and follow an evolve campaign…

MITAuto-check: notesKnowledge Management

Install Kapso

skills CLI
$ npx skills add Leeroo-AI/kapso --skill kapso -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Leeroo-AI/kapso kapso --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Leeroo-AI/kapso.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/codex/kapso .claude/skills/kapso && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kapso
GitHub stars
121
Token cost
~3.6k tokens
SKILL.md length
1,655 words
Files
2
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Operate Kapso (PyPI leeroo-kapso), the long-running agents that optimize AI and Data systems, from a coding session — verify the install with kapso doctor, launch and follow an evolve campaign…

  • The user mentions Kapso
  • SKILL.md covers Facts that prevent the common…, Evolve: run a campaign, Inbox: the campaign is waiting… and Resume an interrupted campaign, plus 7 more sections
  • Calls python, docker and modal; reaches docs.leeroo.com and github.com; needs OPENAI_API_KEY and LEEROOPEDIA_API_KEY
  • Asks to push a measurable metric (accuracy

What it does

Kapso is an agent skill from Leeroo-AI/kapso. Operate Kapso (PyPI leeroo-kapso), the long-running agents that optimize AI and Data systems, from a coding session — verify the install with kapso doctor, launch and follow an evolve campaign toward a scored goal, answer a campaign that is WAITING ON YOU, resume an interrupted run, learn from a finished campaign into the lesson bank, ingest repos and research into the knowledge graph, and deploy the winner. Use when the user mentions Kapso or kapso evolve, learn, research, deploy, doctor, watch, inbox, bank, or…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Knowledge Management, covering Autonomous loops, Knowledge graphs and Email management. The repository describes itself as: Kapso by Leeroo: Long-running agents that optimize AI and Data systems, and learn from every experience. 1 open-source on MLE-Bench; ALE-Bench; RelBench. The licence is MIT.

When your agent uses it

  • The user mentions Kapso
  • Asks to push a measurable metric (accuracy
  • Score) in a project where kapso is installed
  • Ordinary code edits

Example prompts

  • “/kapso”

Requirements

  • Python 3
  • Docker
  • A credential in OPENAI_API_KEY
  • A credential in LEEROOPEDIA_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 6871774. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • docker
    • modal
    • pip
    • bash
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.leeroo.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • LEEROOPEDIA_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kapso loads about 3.6k tokens when it runs. Until then it costs about 183 tokens; SKILL.md has 1,655 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~183
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:22
    - `LEEROOPEDIA_API_KEY` in `.env` connects every campaign session to Leeroopedia,
  • NoteMentions a .env fileSKILL.md:26
    - Secrets come from `.env` in the directory you run from (`find_dotenv(usecwd=True)`);
  • NoteMentions a .env fileSKILL.md:107
    everything in the directory, `.env` included.
  • NoteMentions a .env fileSKILL.md:134
    box reply ./campaign 1 "added the key to .env"
  • NoteMentions a .env fileSKILL.md:140
    `.env`) and reply with a note.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Leeroo-AI/kapso at commit 6871774, republished under its MIT licence (© Leeroo-AI). 1,655 words, ~3,610 tokens.

Download SKILL.mdSave it as .claude/skills/kapso/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
kapso
description
Operate Kapso (PyPI leeroo-kapso), the long-running agents that optimize AI and Data systems, from a coding session — verify the install with kapso doctor, launch and follow an evolve campaign toward a scored goal, answer a campaign that is WAITING ON YOU, resume an interrupted run, learn from a finished campaign into the lesson bank, ingest repos and research into the knowledge graph, and deploy the winner. Use when the user mentions Kapso or kapso evolve, learn, research, deploy, doctor, watch, inbox, bank, or asks to push a measurable metric (accuracy, latency, score) in a project where kapso is installed. Do not use for ordinary code edits, or for hand-tuning a metric when Kapso is neither installed nor asked for.

Kapso

Kapso runs experiment campaigns: propose a candidate, implement it on its own git branch, score it with the judge, keep the best, repeat. Everything below is 0.6.x. Prefer the facts here and the docs links at the end over reading the package source; the source is large and a session that greps it runs out of turns before answering.

Facts that prevent the common mistakes

  • Install with pip install leeroo-kapso. The PyPI package named kapso is an unrelated WhatsApp tool that shadows the kapso command. Python 3.10+.
  • Every model call runs through a coding-agent CLI that must be logged in: claude (ideation, implementation, judging) and codex (research, utilities). There is no API-key fallback. OPENAI_API_KEY is optional and only ever embeds text: the knowledge graph, and semantic search over past experiments once models.embedding is set in the config (off in the shipped config).
  • LEEROOPEDIA_API_KEY in .env connects every campaign session to Leeroopedia, a hosted knowledge base of ML and AI frameworks (nothing to install; a new account at https://app.leeroopedia.com/dashboard has $20 of free credit). Without it the campaign still runs and doctor shows the row as [-- ].
  • Secrets come from .env in the directory you run from (find_dotenv(usecwd=True)); a shell export works only if it is actually exported. Config never holds secrets.
  • All knobs live in one YAML. To change anything, copy the packaged file, edit, and pass it with --config / config_path=. Never set Kapso behaviour through environment variables.
  • Run kapso doctor (or kapso doctor evolve|learn|research|learn_knowledge|deploy) before any run. Every required row must read [OK ]; [-- ] rows are optional. Add --models to live-probe every configured model with one token each; a model the login cannot serve fails here in seconds instead of hours in. A usage cap on a model it can serve is not visible to the probe.
  • A campaign takes tens of minutes to hours. Launch it in the background with its output in a log file, check once that the process is alive, and hand off in your reply right away: the goal you passed, the metric and target you assumed and that the user's own can replace them, the kapso watch <campaign> --follow command, and that WAITING ON YOU is a pause answered through kapso inbox. The status file appears only after the seed copy and repo-memory bootstrap, usually a few minutes in, so do not wait for it, poll the log, or arm monitors unless asked. Never run kapso evolve in a foreground tool call.
  • A run that ends with WAITING ON YOU has paused, not failed (exit code 0). Do not restart it; reply through the inbox (below).

Evolve: run a campaign

The goal text

The judge stops the campaign only when the goal reads as fully achieved, and it scores every experiment from what the evaluation prints. So the goal has to name a success metric and a number, and it helps to name the judge: "accuracy above 0.85 as measured by eval/evaluate.py". Both are the user's to give and optional; whatever the user gives goes into the goal verbatim.

First look for an evaluation the repo already has: a scoring script, a test suite, a benchmark command. Then:

  • The repo has one but the request names no metric or target: run it once for the baseline, write the goal with that metric and a target you propose, and tell the user in one line what you assumed and that their own success metric goes straight into the goal. Leave headroom below any ceiling you can see; a target the data cannot reach ends in the inbox, not in a result.
  • The repo has none and the request names none: your whole reply is one question, one line, asking whether they have a script or command that scores this and what "good" means; nothing is launched, measured or written before the answer. A script they have goes in with --eval-dir, protected. If they answer that they have none or do not know, launch with the best metric you can state in the goal and no --eval-dir: the campaign builds its own evaluation in kapso_evaluation/ from the goal, and the judge checks that it is fair. Never write an evaluator on the user's behalf, and never ask twice.
  • Put rules in the goal, not methods. The judge enforces data rules and prohibitions ("do not modify eval/evaluate.py", "do not fit on data/test.csv") as invariants for the whole campaign, and would freeze a method preference the same way.
Launch
bash
kapso doctor evolve
setsid -f kapso evolve \
  --goal "Get the test accuracy of the churn model in train.py above 0.85 as measured by eval/evaluate.py. The evaluator must not be modified." \
  --initial-repo . \
  --eval-dir eval \
  --data-dir data \
  --output ../churn-campaign \
  --time-budget-minutes 60 --iterations 10 \
  > ../churn-campaign.log 2>&1 < /dev/null
kapso watch ../churn-campaign --follow
  • Codex ends a plain nohup … & child when the shell call returns; setsid -f starts Kapso in its own session, so it outlives the call and the session (where setsid is missing: python3 -c 'import subprocess, sys; subprocess.Popen(sys.argv[1:], start_new_session=True)' kapso evolve …). The launch line is a tool call of its own, nothing chained in front of it.
  • --initial-repo <path|github url> seeds the campaign; a non-git directory is copied and committed as the baseline. The original is never modified. Every experiment lands on branch generic_exp_N inside --output, an empty directory outside the repo being seeded (a sibling), so the seed copy never contains the campaign.
  • --eval-dir is copied to kapso_evaluation/ and integrity-protected: any candidate that edits it is rejected and unscored. Use it whenever the user has a judge. --data-dir is copied to kapso_datasets/. The seed copy includes everything in the directory, .env included.
  • Bound the run: --iterations (default 10), --time-budget-minutes (durable across resumes), --cost-budget (best effort; codex sessions report no cost).
  • -m GENERIC (default, knowledge search on) or -m MINIMAL; -a picks the coding agent (claude_code default, codex, gemini, openhands, oss_claude_code).
  • The summary block ends COMPLETED with score and stop reason, or WAITING ON YOU.

Python, same thing:

python
from kapso import Kapso
solution = Kapso().evolve(
    goal="...", initial_repo=".", eval_dir="eval", data_dir="data",
    output_path="./campaign", time_budget_minutes=60,
)
print(solution.explain())   # .final_score .succeeded .code_path .requests

evolve() blocks for the whole campaign, so run scripts with setsid -f as well.

Inbox: the campaign is waiting on you

A session that needs something only a person can provide (a credential, access, a file) records a request and the campaign pauses.

bash
kapso inbox ./campaign                 # the open requests: key, hit, tried, fix, next
kapso inbox reply ./campaign 1 "added the key to .env"
  • The id may be omitted when one request is open; an empty note means done.
  • The reply resumes the very session that asked, in the foreground of that command.
  • A credential-shaped reply is refused: put the value where fix says (usually .env) and reply with a note.
Show full SKILL.md (627 more words)Show less

Resume an interrupted campaign

bash
kapso watch ../churn-campaign                  # DEAD (pid gone) or STALLED means it died
setsid -f kapso evolve --output ../churn-campaign --resume > ../churn-campaign.log 2>&1 < /dev/null

A resume is a launch: run it in the background the same way, check once that the process is alive, and hand off with kapso watch … --follow. Do it in the same turn you diagnose the death in; a reply that only announces a resume leaves the campaign dead. Other kapso evolve processes on the machine belong to other campaigns; never kill one. The checkpoint in .kapso/run_state.json carries the goal and the search state; the launch record .kapso/launch.json carries every flag of the launch, and a resume reads both, so pass only what you mean to change (a bigger --time-budget-minutes, say). A resume that changes the mode, coding agent, eval-dir, config or knowledge index is refused by name; start a new campaign in a new output path instead. A campaign paused by the inbox resumes through kapso inbox reply, not --resume.

Learn: bank what a campaign taught

learn() mines one finished campaign into the lesson bank, a local git repo of evidence-priced cards. It runs crews for an hour or more and needs both CLIs.

python
from kapso import Kapso
k = Kapso()
lesson = k.learn("./campaign")      # or learn(solution); a trajectory id also works
print(lesson.explain())             # cards created/updated, admitted, report paths
print(k.memory.explain())           # bank head, active cards, serving flag
  • The bank is created on first use at learning.bank.local_path (~/.kapso/bank.git). Share it with kapso bank connect <git-url> or kapso bank create org/name; after that every learn() pushes.
  • Serving is off by default. For the next campaign to read the bank, set learning.serving.enabled: true in your config and pass it. Without that, banked cards have no effect on evolve().
  • Follow a learn with kapso watch learning/status.
  • The kapso learn ... CLI subcommands (import, mine, grade, update, develop, codify, gauntlet, behave) are the learner-development regime, not the everyday path. From a shell the everyday path is the two Python lines above.

learn_knowledge: import outside knowledge

Repos and research become wiki pages in a knowledge graph. This is a different memory from the lesson bank.

python
from kapso import Kapso, Source
k = Kapso()
findings = k.research("gradient boosting for churn", mode=["idea", "implementation"], depth="deep")
k.learn_knowledge(Source.Repo("https://github.com/org/repo"), findings, wiki_dir="data/wikis")
index = k.index_kg(wiki_dir="data/wikis", save_to="data/indexes/churn.index")
Kapso(kg_index=index).evolve(goal="...")

Needs Weaviate on 8080 and Neo4j on 7687 (from a source checkout: bash scripts/start_infra.sh). Ingest time scales with the material, not with depth, and can run for hours.

Research

bash
kapso research --objective "..." --mode idea --mode implementation --depth deep -o findings.json

mode is idea, implementation, or study (repeat the flag); depth is light or deep. In Python research(objective, mode=[...], depth=...) with keyword-only mode and depth; pass results into a campaign with evolve(context=[findings.to_string()]). Runs on codex with web search.

Deploy

Deploy runs one claude session that adapts a copy of the solution (<path>_adapted_<strategy>) for the target and returns when it is ready. It takes a few minutes, not hours: run it in the foreground with a long tool timeout and use the result, rather than backgrounding it and polling.

bash
kapso doctor deploy
kapso deploy --solution-path ../churn-campaign --strategy local   # auto|local|docker|modal|bentoml|langgraph

To call what was deployed, use the Python API, which hands back the running software:

python
from kapso import Kapso, DeployStrategy, SolutionResult
solution = SolutionResult(goal="churn model", code_path="../churn-campaign")  # or the object evolve() returned
software = Kapso().deploy(solution, strategy=DeployStrategy.LOCAL)
print(software.run({"tenure_months": 3, "monthly_spend": 92.0, "support_tickets": 4, "logins_per_week": 1.0, "plan": "basic"}))
software.stop()

The original solution is untouched. AUTO lets a selector pick the target.

Models and config

bash
python -c "from kapso.kapso import DEFAULT_CONFIG_PATH; print(DEFAULT_CONFIG_PATH)"
cp "$(python -c 'from kapso.kapso import DEFAULT_CONFIG_PATH; print(DEFAULT_CONFIG_PATH)')" kapso-config.yaml
# edit, then:
kapso doctor learn --models --config kapso-config.yaml

Shipped models: campaigns on claude-opus-5, learning crews on claude-fable-5, codex roles on gpt-5.6-sol. Swap a model by editing every occurrence in the block you care about (learning.* for the crews, modes.<MODE>.* for campaigns). The timeout_minutes caps were calibrated on the shipped models; tell the user a slower model may need them raised, but change only what was asked. Kapso(config_path="kapso-config.yaml") and --config on every CLI verb select your file.

When Kapso is the wrong tool

A one-file change with an obvious fix is faster by hand. Reach for a campaign when there is a judge to beat, many candidates worth trying, and the user accepts an unattended run of an hour or more on their coding-agent subscription. Say which you are doing and why; never launch a campaign the user did not ask for.

Docs, served as Markdown

The whole site is also an MCP server: claude mcp add --transport http kapso-docs https://docs.leeroo.com/mcp.

© Leeroo-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/codex/kapso of Leeroo-AI/kapso.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 6871774

Compare with similar skills

Kapso next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kapso compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kapso this skillLeeroo-AI/kapso121—~3.6kAutomated safety check: NotesMIT
Ontology1mancompany/OneManCompany4422 repos~1.5kAutomated safety check: PassApache-2.0
DreamingSignet-AI/signetai305—~2.9kAutomated safety check: PassCustom licence
Memory Tasksbasicmachines-co/basic-memory4.1k1 repos~1.4kAutomated safety check: PassAGPL-3.0
Lat Md Knowledge Graphstevesolun/ctx588—~466Automated safety check: PassMIT
Basic Memorybasicmachines-co/basic-memory4.1k—~2.9kAutomated safety check: PassAGPL-3.0

Similar skills

  • Ontology

    1mancompany/OneManCompany

    Typed knowledge graph for structured agent memory and composable skills.

    442 GitHub starsUsed in 2 repos~1.5k tokens
    Knowledge ManagementAuto-check passed
  • Dreaming

    Signet-AI/signetai

    Maintain Signet's living ontology and memory substrate from transcripts, memory artifacts, source artifacts, notes, summaries, and imported records.

    305 GitHub stars~2.9k tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Memory Tasks

    basicmachines-co/basic-memory

    Task management via Basic Memory schemas: create, track, and resume structured tasks that survive context compaction.

    4.1k GitHub starsUsed in 1 repo~1.4k tokens
    Knowledge ManagementAuto-check passed
  • Design or audit a repo-local markdown knowledge graph with wiki links, source-code backlinks, drift checks, and searchable sections.

    588 GitHub stars~466 tokensUpdated 6 days ago
    Knowledge ManagementAuto-check passed
  • Basic Memory

    basicmachines-co/basic-memory

    Use the Basic Memory knowledge graph for persistent memory across sessions.

    4.1k GitHub stars~2.9k tokensUpdated today
    Knowledge ManagementAuto-check passed
  • Ade Perf Boot

    arul28/ADE

    Performance patterns discovered for ADE's cold launch and "main screen" surfaces — welcome / project picker, recent projects list, project open flow, remote runtime connect, iOS pairing.

    114 GitHub stars~707 tokensUpdated today
    Knowledge ManagementAuto-check passed

More from Leeroo-AI/kapso

  • Kapso

    Leeroo-AI/kapso

    Optimize code using KAPSO (Knowledge-Grounded Optimization).

    121 GitHub stars~642 tokensUpdated today
    Auto-check passed

Questions about Kapso

What does Kapso do?

Operate Kapso (PyPI leeroo-kapso), the long-running agents that optimize AI and Data systems, from a coding session — verify the install with kapso doctor, launch and follow an evolve campaign…. Kapso is an agent skill from Leeroo-AI/kapso. Operate Kapso (PyPI leeroo-kapso), the long-running agents that optimize AI and Data systems, from a coding session — verify the install with kapso doctor, launch and follow an evolve campaign toward a scored goal, answer a campaign that is WAITING ON YOU, resume an interrupted run, learn from a finished campaign into the lesson bank, ingest repos and research into the knowledge graph, and deploy the winner.

When should I use Kapso?

Kapso fits situations like: the user mentions Kapso; asks to push a measurable metric (accuracy; score) in a project where kapso is installed; ordinary code edits.

How do I install Kapso in Claude Code?

Run `npx skills add Leeroo-AI/kapso --skill kapso -a claude-code`. Or copy the skill folder (skills/codex/kapso in Leeroo-AI/kapso) into .claude/skills/kapso in your project. Claude Code loads it when a task matches its description.

How do I install Kapso in Codex?

Run `npx skills add Leeroo-AI/kapso --skill kapso -a codex`. Or copy the skill folder (skills/codex/kapso in Leeroo-AI/kapso) into .agents/skills/kapso in your project. Codex loads it when a task matches its description.

Can I use Kapso in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Leeroo-AI/kapso --skill kapso -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kapso, .gemini/skills/kapso, .github/skills/kapso and .opencode/skills/kapso in your project.

What does Kapso need to run?

Going by SKILL.md and its folder, Kapso needs the command-line tools its instructions call (python, docker, modal, pip, bash and claude) and credentials named OPENAI_API_KEY and LEEROOPEDIA_API_KEY. Our summary lists: Python 3; Docker; A credential in OPENAI_API_KEY; A credential in LEEROOPEDIA_API_KEY.

Does Kapso access the network?

SKILL.md names 2 domains. In commands or code: docs.leeroo.com and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Kapso safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Kapso use?

Kapso is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kapso use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kapso?

Skills that share tags, products or a category with Kapso: Ontology (1mancompany/OneManCompany, 442 stars), Dreaming (Signet-AI/signetai, 305 stars), Memory Tasks (basicmachines-co/basic-memory, 4.1k stars) and Lat Md Knowledge Graph (stevesolun/ctx, 588 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kapso?

Leeroo-AI (a GitHub organization) maintains it in Leeroo-AI/kapso, which has 121 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 9, 2026.

Source: Leeroo-AI/kapso on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.