Agent skill

Kayba Pipeline

by kayba-ai in kayba-ai/agentic-context-engine

End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.

Apache-2.0Auto-check passedAgent Workflows

Install Kayba Pipeline

skills CLI
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kayba-ai/agentic-context-engine kayba-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kayba-pipeline .claude/skills/kayba-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kayba-pipeline
GitHub stars
2.6k
Token cost
~1.4k tokens
SKILL.md length
522 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.

  • Works in 5 steps: sequential → sequential → sequential → …
  • The user says run the pipeline
  • SKILL.md covers Inputs, Pipeline overview, Orchestration instructions and Error handling, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Kayba Pipeline is an agent skill from kayba-ai/agentic-context-engine. End-to-end agent evaluation and improvement pipeline. Takes a traces folder and optional HITL flag, then orchestrates sub-agents through 7 stages — each stage is its own skill invoked by a dedicated sub-agent. Trigger when the user says "run the pipeline", "kayba pipeline", "evaluate and fix", "full eval", "analyze traces and fix", or provides a traces folder with intent to improve their agent.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Subagents and Agent evaluation and testing. The repository describes itself as: 🧠 Make your agents learn from experience. Now available as a hosted solution at kayba.ai. The licence is Apache-2.0.

When your agent uses it

  • The user says run the pipeline
  • Evaluate and fix
  • Analyze traces and fix
  • Provides a traces folder with intent to improve their agent

Example prompts

  • “run the pipeline”
  • “kayba pipeline”
  • “evaluate and fix”
  • “/kayba-pipeline”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. sequential
  2. sequential
  3. sequential
  4. HITL Gate
  5. sequential

What it can do on your machine

Read from SKILL.md and the folder at commit 3a31983. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kayba Pipeline loads about 1.4k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 522 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kayba-ai/agentic-context-engine at commit 3a31983, republished under its Apache-2.0 licence (© kayba-ai). 522 words, ~1,420 tokens.

Download SKILL.mdSave it as .claude/skills/kayba-pipeline/SKILL.md (or your agent's skills folder).
name
kayba-pipeline
description
End-to-end agent evaluation and improvement pipeline. Takes a traces folder and optional HITL flag, then orchestrates sub-agents through 7 stages — each stage is its own skill invoked by a dedicated sub-agent. Trigger when the user says "run the pipeline", "kayba pipeline", "evaluate and fix", "full eval", "analyze traces and fix", or provides a traces folder with intent to improve their agent.

kayba-pipeline

End-to-end pipeline: analyze traces → define metrics → build rubric → plan fixes → implement fixes.

Each stage is a separate skill file that can be run independently or as part of this pipeline.

Inputs

The user provides two things:

  1. TRACES_FOLDER — path to a directory containing trace JSON files
  2. HITL — true or false — whether to pause for human review before implementing fixes

If the user doesn't specify HITL, default to true (safe default).


Pipeline overview

┌─────────────────────────────────────────────────────────────────────┐
│  Stage 1: Kayba API Analysis        → skill: kayba-pipeline:stage-1-api-analysis   │
│  Stage 2: Domain Context Gathering  → skill: kayba-pipeline:stage-2-domain-context │
│  ─── stages 1 & 2 run in parallel ───                                              │
│  Stage 3: Metrics & Analysis        → skill: kayba-pipeline:stage-3-metrics        │
│  Stage 4: Rubric Definition         → skill: kayba-pipeline:stage-4-rubric         │
│  Stage 5: Action Plan               → skill: kayba-pipeline:stage-5-action-plan    │
│  Stage 6: HITL Gate                 → skill: kayba-pipeline:stage-6-hitl           │
│  Stage 7: Fix Implementation        → skill: kayba-pipeline:stage-7-fixer          │
└─────────────────────────────────────────────────────────────────────┘

Orchestration instructions

You are the orchestrator. Your job is to:

  1. Create the eval/ directory and eval/pipeline_log.md
  2. Spawn sub-agents that invoke stage skills via the Skill tool
  3. Coordinate stage ordering and handle the HITL gate
Setup

Create eval/ directory and initialize eval/pipeline_log.md:

markdown
# Pipeline Log

| Stage | Name | Status | Started | Completed | Notes |
|-------|------|--------|---------|-----------|-------|
| 1 | Kayba API Analysis | pending | | | |
| 2 | Domain Context | pending | | | |
| 3 | Metrics & Analysis | pending | | | |
| 4 | Rubric Definition | pending | | | |
| 5 | Action Plan | pending | | | |
| 6 | HITL Gate | pending | | | |
| 7 | Fix Implementation | pending | | | |
Stages 1 & 2 — run in parallel

Spawn two sub-agents in parallel using the Agent tool:

Agent 1:

  • Name: api-analyst
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-1-api-analysis" using the Skill tool. The traces folder is: {TRACES_FOLDER}. Follow the skill instructions completely.

Agent 2:

  • Name: domain-scout
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-2-domain-context" using the Skill tool. The traces folder is: {TRACES_FOLDER}. Follow the skill instructions completely.

Wait for both to complete before proceeding.

Stage 3 — sequential

Spawn one sub-agent after stages 1 & 2 complete:

  • Name: metric-engineer
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-3-metrics" using the Skill tool. The traces folder is: {TRACES_FOLDER}. Follow the skill instructions completely — this includes iterating on the metrics until you're satisfied.
Stage 4 — sequential

Spawn one sub-agent after stage 3 completes:

  • Name: rubric-builder
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-4-rubric" using the Skill tool. Follow the skill instructions completely.
Stage 5 — sequential

Spawn one sub-agent after stage 4 completes:

  • Name: action-planner
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-5-action-plan" using the Skill tool. Follow the skill instructions completely.
Show full SKILL.md (233 more words)Show less
Stage 6 — HITL Gate

If HITL is true:

Spawn one sub-agent after stage 5 completes:

  • Name: hitl-reviewer
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-6-hitl" using the Skill tool. Follow the skill instructions completely. Present the full review to the user and collect their decision before proceeding.

Wait for the sub-agent to complete. Check eval/stage6_decision.md for the outcome:

  • If decision is "Approve all" or "Approve with modifications" — proceed to Stage 7
  • If decision is "Reject" — re-run Stage 5 with the user feedback recorded in eval/stage6_decision.md, then re-run Stage 6
  • Only proceed to Stage 7 after a clear approval is recorded

If HITL is false:

  • Skip to Stage 7
  • Log "HITL skipped" in eval/pipeline_log.md
Stage 7 — sequential

Spawn one sub-agent after stage 6 completes (or is skipped):

  • Name: fixer
  • Type: general-purpose
  • Prompt: Invoke the skill "kayba-pipeline:stage-7-fixer" using the Skill tool. Follow the skill instructions completely.

Error handling

  • If any stage fails, log the failure in eval/pipeline_log.md with the stage number and error
  • Do not proceed to dependent stages if a prerequisite failed
  • If Stage 1 fails (kayba CLI issues), ask the user whether to proceed without API insights — if yes, skip Stage 1 and have Stage 3 work from domain context + raw traces only

After completion

Update eval/pipeline_log.md with final status for all stages. Report to the user:

  • How many stages completed successfully
  • Summary of metrics (from rubric)
  • Summary of fixes applied (from changes log)

© kayba-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/kayba-pipeline of kayba-ai/agentic-context-engine.

Open the folder on GitHubat commit 3a31983

Compare with similar skills

Kayba Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kayba Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kayba Pipeline this skillkayba-ai/agentic-context-engine2.6k—~1.4kAutomated safety check: PassApache-2.0
Shipszymdzum/browser-debugger-cli173—~4.4kAutomated safety check: PassMIT
Eval Answermalloydata/publisher116—~4.3kAutomated safety check: PassMIT
Skill Forge EvalAgriciDaniel/skill-forge179—~1.7kAutomated safety check: PassMIT
Diagnosing Superpowers SessionsjnMetaCode/superpowers-zh8.3k—~858Automated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Ship

    szymdzum/browser-debugger-cli

    bdg's shipping workflow: take verified issues from the queue to merged PRs with coding subagents (implementer, fresh reviewer, fresh-agent test before merge), with bdg's commands, gates, conventions…

    173 GitHub stars~4.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Skill Forge Eval

    AgriciDaniel/skill-forge

    Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality.

    179 GitHub stars~1.7k tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check passed
  • Diagnosing Superpowers Sessions

    jnMetaCode/superpowers-zh

    Investigates what went wrong in a superpowers session by reading its transcript, reports findings with path and line citations, and can draft a GitHub issue or redacted bundle.

    8.3k GitHub stars~858 tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from kayba-ai/agentic-context-engine

All 8 skills in this repo
  • Kayba Stage 1 API Analysis

    kayba-ai/agentic-context-engine

    Fetch pre-computed insights from the Kayba API and build a structured summary.

    2.6k GitHub stars~1.1k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 2 Domain Context

    kayba-ai/agentic-context-engine

    Gather domain context about the repository and agent — system prompt, tool definitions, domain docs, and behavior patterns from traces.

    2.6k GitHub stars~1.9k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 3 Metrics

    kayba-ai/agentic-context-engine

    Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.

    2.6k GitHub stars~3.4k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 4 Rubric

    kayba-ai/agentic-context-engine

    Organize computed metrics into a tiered evaluation rubric with leading, lagging, and quality indicators.

    2.6k GitHub stars~2.2k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 5 Action Plan

    kayba-ai/agentic-context-engine

    Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.

    2.6k GitHub stars~2.9k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 6 Hitl

    kayba-ai/agentic-context-engine

    Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.

    2.6k GitHub stars~2.6k tokensUpdated 17 days ago
    Auto-check passed

Categories

Questions about Kayba Pipeline

What does Kayba Pipeline do?

End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine. Kayba Pipeline is an agent skill from kayba-ai/agentic-context-engine. End-to-end agent evaluation and improvement pipeline.

When should I use Kayba Pipeline?

Kayba Pipeline fits situations like: the user says run the pipeline; evaluate and fix; analyze traces and fix; provides a traces folder with intent to improve their agent.

How do I install Kayba Pipeline in Claude Code?

Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-pipeline -a claude-code`. Or copy the skill folder (.claude/skills/kayba-pipeline in kayba-ai/agentic-context-engine) into .claude/skills/kayba-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Kayba Pipeline in Codex?

Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-pipeline -a codex`. Or copy the skill folder (.claude/skills/kayba-pipeline in kayba-ai/agentic-context-engine) into .agents/skills/kayba-pipeline in your project. Codex loads it when a task matches its description.

Can I use Kayba Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kayba-ai/agentic-context-engine --skill kayba-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kayba-pipeline, .gemini/skills/kayba-pipeline, .github/skills/kayba-pipeline and .opencode/skills/kayba-pipeline in your project.

What does Kayba Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Kayba Pipeline is instructions for the agent only.

Does Kayba Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kayba Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kayba Pipeline use?

Kayba Pipeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kayba Pipeline use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kayba Pipeline?

Skills that share tags, products or a category with Kayba Pipeline: Ship (szymdzum/browser-debugger-cli, 173 stars), Eval Answer (malloydata/publisher, 116 stars), Skill Forge Eval (AgriciDaniel/skill-forge, 179 stars) and Diagnosing Superpowers Sessions (jnMetaCode/superpowers-zh, 8.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kayba Pipeline?

kayba-ai (a GitHub organization) maintains it in kayba-ai/agentic-context-engine, which has 2,590 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 24, 2026.

Source: kayba-ai/agentic-context-engine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.