Agent skill

Kayba Stage 5 Action Plan

by kayba-ai in kayba-ai/agentic-context-engine

Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.

Apache-2.0Auto-check passedAgent Workflows

Install Kayba Stage 5 Action Plan

skills CLI
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-5-action-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-5-action-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-5-action-plan .claude/skills/kayba-stage-5-action-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kayba-stage-5-action-plan
GitHub stars
2.6k
Token cost
~2.9k tokens
SKILL.md length
1,177 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.

  • The user says run stage 5
  • SKILL.md covers Inputs, Process, Output format and Outputs
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Make action plan

What it does

Kayba Stage 5 Action Plan is an agent skill from kayba-ai/agentic-context-engine. Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations. Trigger when the user says "run stage 5", "make action plan", "triage skills", or when invoked by the kayba-pipeline orchestrator. Requires eval outputs from stages 1-4.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. The repository describes itself as: 🧠 Make your agents learn from experience. Now available as a hosted solution at kayba.ai. The licence is Apache-2.0.

When your agent uses it

  • The user says run stage 5
  • Make action plan
  • Invoked by the kayba-pipeline orchestrator

Example prompts

  • “run stage 5”
  • “make action plan”
  • “triage skills”
  • “/kayba-stage-5-action-plan”

What it can do on your machine

Read from SKILL.md and the folder at commit 3a31983. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kayba Stage 5 Action Plan loads about 2.9k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,177 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kayba-ai/agentic-context-engine at commit 3a31983, republished under its Apache-2.0 licence (© kayba-ai). 1,177 words, ~2,850 tokens.

Download SKILL.mdSave it as .claude/skills/kayba-stage-5-action-plan/SKILL.md (or your agent's skills folder).
name
kayba-stage-5-action-plan
description
Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations. Trigger when the user says "run stage 5", "make action plan", "triage skills", or when invoked by the kayba-pipeline orchestrator. Requires eval outputs from stages 1-4.

Stage 5: Action Plan

Triage each insight and produce a concrete, prioritized action plan.

Inputs

  • eval/stage1_insights_summary.md — insights from Kayba
  • eval/stage2_domain_context.md — domain context
  • eval/baseline_metrics.md — the evaluation rubric
  • eval/baseline_metrics.json — baseline values
  • eval/compute_baselines.py — measurement code

Read all files before starting.

Process

1. Triage each insight

For each insight/skill, answer three questions in order: Is it valid? Is it already handled? Is it a code fix or prompt fix?

1a. Validity check
  • Does it describe a real, recurring problem visible in traces — or noise from a one-off edge case?
  • Is it actionable — can the agent actually change this behavior given its tools and context?
  • If not valid → verdict: discard with a one-sentence reason.
1b. "Already handled" verification

Do not rely on memory or assumption. Run these checks and cite what you find:

  1. Grep the codebase for 2-3 key terms from the insight (tool names, error strings, behavioral keywords). Example: for an insight about cancellation eligibility, grep for cancel, eligibility, criteria.
  2. Read the existing system prompt text — check AGENT_INSTRUCTION in the agent file and the domain policy file. Quote any existing language that addresses this behavior.
  3. Verdict:
    • If existing text partially covers it → keep as a strengthening fix, note what's missing.
    • If no existing coverage → keep.
    • If existing prompt text already covers the behavior thoroughly AND the baseline metric is >= 95% → discard (cite the existing text and metric). A high baseline alone is NOT sufficient to discard — if the metric is below 95%, there are still failures to fix. An 87% baseline means 1 in 8 attempts still fails; that is worth fixing.
1c. Code-vs-prompt decision tree

Walk through this tree for every non-discarded insight:

Q1: Can the agent fix this by following different instructions?
    (Does it have the right tools, correct data in tool responses,
     and sufficient context to behave correctly?)
  │
  ├─ YES → PROMPT FIX
  │        The agent has everything it needs but acts wrong.
  │        A system prompt addition would fix it.
  │
  └─ NO → Q2: What is the agent missing?
           │
           ├─ Tool doesn't exist, schema is wrong, API returns
           │  incomplete data, infrastructure drops information,
           │  timeout/error not surfaced to agent
           │  → CODE FIX
           │    Name the file, function, and specific change.
           │
           └─ The agent has partial information but the prompt
              can't fully compensate (e.g., needs a new tool
              but a heuristic prompt workaround exists)
              → PROMPT FIX (primary) + CODE FIX (optional)
                Note both. Mark the code fix as "optional" with
                a one-sentence justification for why it's lower priority.

Ambiguity default: When genuinely uncertain, default to prompt fix and add a note: "Classification uncertain — defaulting to prompt fix. Revisit if prompt change doesn't move metrics." This is safer because prompt fixes are cheaper to test and revert, and Stage 7 handles prompt fixes and code fixes through different paths.

Use the reflector's reasoning from Stage 1 insights — it often explicitly identifies root causes that clarify the code-vs-prompt distinction.

Before writing recommendations, merge insights that are redundant. Two insights should merge when ALL three conditions hold:

  1. Same target behavior — they describe the agent doing (or failing to do) the same thing.
  2. Overlapping fix text — the prompt instructions you'd write for each would share >50% of their content.
  3. Addressing one substantially addresses the other — fixing insight A would fix >80% of the cases described by insight B.

When NOT to merge — two insights about the same tool or domain area but different failure modes should remain separate. Example: "agent doesn't check cancellation eligibility" and "agent doesn't execute cancellation after user confirms" both involve cancel_reservation but are completely different behavioral failures with different prompt fixes. Keep them separate.

For each merge, document:

  • Which insight IDs are combined
  • Which insight's framing is primary (use the one with stronger trace evidence)
  • What, if anything, is lost from the secondary insight (add it as a sub-point)
3. Write specific recommendations

For each insight (after merging):

  • Discards: one sentence on why it's not valid or actionable.
  • Code fixes: what code/schema/infrastructure to change. Name the file, the function, the specific change. If Stage 7 needs to find the right code location, give it enough to grep for.
  • Prompt fixes: the exact instruction text to add to the system prompt, where it should go (e.g., appended to AGENT_INSTRUCTION, added to domain policy, or as a standalone skill block), and why this wording over alternatives.
4. Assess risk per fix

For each non-discarded fix, assess whether the change could break currently-working behaviors:

RiskDefinitionExample
NoneChange is additive; no existing behavior could be affectedAdding a new metric to compute_baselines.py
LowChange targets a behavior that is currently failing; working cases are unrelatedAdding a cancellation checklist when current cancellation compliance is 0%
MediumChange modifies a behavior where some cases already work correctlyStrengthening confirmation protocol when 28.6% already succeed — could the new wording break the working 28.6%?
HighChange rewrites or constrains a behavior that mostly worksRestricting tool-call patterns when 41.4% already comply — overly rigid wording could cause the agent to under-call tools

For Medium and High risk fixes, add a one-sentence mitigation: what to watch for, or how to word the prompt to preserve working cases.

Show full SKILL.md (455 more words)Show less
5. Handle qualitative-only insights — STILL PRODUCE FIXES

Some insights from Stage 3 may be flagged as "unmeasurable." These still get fixes. An insight that the agent fabricates data or violates policy is a real problem whether or not we can measure it programmatically. Treat them the same as any other insight:

  • Run the same triage (validity → already-handled → code-vs-prompt) as every other insight.
  • Include them in the priority-ranked implementation list alongside all other fixes. They are NOT second-class.
  • Use the trace evidence from the insight (not the metric) to assess impact and priority. If the insight has strong trace evidence showing clear failures, rank it accordingly.
  • For prioritization: since there is no metric denominator, use confidence = 0.5 and estimate impact from the severity described in the insight evidence.
  • In the fix entry, note that this fix has no programmatic metric for automated before/after comparison, so improvement should be verified via manual trace review or LLM-as-judge after generating new traces.

Only relegate an insight to a non-actionable "Monitor Items" section if the triage concludes it should be discarded (not valid or not actionable). Being unmeasurable is NOT a reason to skip fixing it.

For each non-discarded fix, identify which metric(s) from the rubric would move if this fix is implemented. Use the metric IDs from eval/baseline_metrics.md (e.g., M1, M2).

7. Prioritize

Rank non-discarded fixes using this formula:

Priority Score = Impact × Confidence × Tier Bonus ÷ Risk Factor

Where:

  • Impact = estimated metric delta. Use the gap between baseline and 100% as the ceiling. A fix expected to close 50% of that gap on M1 (baseline 41.4%) has impact = 0.5 × (1.0 - 0.414) = 0.293.
  • Confidence = sample size reliability. Use the denominator from baseline_metrics.json:
    • denominator >= 20: confidence = 1.0
    • denominator 10-19: confidence = 0.8
    • denominator 5-9: confidence = 0.6
    • denominator < 5: confidence = 0.3
  • Tier Bonus = leading metrics get a 1.5x multiplier (they validate adoption), lagging and quality get 1.0x. Rationale: leading metrics move first and tell you if your fix is even being adopted — you want those signals early.
  • Risk Factor = None: 1.0, Low: 1.0, Medium: 1.5, High: 2.0

You do not need to compute exact scores to three decimal places. The formula is a tiebreaker and sanity check. The point is:

  • High-impact, high-confidence, leading-metric fixes with low risk go first.
  • Low-confidence fixes (small denominators) get deprioritized even if the metric is at 0%.
  • High-risk fixes get deprioritized unless impact is overwhelming.

After scoring, apply one manual adjustment pass: if a fix is a prerequisite for another fix (e.g., "confirmation protocol" must exist before "post-confirmation execution" can be measured), promote the prerequisite even if its standalone score is lower.

Output format

Write to eval/action_plan.md:

markdown
# Action Plan

## Summary
- Total insights: N
- Discarded: X (with reasons)
- Code fixes: Y
- Prompt fixes: Z
- Fixes without programmatic metric (verify manually): Q

## Implementation Priority
| Rank | Fix | Type | Metrics | Risk | Score rationale |
|------|-----|------|---------|------|-----------------|
| 1 | [name] | prompt | M1, M2 | Low | [one-line: why this ranks here] |
| 2 | ... | ... | ... | ... | ... |

---

## Skill: [insight ID(s)] — [title]
**Summary:** [one-line description of what the skill addresses]
**Verdict:** `prompt fix` | `code fix` | `discard`
**Classification path:** [which branch of the decision tree — e.g., "Agent has tools and data but acts wrong → prompt fix"]
**Rationale:** [why this verdict — reference specific trace evidence from insights]
**Risk:** None | Low | Medium | High — [one-sentence justification]
**Risk mitigation:** [for Medium/High only — what to watch for or how to preserve working cases]
**Recommendation:** [specific change to make]
**Files to modify:** [list of files, for code fixes]
**Metric link:** [which metrics would move, with baseline values]
**Already-handled check:** [what you grepped, what existing prompt text you found, verdict]

---
[repeat for each insight]

## Consolidated Prompt Skills
[After all per-insight entries, list the final merged prompt skill texts in priority order, ready for Stage 7 to implement]

## Monitor Items (Non-Actionable Only)
[Only insights that were triaged as genuinely non-actionable — e.g., the agent cannot change this behavior, or the insight is noise. Unmeasurable insights that are still real problems should appear in the priority list above, NOT here.]

Group related insights under cluster headings when they address the same underlying behavior. For merged insights, list all constituent insight IDs in the heading.

Outputs

  • eval/action_plan.md

© kayba-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/kayba-pipeline/stage-5-action-plan of kayba-ai/agentic-context-engine.

Open the folder on GitHubat commit 3a31983

Compare with similar skills

Kayba Stage 5 Action Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kayba Stage 5 Action Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kayba Stage 5 Action Plan this skillkayba-ai/agentic-context-engine2.6k—~2.9kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k36 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79689 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 36 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    796 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from kayba-ai/agentic-context-engine

All 8 skills in this repo
  • Kayba Pipeline

    kayba-ai/agentic-context-engine

    End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.

    2.6k GitHub stars~1.4k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 1 API Analysis

    kayba-ai/agentic-context-engine

    Fetch pre-computed insights from the Kayba API and build a structured summary.

    2.6k GitHub stars~1.1k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 2 Domain Context

    kayba-ai/agentic-context-engine

    Gather domain context about the repository and agent — system prompt, tool definitions, domain docs, and behavior patterns from traces.

    2.6k GitHub stars~1.9k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 3 Metrics

    kayba-ai/agentic-context-engine

    Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.

    2.6k GitHub stars~3.4k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 4 Rubric

    kayba-ai/agentic-context-engine

    Organize computed metrics into a tiered evaluation rubric with leading, lagging, and quality indicators.

    2.6k GitHub stars~2.2k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 6 Hitl

    kayba-ai/agentic-context-engine

    Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.

    2.6k GitHub stars~2.6k tokensUpdated 17 days ago
    Auto-check passed

Categories

Questions about Kayba Stage 5 Action Plan

What does Kayba Stage 5 Action Plan do?

Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations. Kayba Stage 5 Action Plan is an agent skill from kayba-ai/agentic-context-engine. Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.

When should I use Kayba Stage 5 Action Plan?

Kayba Stage 5 Action Plan fits situations like: the user says run stage 5; make action plan; invoked by the kayba-pipeline orchestrator.

How do I install Kayba Stage 5 Action Plan in Claude Code?

Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-5-action-plan -a claude-code`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-5-action-plan in kayba-ai/agentic-context-engine) into .claude/skills/kayba-stage-5-action-plan in your project. Claude Code loads it when a task matches its description.

How do I install Kayba Stage 5 Action Plan in Codex?

Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-5-action-plan -a codex`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-5-action-plan in kayba-ai/agentic-context-engine) into .agents/skills/kayba-stage-5-action-plan in your project. Codex loads it when a task matches its description.

Can I use Kayba Stage 5 Action Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-5-action-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kayba-stage-5-action-plan, .gemini/skills/kayba-stage-5-action-plan, .github/skills/kayba-stage-5-action-plan and .opencode/skills/kayba-stage-5-action-plan in your project.

What does Kayba Stage 5 Action Plan need to run?

SKILL.md names no scripts, command-line tools or credentials: Kayba Stage 5 Action Plan is instructions for the agent only.

Does Kayba Stage 5 Action Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kayba Stage 5 Action Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kayba Stage 5 Action Plan use?

Kayba Stage 5 Action Plan is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kayba Stage 5 Action Plan use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kayba Stage 5 Action Plan?

Skills that share tags, products or a category with Kayba Stage 5 Action Plan: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kayba Stage 5 Action Plan?

kayba-ai (a GitHub organization) maintains it in kayba-ai/agentic-context-engine, which has 2,590 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 24, 2026.

Source: kayba-ai/agentic-context-engine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.