Agent skill

Kayba Stage 6 Hitl

by kayba-ai in kayba-ai/agentic-context-engine

Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.

Apache-2.0Auto-check passedAgent Workflows

Install Kayba Stage 6 Hitl

skills CLI
$ npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kayba-ai/agentic-context-engine kayba-stage-6-hitl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kayba-ai/agentic-context-engine.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kayba-pipeline/stage-6-hitl .claude/skills/kayba-stage-6-hitl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kayba-stage-6-hitl
GitHub stars
2.6k
Token cost
~2.6k tokens
SKILL.md length
877 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.

  • The user says run stage 6
  • SKILL.md covers Inputs, Process, Output format and Rules, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Approve action plan

What it does

Kayba Stage 6 Hitl is an agent skill from kayba-ai/agentic-context-engine. Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome. Trigger when the user says "run stage 6", "HITL review", "approve action plan", or when invoked by the kayba-pipeline orchestrator. Requires eval/actionplan.md and eval/baselinemetrics.md to exist.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Human-in-the-loop approvals. The repository describes itself as: 🧠 Make your agents learn from experience. Now available as a hosted solution at kayba.ai. The licence is Apache-2.0.

When your agent uses it

  • The user says run stage 6
  • Approve action plan
  • Invoked by the kayba-pipeline orchestrator

Example prompts

  • “run stage 6”
  • “HITL review”
  • “approve action plan”
  • “/kayba-stage-6-hitl”

What it can do on your machine

Read from SKILL.md and the folder at commit 3a31983. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kayba Stage 6 Hitl loads about 2.6k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 877 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kayba-ai/agentic-context-engine at commit 3a31983, republished under its Apache-2.0 licence (© kayba-ai). 877 words, ~2,621 tokens.

Download SKILL.mdSave it as .claude/skills/kayba-stage-6-hitl/SKILL.md (or your agent's skills folder).
name
kayba-stage-6-hitl
description
Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome. Trigger when the user says "run stage 6", "HITL review", "approve action plan", or when invoked by the kayba-pipeline orchestrator. Requires eval/action_plan.md and eval/baseline_metrics.md to exist.

Stage 6: Human-In-The-Loop Gate

Present the action plan with enough context for an informed decision, collect the user's approval, and record the outcome.

The goal is not rubber-stamping. The user must receive enough information to genuinely evaluate, modify, or reject the plan -- even if they have not seen Stages 1-5.

Inputs

  • eval/action_plan.md -- the prioritized action plan from Stage 5
  • eval/baseline_metrics.md -- the evaluation rubric with baseline values
  • eval/baseline_metrics.json -- raw metric data (for exact numerator/denominator counts)
  • eval/stage1_insights_summary.md -- original insights (for trace evidence references)

Read all four files before starting.

Process

1. Build the executive summary

Compute and present the following counts from the action plan:

  • Total insights analyzed (raw count before deduplication)
  • Distinct actionable items after deduplication
  • Breakdown: prompt fixes, code fixes, discarded
  • Discard rate with one-line reason per discard (e.g., "5ac7f4ce: efficiency optimization, conflicts with turn discipline constraint")

Format:

EXECUTIVE SUMMARY
-----------------
Insights analyzed:    19 (raw) -> 12 distinct after dedup
Actionable:           9  (8 prompt fixes, 1 code fix)
Discarded:            3  (reasons listed below)

Discards:
  - 5ac7f4ce (Upfront Info Collection): conflicts with higher-priority turn discipline
  - fe2d51cb (Proactive Reservation Lookup): already default behavior, no failure evidence
  - 1fa1b826 (Cancellation Denial Enumeration): subsumed into cancellation checklist
2. Present the top 3 highest-impact changes

For each of the top 3 fixes by priority, present:

Before/after behavior -- use concrete examples from actual traces referenced in the insights. Quote the specific agent behavior that was wrong (before) and describe what the agent should do instead (after). Reference the trace task ID.

Target metric delta -- which metric(s) this fix targets, the current baseline value, and the expected direction. Do not fabricate precise target numbers. Use the format: "M1: 41.4% -> higher (target: 90%+)" only when the action plan provides a target; otherwise use "M1: 41.4% -> up".

Risk rating -- assess each fix:

  • Low -- additive prompt instruction, no behavioral side effects expected
  • Medium -- changes existing behavior, could affect adjacent workflows
  • High -- modifies code/infrastructure, or could degrade a metric while improving another

Format each as a numbered block:

#1: Turn Discipline (covers 55c00c40, d9683144)
    Type:     prompt fix
    Metrics:  M1 (41.4% -> up), M2 (20.7% -> up)
    Risk:     Low

    BEFORE (task_1, task_5, task_7, ...):
      Agent batches 2-3 tool calls per turn (e.g., get_reservation + get_flight_status
      in a single response). Also includes user-facing text alongside tool calls.

    AFTER:
      Exactly one tool call per response. No user-facing content in tool-call turns.
      Agent processes each result before making the next call.
3. Present the full prioritized fix list

Display all non-discarded fixes in a table:

| Priority | Fix Name                          | Type       | Target Metrics  | Risk   | Effort |
|----------|-----------------------------------|------------|-----------------|--------|--------|
| 1        | Turn Discipline                   | prompt fix | M1, M2          | Low    | Low    |
| 2        | Post-Confirmation Execution       | prompt fix | M3              | Low    | Low    |
| 3        | Cancellation Checklist            | prompt fix | M5              | Low    | Low    |
| ...      | ...                               | ...        | ...             | ...    | ...    |

Effort ratings:

  • Low -- single prompt addition, under 5 lines
  • Medium -- multiple prompt additions or minor code change
  • High -- significant code changes, new metric implementation, or architectural changes
4. Present "What we are NOT fixing and why"

List every discarded insight with:

  • Insight ID and name
  • One-line reason for discard
  • What would change your mind (under what conditions should this be revisited)

This section exists so the user can override a discard if they disagree.

5. Flag small-sample and low-confidence items

Any metric with denominator < 5 must be explicitly called out:

LOW-CONFIDENCE METRICS (small sample size):
  - M5 (Cancellation Policy Compliance): based on 2 observations -- directional only
  - M6 (Compensation Execution Rate): based on 1 observation -- directional only

Fixes targeting these metrics (Cancellation Checklist, Compensation Rules) are
still recommended because the policy violations are clear from trace evidence,
but the measured improvement may not be statistically meaningful until the
trace corpus grows.

Also flag any fix where the action plan notes uncertainty or partial evidence.

6. Show the insight-to-fix traceability chain

For each fix, present the chain: insight -> metric -> fix -> expected improvement. This can be a compact list or a table. The purpose is to let the user verify that nothing was lost or invented between stages.

TRACEABILITY:
  55c00c40 (Tool Call Discipline) -> M1, M2 -> Skill 1 (Turn Discipline) -> M1 up, M2 up
  6ea141e1 (Execution Discipline) -> M3 -> Skill 2 (Post-Confirmation) -> M3 up
  0f4a952b + 6ce88ebb (Cancellation) -> M5 -> Skill 3 (Cancellation Checklist) -> M5 up
  ...
7. Collect the decision

Present exactly three options:

OPTIONS:
  [A] Approve all -- implement all 9 fixes as described
  [B] Approve with modifications -- review each fix individually
  [C] Reject -- return to Stage 5 with feedback

Use the appropriate mechanism to collect the user's choice (direct question or AskUserQuestion if available).

If the user selects [A] Approve all

Record the decision and proceed. No further interaction needed.

Show full SKILL.md (407 more words)Show less
If the user selects [B] Approve with modifications

Walk through each fix individually, in priority order. For each fix, present:

  • The fix name, type, and target metrics
  • The recommended prompt/code change (quote the exact text from the action plan)
  • Risk and effort ratings

Then ask: "Approve / Skip / Modify?"

  • Approve -- keep as-is
  • Skip -- remove from the plan, record reason
  • Modify -- ask the user what to change, record the original and the modification

After walking through all fixes, present a summary of changes:

  • Fixes approved as-is: N
  • Fixes skipped: M (list with reasons)
  • Fixes modified: K (list with what changed)

Ask for final confirmation: "Proceed with this modified plan?"

Then update eval/action_plan.md:

  • Remove skipped fixes (move to a "Skipped by HITL" section at the bottom with reasons)
  • Update modified fixes with the user's changes, preserving the original recommendation in a "Original recommendation" sub-field
  • Add a header note: "Modified during HITL review on [date]. See eval/stage6_decision.md for details."
If the user selects [C] Reject

Ask the user for specific feedback:

  • What was wrong with the plan?
  • Which insights or metrics should be reconsidered?
  • Any new constraints or priorities?

Record the feedback in eval/stage6_decision.md and signal that Stage 5 should be re-run with the user's feedback incorporated.

Output format

eval/stage6_decision.md

Write this file regardless of which option was selected.

markdown
# Stage 6: HITL Decision Record

## Date
[timestamp]

## Decision
[Approve all | Approve with modifications | Reject]

## What was presented
- Total insights: N (M distinct after dedup)
- Actionable fixes: X (Y prompt, Z code)
- Discarded: W
- Metrics: [list metric IDs and baselines]
- Low-confidence flags: [list metrics with small denominators]

## Top 3 changes presented
1. [fix name] -- [type] -- targets [metrics] -- risk [rating]
2. ...
3. ...

## Decision details

### If Approve all:
User approved all N fixes without modification.
Reasoning: [any reasoning the user provided, or "No additional reasoning provided"]

### If Approve with modifications:
| Fix | Original Status | Decision | Reason |
|-----|----------------|----------|--------|
| Turn Discipline | Priority 1 | Approved | -- |
| Compensation Rules | Priority 5 | Modified | User changed wording to... |
| Cabin Change Rules | Priority 8 | Skipped | User considers low priority |

Modifications detail:
- [Fix name]: Original: "..." -> Modified: "..." -- User rationale: "..."

### If Reject:
User feedback: [verbatim feedback]
Specific concerns: [list]
Re-run instructions for Stage 5: [what to change]

## Traceability snapshot
[Copy of the traceability chain from step 6, so the decision record is self-contained]
eval/action_plan.md (updated, only if modifications were made)

If the user selected [B] and made changes:

  • Add a modification header at the top of the file
  • Update individual fix entries with user changes
  • Move skipped fixes to a "Skipped by HITL" section
  • Preserve original recommendations as sub-fields for auditability

Rules

  • Do NOT auto-approve. The entire point of this stage is human judgment.
  • Do NOT summarize so aggressively that the user cannot evaluate. When in doubt, include more context.
  • Do NOT proceed to Stage 7 until a clear approval (full or modified) is recorded.
  • Do NOT modify eval/action_plan.md unless the user explicitly requests modifications.
  • Do NOT skip the small-sample warnings. If M5 has denominator 2 and M6 has denominator 1, the user must know this.
  • Do NOT fabricate target metric values. Use targets from the action plan when available; otherwise state direction only.
  • Always present the "What we are NOT fixing" section. Omitting discards hides information the user needs.
  • If the user asks clarifying questions, answer them fully before re-presenting the decision options.

Outputs

  • eval/stage6_decision.md -- full record of what was presented, decided, and why
  • eval/action_plan.md -- updated only if the user selected "Approve with modifications"

© kayba-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/kayba-pipeline/stage-6-hitl of kayba-ai/agentic-context-engine.

Open the folder on GitHubat commit 3a31983

Compare with similar skills

Kayba Stage 6 Hitl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kayba Stage 6 Hitl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kayba Stage 6 Hitl this skillkayba-ai/agentic-context-engine2.6k—~2.6kAutomated safety check: PassApache-2.0
Show Me Your Work Decision Logcursor/plugins11k8 repos~1.6kAutomated safety check: PassNone
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Loop Constraints Enforcercobusgreyling/loop-engineering11k1 repos~475Automated safety check: NotesMIT
Ask User QuestionMemTensor/MemOS12k—~1kAutomated safety check: PassApache-2.0
PUA High-Agency Governancetanweai/pua20k—~502Automated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    11k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Loop Constraints Enforcer

    cobusgreyling/loop-engineering

    Loads a project's loop-constraints.md before any other action and blocks pushes, edits or merges that violate the rules it defines.

    11k GitHub starsUsed in 1 repo~475 tokens
    Agent WorkflowsAuto-check: notes
  • Ask User Question

    MemTensor/MemOS

    Shows a question as a modal in the interface to clarify a task, collect a preference or get approval, since the user cannot see terminal output.

    12k GitHub stars~1k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Pushes an agent to keep verifying and changing approach after repeated failures, using a diagnosis line, evidence-based completion and confirmation before risky edits.

    20k GitHub stars~502 tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Agentmemory Forget

    rohitg00/agentmemory

    Deletes chosen memories from agentmemory only after showing the matches and getting an explicit yes, for privacy requests and cleanup of outdated notes.

    29k GitHub stars~612 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from kayba-ai/agentic-context-engine

All 8 skills in this repo
  • Kayba Pipeline

    kayba-ai/agentic-context-engine

    End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.

    2.6k GitHub stars~1.4k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 1 API Analysis

    kayba-ai/agentic-context-engine

    Fetch pre-computed insights from the Kayba API and build a structured summary.

    2.6k GitHub stars~1.1k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 2 Domain Context

    kayba-ai/agentic-context-engine

    Gather domain context about the repository and agent — system prompt, tool definitions, domain docs, and behavior patterns from traces.

    2.6k GitHub stars~1.9k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 3 Metrics

    kayba-ai/agentic-context-engine

    Define metrics from Kayba insights, implement them as Python measurement code, run against traces, and iterate until the metrics are clean and meaningful.

    2.6k GitHub stars~3.4k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 4 Rubric

    kayba-ai/agentic-context-engine

    Organize computed metrics into a tiered evaluation rubric with leading, lagging, and quality indicators.

    2.6k GitHub stars~2.2k tokensUpdated 17 days ago
    Auto-check passed
  • Kayba Stage 5 Action Plan

    kayba-ai/agentic-context-engine

    Triage each insight into discard/code-fix/prompt-fix and produce a prioritized action plan with specific recommendations.

    2.6k GitHub stars~2.9k tokensUpdated 17 days ago
    Auto-check passed

Categories

Questions about Kayba Stage 6 Hitl

What does Kayba Stage 6 Hitl do?

Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome. Kayba Stage 6 Hitl is an agent skill from kayba-ai/agentic-context-engine. Human-In-The-Loop gate that presents the action plan with full context, collects an informed approval/modification/rejection decision, and records the outcome.

When should I use Kayba Stage 6 Hitl?

Kayba Stage 6 Hitl fits situations like: the user says run stage 6; approve action plan; invoked by the kayba-pipeline orchestrator.

How do I install Kayba Stage 6 Hitl in Claude Code?

Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a claude-code`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-6-hitl in kayba-ai/agentic-context-engine) into .claude/skills/kayba-stage-6-hitl in your project. Claude Code loads it when a task matches its description.

How do I install Kayba Stage 6 Hitl in Codex?

Run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a codex`. Or copy the skill folder (.claude/skills/kayba-pipeline/stage-6-hitl in kayba-ai/agentic-context-engine) into .agents/skills/kayba-stage-6-hitl in your project. Codex loads it when a task matches its description.

Can I use Kayba Stage 6 Hitl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kayba-ai/agentic-context-engine --skill kayba-stage-6-hitl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kayba-stage-6-hitl, .gemini/skills/kayba-stage-6-hitl, .github/skills/kayba-stage-6-hitl and .opencode/skills/kayba-stage-6-hitl in your project.

What does Kayba Stage 6 Hitl need to run?

SKILL.md names no scripts, command-line tools or credentials: Kayba Stage 6 Hitl is instructions for the agent only.

Does Kayba Stage 6 Hitl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kayba Stage 6 Hitl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kayba Stage 6 Hitl use?

Kayba Stage 6 Hitl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kayba Stage 6 Hitl use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kayba Stage 6 Hitl?

Skills that share tags, products or a category with Kayba Stage 6 Hitl: Show Me Your Work Decision Log (cursor/plugins, 11k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Loop Constraints Enforcer (cobusgreyling/loop-engineering, 11k stars) and Ask User Question (MemTensor/MemOS, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kayba Stage 6 Hitl?

kayba-ai (a GitHub organization) maintains it in kayba-ai/agentic-context-engine, which has 2,590 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 24, 2026.

Source: kayba-ai/agentic-context-engine on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.